Week 22. Infrastructure and research craft
Learn in the app: tutor, coding problems →
Core: torch.profiler and MFU, reproducibility and seed noise, reading papers · Depth: expressivity (DFA, the Dyck language), interpretability and alignment, tools (git, Docker, SLURM), Gaussian processes (05-ГЛУБИНА) · ≈ 10 h core / 21 h total
This week is about iteration speed. A researcher who runs ten experiments a day overtakes one who runs a single experiment, even if the second is smarter. Speed comes from three things: measure, do not guess (profiling), trust your numbers (reproducibility) and drop dead ideas in time (research taste).
Part 1. Profiling
In plain terms. A training step takes 100 ms. The profile: for 50 ms the GPU waits for the next batch from the dataloader.
For another 30 ms it sits idle because of loss.item() on every step: the CPU waits for the computation to finish
and only then queues the next kernels. Forward, backward and the optimizer take only 20 ms.
For 80% of the time the GPU does nothing useful. Make the matmuls twice as fast: the step becomes 90 ms, a 10% gain.
Remove the sync and hide loading behind compute: the step is about 20 ms, five times faster.
Speed up what takes the time, not what is easiest to speed up.
- The first question is always the same: is the GPU computing or waiting? Three typical bottlenecks: the dataloader (the CPU cannot prepare batches fast enough), CPU–GPU syncs (the CPU stands and waits until the GPU finishes its queue), small memory-bound kernels (limited by memory reads, week 13)
- CUDA is asynchronous.
time.time()withouttorch.cuda.synchronize()measures queuing kernels, not running them..item(),.cpu(),print(loss)on every step create hidden syncs torch.profiler:record_functionto mark phases, a table by self time, a trace in Perfetto (a browser-based trace viewer).schedule(wait, warmup, active)skips the first steps: those are about the allocator, warmup and compilation, not the steady state- How to read a trace: gaps between kernels on the GPU track = the GPU is idle → the bottleneck is on the CPU. A dense GPU track → look at the top kernels
- Fixes in order of cost, cheapest first: remove syncs →
num_workers,pin_memory→ bf16 autocast →torch.compile(fusing elementwise kernels, week 13) → a larger batch
In plain terms. A 124M-parameter model trains at 200 000 tokens/s.
Useful work: 6 · 1.24·10⁸ · 2·10⁵ ≈ 1.5·10¹⁴ FLOP/s. The H100 bf16 peak is about 9.9·10¹⁴.
MFU = 15%: 85% of the hardware's capacity is not being used, and the profile will tell you where it went.
- MFU (model FLOPs utilization) =
6N · tokens/s/ peak FLOPs (week 9). One number that shows how much performance is still on the table
Part 2. Reproducibility
In plain terms. A baseline run on three seeds gives val loss 3.21, 3.25 and 3.23, a spread of 0.04. A new method on one seed gives 3.22. It is "better" than the mean by 0.01, but that is within the spread: another seed could just as well give 3.26. Measure the noise first, then compare.
- A seed does not pin everything down: nondeterministic kernels (atomic additions in a different order give different rounding), data
order with several workers.
torch.use_deterministic_algorithms(True)fixes this at the cost of speed - Seed noise is measured before comparing. A difference smaller than the spread across three seeds does not count as a result (week 18)
- In every log: the config, the commit hash, the data version. wandb or an equivalent
git bisectfor quality regressions: binary search over commits with a check script. Docker pins the environment; tmux, ssh and SLURM make up the cluster working environment
Part 3. Research craft
- Reading papers in three passes: (1) abstract, figures, conclusions: 10 minutes to decide whether to read further; (2) method and experiments: what they compare against, and whether it is fair; (3) derive the key result yourself and find the weak spot
- Three questions for any paper. A made-up example: a paper called "LowKV" promises to compress the KV cache 8 times losslessly,
"as guaranteed by the Eckart–Young theorem".
- Are the theorem's conditions met? The theorem says that truncated SVD gives the best approximation of the key matrix at a given rank,
not that this approximation is good. If rank 16 out of 128 keeps 70% of the spectral energy (the sum of squared singular values),
"lossless" does not follow from any theorem. Moreover, a small error in
Kcan grow after softmax - Do the baselines get the same information? LowKV sees 32k tokens, while the baseline is truncated to 4k to fit into the same memory. They compared input length, not compression method. An honest baseline: the same 32k, compressed differently (GQA, cache quantization), at equal memory
- Is there prior work? A low-rank latent instead of the full KV already exists: MLA from DeepSeek-V2 (week 11). So the novelty lies elsewhere, and the paper has to say where
- Are the theorem's conditions met? The theorem says that truncated SVD gives the best approximation of the key matrix at a given rank,
not that this approximation is good. If rank 16 out of 128 keeps 70% of the spectral energy (the sum of squared singular values),
"lossless" does not follow from any theorem. Moreover, a small error in
- Choosing a problem: importance × chance of success × your advantage. A cheap early signal matters more than an elegant setup
- The kill criterion for an idea is written down in advance: "if the effect is smaller than the noise at two scales, I close it". Without it, an idea lives until your patience runs out
Part 4. Expressivity (briefly)
- Regular and context-free languages, DFA (deterministic finite automaton). An RNN with finite precision
is essentially a finite automaton: a DFA can be encoded as a ReLU-RNN (the state is stored as one-hot, and each transition is given
by a matrix per symbol). Dyck languages (balanced bracket sequences) serve as a test of hierarchy,
which an automaton lacks: to check depth
k, you need to rememberkopen brackets
Part 5. How to look inside a model and what alignment is (extension)
Research rounds increasingly ask not only "how do you train it" but also "how do you tell what the model has learned"
and "how do you make it do what was meant". Below is a minimal toolkit: each tool
can be tried in code on nanolm in an evening.
In plain terms. A model with 4 layers and a vocabulary of three tokens: "4", "5", "number". Take the residual stream (week 6) after each layer at the last position of the phrase "Two plus two equals", pass it through the final RMSNorm and the output matrix as if the model ended at that layer, and look at the top-1 token. Layer 1: "number", layer 2: "number", layer 3: "4" with probability 0.4, layer 4: "4" with 0.9. You can see at which layer the answer appeared and how the confidence grew. This is the logit lens.
- Logit lens. The hidden state of layer
l, of shape[D], is normalized with the final RMSNorm (with its weightγ) and multiplied by the output matrixW_Uof shape[D, V]:logits_l = RMSNorm(h_l) · W_U. Without the final normalization the numbers at early layers are not comparable with the model's output. A limitation: early layers live in their own "basis", and the logit lens often shows garbage there; the tuned lens (Belrose et al., 2023) learns a small linear map per layer and reads early layers more reliably - Linear probe (logistic regression on a layer's activations). The question: does the layer contain a feature, for example "the sentence is in the past tense"? Collect the layer's activations on labeled examples, train a logistic regression on one part and measure accuracy on the held-out part. In plain terms. 200 examples, a layer of dimension 64. A probe on layer 6 gives 0.95 on the held-out part, on layer 1 it gives 0.55 with a class share of 0.5: the feature is linearly readable at layer 6. Accuracy on the training part proves nothing: with 60 examples and 64 dimensions logistic regression will fit even random labels. And even high held-out accuracy does not mean the model uses the feature: that is checked by intervention (remove the feature direction and see whether the output changes)
- SAE (sparse autoencoder). Activations of dimension
Dare decomposed over a dictionary ofM ≫ Ddirections so that only a few are active at each point:f = ReLU(W_enc · (x − b) + b_enc),x̂ = W_dec · f + b, loss‖x − x̂‖² + λ‖f‖₁. In plain terms.D = 512, a dictionary ofM = 16 384, and on one token about 20 of the 16 384 features are active. A single feature is often readable by a human ("Python code", "mention of the Golden Gate Bridge"), while a single neuron usually mixes several concepts (it is polysemantic). The price is reconstruction error and features nobody can name - Steering (controlling activations). A feature direction (a column of the SAE decoder, or the difference of mean
activations on two sets of prompts, for example "happy" minus "sad") is multiplied by a coefficient and added
to the residual stream at one layer during generation:
h ← h + α·v. In plain terms. Atα = 0the model writes neutrally, atα = 4noticeably happier, atα = 20the text falls apart. Choosingαand the layer is an experiment with a metric, not tuning by eye - Sycophancy (the model agrees with the user against the facts). It is measured with pairs of prompts: the same question without the user's opinion and with it ("I think the answer is B, what do you think?"). In plain terms. 200 questions with known answers, and the model is right on 160 without an opinion. With a wrong user opinion it is right on 110: 50 answers out of 160 flipped, sycophancy is 31%. A second variant: the model answered correctly, the user writes "Are you sure? I don't think so", and you count the share of answers the model changed to wrong ones. One of the causes: human annotators in RLHF more often prefer answers that agree with them, and the reward model learns this preference
- RLAIF and constitutional AI. RLAIF (RL from AI feedback): preferences between pairs are labeled by a model, not a human. Constitutional AI: the model is given a list of principles (a "constitution", for example "choose the less harmful and more honest answer"), it critiques and rewrites its own answers according to these principles (the SFT stage), and then compares pairs of answers itself by the same principles, and a reward model for RL is trained on these judgments. The gain: cheap, scalable, the principles are written down explicitly and can be checked. The price: the judge model carries its own biases into the data (length, the self-preference of week 18)
- The link with the reward hacking of week 17. Sycophancy is reward hacking: the policy found that agreeing with the user raises the reward, while the truth was what was meant. With RLAIF it is the same: the policy optimizes the judge model's score, not the principles. Interpretability offers a way to catch this: a probe or an SAE feature for "the model knows the right answer" firing on an answer that contradicts it shows a gap between what the model "knows" and what it says
In the trainer these are the logit_lens problem (the top-1 token per layer through the final RMSNorm and W_U) and linear_probe
(logistic regression on activations, accuracy on the held-out part).
Code → scripts/profile_train.py: nanolm training steps under torch.profiler
after warmup, with phases marked by record_function (forward, backward, optimizer);
the output is the top operations by self time, grouped by category,
and a plain-words verdict on "where the bottleneck is". Task: compute the MFU of your run with the formula
above, make one change, repeat the profile, and record the result in a
"before / after / why" table. On CPU and MPS the script profiles the CPU and says so explicitly.
Math (Track D): D35: gambler's ruin two ways; D36: stars and bars and inclusion-exclusion.
Interview question of the week: "Training is 3 times slower than the estimate. What do you do?"
Structure: (1) where the expectation comes from: 6N × tokens / (peak × MFU); (2) isolate:
the dataloader alone, the model on synthetic data; (3) the profiler: idle gaps,
syncs, top kernels; (4) one hypothesis → one change → one measurement; (5) what
remains and why (communication, if training is distributed).
Sources: PyTorch Profiler documentation; Keshav, How to Read a Paper (2007); Hamming, You and Your Research (1986). For part 5: nostalgebraist, interpreting GPT: the logit lens (2020); Belrose et al., Tuned Lens (2023); Alain, Bengio, Understanding intermediate layers using linear classifier probes (2016); Bricken et al., Towards Monosemanticity (2023); Templeton et al., Scaling Monosemanticity (2024); Turner et al., Activation Addition (2023); Sharma et al., Towards Understanding Sycophancy in Language Models (2023); Bai et al., Constitutional AI: Harmlessness from AI Feedback (2022). Links are in 04-РЕСУРСЫ, section "Interpretability and alignment".
Deeper: 05-ГЛУБИНА, section "Weeks 19 and 22. Gaussian process regression".
Week outcomes
- I can run
torch.profileron someone else's training loop and name the bottleneck within 15 minutes, with evidence from the trace. - I can compute the MFU of my training run and explain where the rest goes.
- I can make an experiment deterministic and measure seed noise before comparing methods.
- I can read a paper in three passes and state its contribution and weak spot in 5 minutes.
- I can apply the logit lens and a linear probe to my model and explain why a probe is measured on a held-out part and why high accuracy does not yet prove the model uses the feature.
- I can measure sycophancy with pairs of prompts and explain how RLAIF and constitutional AI replace annotators and why sycophancy is reward hacking.
Self-check
- Why does timing without
torch.cuda.synchronize()lie, and in which direction? - How do you tell CPU-bound training from GPU-bound training in a trace?
- Why can an RNN with finite precision not recognize a Dyck language of arbitrary depth?
- Why does the logit lens apply the final RMSNorm before the output matrix? A linear probe on a layer gave 0.98 on the training part and 0.52 on the held-out part with a class share of 0.5: what does that mean?
- How do you measure sycophancy with a pair of prompts, and what does it have to do with the reward hacking of week 17? How does constitutional AI differ from RLHF with human annotators?
Mock interviews of the week (7 and 8 of 12). (7) ML debugging, session E: the partner plants 5–6 bugs in your transformer block; without a partner, do the three "Find and fix" problems in the trainer. (8) Rapid-fire, session A: 20 questions on weeks 19–22 and 10 from earlier weeks; it is also item 5 of checkpoint 5.
✅ Checkpoint 5
No hints, out loud and recorded, about two hours in one day:
- ML system design in 45 minutes with rubric 7: a RAG assistant or an agent, a new prompt. Threshold: at least 6 on every criterion and at least 7 on "Data and quality evaluation"
- In 10 minutes, compare the transformer, Mamba and MoE: what each pays for context and memory, with numbers
(an
O(S)cache versus anO(1)state, total and active MoE parameters) - In 10 minutes: how an image gets into an LLM, and why good perplexity on a long context does not yet prove that the model uses it
- In 15 minutes, from your own
torch.profilertrace, name the bottleneck, one change and a before/after measurement; seed noise measured before the comparison - Rapid-fire (mock 8 counts): at least 20 ✓ and at most 3 ✗
Failed an item? Go back to its week before the week 23 sprint.