Skip to content
Snula
Curriculum
RU Open

Curriculum

Week 15. SFT and data

Phase 4. Post-training · week 15 of 24

Learn in the app: tutor, coding problems →

Core: SFT and loss masking, chat templates, LoRA and QLoRA, data deduplication and filtering · Depth: distillation and pruning, the data lab of week 15b · ≈ 11 h core / 21 h total

From this week on, we no longer pretrain the model but refine it: a pretrained LLM continues text rather than answering questions. SFT is the first stage of post-training, and the RL of weeks 16–17 starts from the SFT model. This week also covers two tools without which fine-tuning does not fit the budget and a benchmark proves nothing (week 18): LoRA, and clean data with decontamination.

  • Instruction tuning (SFT, fine-tuning on "instruction → response" pairs), chat templates (the format that glues a dialogue into one string), special role tokens (they mark where the user's turn is and where the assistant's is)

In plain terms. Example: <user> What is 2 + 2? <assistant> 4 <end>. Say this is 8 prompt tokens and 2 response tokens. Without a mask the loss is computed at every position, and about 80% of the signal goes into predicting the user's question. The model learns to write like the user instead of answering. With a mask, the prompt labels are replaced by ignore_index (−100), and the loss is computed only where 4 and <end> are predicted. With long prompts the skew is worse: a 900-token document and a 100-token response: without a mask the response would get 10% of the gradient.

  • Masking the loss on the prompt, and why without it the model learns the wrong thing

In plain terms. Asked "capital of Australia?", the teacher outputs a distribution: Canberra 0.7, Sydney 0.2, Melbourne 0.1. A one-hot label says only "Canberra". A soft label adds more: Sydney is a plausible mistake, Melbourne a less plausible one. A student trained on (0.7, 0.2, 0.1) gets more information from a single example. The loss here is cross-entropy with the teacher's distribution, and the student's log softmax is computed via logsumexp (week 4).

  • Synthetic data, distillation (the student learns from the teacher's outputs, including from logits; this is exactly where the instability of logsumexp shows up)
  • Data curation: deduplication (MinHash/LSH), quality filters, decontamination (removing training texts that overlap with test sets)

In plain terms. A layer matrix W of size 4096×4096 holds 16.8 million numbers, and it is left untouched. On top of it goes a correction ΔW = B·A: B of size 4096×8, A of size 8×4096, 65,536 numbers together, 0.4% of W. The correction is constrained in shape. Rank 1 is b·aᵀ: with b = (1, 2) and a = (3, 0, 1) you get [[3, 0, 1], [6, 0, 2]], and the second row must be a multiple of the first. Rank 8 means at most 8 independent directions of change for the whole layer. That is enough if the needed change is nearly low-rank, and not enough if it is not.

  • LoRA / QLoRA: low-rank decomposition (a weight correction as the product of two narrow matrices), where to insert adapters, what memory it saves. QLoRA is the same on top of a frozen model compressed to 4 bits (quantization, details in week 14)
LoRA: frozen W and the correction B·ALoRA: frozen W and the correction B·A
Diagram 21. W is frozen; a rank-r correction B·A with scale α/r is trained. At the start B = 0, so the output does not change; after training the correction is merged into W.

Code → nanolm/lora.py: LoRALinear, apply_lora, mark_only_lora_trainable, merge_lora, parameter_summary. The task is in exercises_en/lora.py, check: NANOLM_IMPL=exercises_en pytest tests/test_lora.py -v.

Check three things by hand: at the start the model's output does not change (B is initialized with zeros); after mark_only_lora_trainable the frozen parameters' .grad stays None; merging does not change the output.

And separately, test_lora_quality_is_bounded_by_rank: rank is a ceiling, not a "bigger is better" knob. If ΔW is full-rank, a low rank will never cover it.

Code → nanolm/compress.py: magnitude_prune (pruning, that is, zeroing the smallest weights: global and per layer), prune_ffn_neurons (removing whole SwiGLU neurons, so the matrices get narrower), low_rank_approximate via SVD (singular value decomposition), distillation_loss with a temperature and a T² factor. The task is in exercises_en/compress.py, check: NANOLM_IMPL=exercises_en pytest tests/test_compress.py -v. The distillation from the list above is in code here: with alpha = 0 the loss equals plain cross-entropy, and without the T² factor the gradient of the soft part would fall as 1/T². The low-rank error matches the Eckart–Young theorem.

Code → nanolm/data.py: a pretraining corpus pipeline built from the functions normalize_unicode, extract_text, detect_language, quality_filters, mask_pii, exact_dedup, minhash_signature, lsh_buckets, fuzzy_dedup, decontaminate, run_pipeline, with a per-stage report. The task is in exercises_en/data.py, check: NANOLM_IMPL=exercises_en pytest tests/test_data.py -v. python scripts/data_pipeline.py gives a pipeline report and trains the same TINY model on the raw and on the clean corpus. This is a separate week 15b: data lab: theory (NFKC, MinHash, choosing b and r, decontamination, PII and licenses), tasks and interview questions. If you are following the 24 weeks strictly, do its condensed 8-hour version in week 15 (described in the same place).

Math (Track D): D21: MinHash and LSH for deduplication; D22: bias, variance and shrinkage.

Interview question of the week: "How does LoRA work, and when is it not enough?" A 3-minute structure: (1) the formula first: W is frozen, a rank-r correction ΔW = B·A is trained; (2) a number: W at 4096×4096 is 16.8 million numbers, the adapter at r = 8 only 65,536, 0.4%; gradients and Adam state are needed only for the adapters, so of the 16 bytes per parameter (week 9) a frozen weight keeps 2, and with QLoRA half a byte; (3) initialization: B = 0, so the output does not change at the start; after training B·A is merged into W, and inference costs nothing extra; (4) the key takeaway: rank is a ceiling, not a "bigger is better" knob: if the needed change is full-rank, a low rank will never cover it; (5) expect "what if you zero both A and B?": the gradient of each is proportional to the other, both are zero, and training never starts.

Week outcomes

  • I can implement LoRALinear, apply_lora, merge_lora and pass all the module's tests.
  • I can compute the number of trainable parameters and the memory savings of LoRA for a given rank and set of layers.
  • I can explain exactly what the model will learn if the loss on the prompt is not masked.
  • I can describe a near-duplicate deduplication pipeline using MinHash and LSH.

Self-check

  1. Why is B initialized with zeros and A randomly, and what happens if you do it the other way round?
  2. Why is rank a ceiling on LoRA quality rather than a "bigger is better" knob?
  3. What is decontamination, and why does a benchmark prove nothing without it?

In the app each week has skills to rate yourself on, questions with answer checking, Python coding problems and a tutor grounded in the course.

Learn in the app: tutor, coding problems
← PreviousWeek 14. Scaling laws, precision, parallelism Next →Week 16. RL: from REINFORCE to PPO

Snula
Snula: LLMs from scratch

  • Home
  • Curriculum
  • App
  • Privacy
  • Terms

The course text is licensed under CC BY-NC-SA 4.0, nanolm code under Apache-2.0.