Week 23. Technical sprint
Learn in the app: tutor, coding problems →
Core: three mocks (two ML coding and the final rapid-fire), 10 timed track C problems · Depth: the mock online assessment, the math session, repeating weak formats · ≈ 14 h core / 20 h total
The goal of this week is not to learn anything new but to make what you have already learned available under pressure. Three formats train three different skills: rapid-fire trains fluency, ML coding trains execution, and experiment design trains reasoning. They cannot substitute for one another, so each one has its own protocol. One rule for all of them: no AI, timer on, everything recorded.
Of the course's 12 mock interviews, eight are already done in weeks 19–22, two a week; three are here and the last one is in week 24. The protocols of every format are collected below, and weeks 19–22 refer to them.
| Interview type | Already done | Load this week |
|---|---|---|
| ML coding (the most common) | 3 mocks (weeks 19, 20, 21) | 2 mocks, 45 minutes each: a transformer block; k-means or PCA |
| General coding | 10 timed problems from Track C and 1 mock 90-minute online assessment | |
| Rapid-fire | 2 sessions (weeks 19, 22) | 1 final session of 30 questions |
| Experiment design | 1 session (week 20) | one more in week 24, together with a paper review |
| Math | 1 session on paper from Track D | |
| ML system design | 1 session (week 21, session D) | repeat if the rubric score is below 6 |
| ML debugging | 1 session (week 22, session E) | repeat if the rubric score is below 6 |
Session A. Rapid-fire: 30 questions in 30 minutes
- Questions are drawn at random from the rapid-fire section of the question bank and from the "Self-check" blocks of weeks 1–22. The order is random: in an interview, topics do not come in phase order
- 60 seconds per answer, a hard limit of 2 minutes. Answer structure: a one-sentence definition → the mechanism or formula → the trade-off or the "why". Did not fit? Next question
- Record yourself on a voice recorder; log and score using the rapid-fire rubric and session template, marking each answer ✓ / ~ / ✗
- Split misses into three kinds: do not know, know but did not recall, said it imprecisely. The first go into the list of gaps for Sunday, the second into Anki, and the third need to be rephrased in writing
- Target by the third session: the checkpoint 3 threshold (≥ 20 ✓ and ≤ 3 ✗), with 24 ✓ as the benchmark. Why: rapid-fire tests not depth but the absence of gaps; one failure on "what is the KV cache" costs more than ten good answers
In plain terms. The question: "Why does attention divide by √H?"
Bad: "For stability, that is how it is done in the paper." There is no definition and no mechanism, so it gets ✗,
even though the word "stability" is not false.
Good (about 40 seconds): "So that softmax does not saturate. The dot product q·k is a sum of H terms,
and with unit-variance coordinates its variance is H. At H = 64 the logits are on the order of ±8,
softmax is almost one-hot, and the gradient through it is almost zero. Dividing by √H brings the variance back to 1.
Large models add QK-norm for the same reason." Definition, mechanism and "why" in under a minute: ✓.
Session B. Mock ML coding: 45 minutes
- Timing: 5 minutes to clarify the problem and the shapes out loud; 25 for the code; 10 to check it on a small example and on shapes; 5 for complexity and what to improve
- Problems: a transformer block; attention with a mask and GQA; beam search; generation with a KV cache; k-means or PCA (Module 0, F6)
- An empty file, no autocomplete, think out loud. In an interview, silence reads as "stuck", even if you are thinking productively
- Stuck for 5 minutes? Write down where, simplify the problem (no batch, no mask), get it working, and come back to the full version
- Common ways to lose points: softmax along the wrong axis, a forgotten mask, an unstable softmax,
mixing up
KandNin GQA, not a single shape check
Session C. Experiment design: 45 minutes
- Template: question → hypothesis and what would refute it → a trivial and a strong baseline → primary and secondary metrics → data and splits → noise control (seeds, intervals, McNemar) → ablations → budget → what could go wrong
- Examples: "Is it worth switching from MHA to GQA?", "Does a length curriculum help?", "Does reasoning suffer more from int4 than factual knowledge does?"
- Score it with the experiment design rubric. The most common mistake: starting with the method instead of with what result would refute the hypothesis
Session D. ML system design: 45 minutes
- Timing: 8 minutes on requirements and clarifying questions; 5 on the back-of-the-envelope estimate; 15 on architecture and data; 10 on quality evaluation and serving; 7 on failure modes, safety and the interviewer's questions
- A prompt from section 8 of the question bank, but not the one from part 5 of week 21 or from its mock. The partner answers clarifying questions from a "backstory" prepared in advance: how many users, what budget
- Scored with the ML system design rubric. The most common mistake: drawing a diagram after 2 minutes without finding out what matters most: latency, cost or accuracy. The second most common: not saying how you'll know the system works
Session E. ML debugging: 45 minutes
- The partner takes your working transformer block implementation and plants 5–6 bugs from section 9 of the question bank. 30 minutes to find and fix them, 15 for the extension: add a KV cache and check that it matches a full pass
- Witness checks first, then reading the code: shapes, causality, attention weight sums, overfitting one batch. Name each bug you find out loud together with the symptom it caused
- Preparation: the trainer has a problem on causal attention. Solve it, then break the mask, the softmax axis and the scaling in your solution one at a time, and see which test fails and with what message. This is how you learn to recognize a bug by its symptom rather than by its line of code. Then do the three "Find and fix" problems in the trainer: the attention block, the training loop and the top-k sampler. In each the code is already written and several bugs are hidden in it
- Scored with the ML debugging rubric
Weekly review. On Sunday, fill in a "format → misses → cause → action" table. A repeated miss means a topic, not an accident: it goes back into the core weeks.
Interview question of the week: "You are stuck in ML coding. What do you do?" This is answered with behavior, not words, so you practice it in every mock. Structure: say out loud exactly where the difficulty is → simplify the problem → check the simplified version on a small example → ask a clarifying question → return to the full version.
Week outcomes
- I can answer 30 random rapid-fire questions in 30 minutes with at least 20 ✓ and at most 3 ✗.
- I can write attention with a mask and GQA, or beam search, from an empty file in 25 minutes while thinking out loud.
- I can design an experiment from the template in 45 minutes, starting with the refutation criterion.
- I can sort my misses into the three kinds and name an action for each.
- I can run an ML system design session in 45 minutes from requirements to failure modes and justify choices with numbers.
- I can find at least 5 of 6 planted bugs in a transformer block and name the symptom of each.
Self-check
- In 60 seconds: what is the KV cache, how much memory does it take, and how is it reduced?
- What checks do you do in the last 10 minutes of a mock ML coding session, and in what order?
- For the question "does a length curriculum help?", what result would refute the hypothesis?
- Which witness checks do you write before reading someone else's broken code?