Study Program: From Zero to Research Scientist / MTS
24 weeks. Built on open primary sources: Stanford CS336, original papers and model technical reports. The topic list and priorities were cross-checked against public lecture notes and an analysis of 57 interviews by an OpenAI researcher (see Resources), plus skills that are missing there but come up in interviews and matter on the job.
Workload: the whole program takes 20–25 hours a week, about 520 hours over 24 weeks. The core of each week takes 10–14 hours: if you only have 10–15 hours a week, do the core and leave the depth for a second pass (the route is in the Program map). Keep the order of the weeks.
Before and after the core. Week 1 is preceded by the entry diagnostic and Module 0 (2–6 weeks, only if the diagnostic showed gaps). After week 22 comes an optional research Capstone of 4–6 weeks.
Every week ends with two blocks. Week outcomes: what you should be able to do. Self-check: 3 to 11 questions (depending on how many topics the week has) that you must answer out loud without peeking, about a minute per question. If you could not answer, the week is not done.
How a week works
Every week = 1 topic + 5 parallel tracks that run through the entire program.
| Track | What | Per week |
|---|---|---|
| A. Theory | The week's topic: notes in your own words + deriving formulas on paper | 6–8 h |
| B. ML code from scratch | 1 implementation per week, no AI assistant | 4–6 h |
| C. LeetCode | 3 problems a week from the Track C list | 3 h |
| D. Math | 2 problems a week from Track D | 2 h |
| E. Review | Anki + "a transformer from memory in 25 minutes" | 3 h |
The iron rule: tracks B and E are written with AI completely turned off. Otherwise you will overestimate yourself by exactly as much as you lean on autocomplete.
Core and depth. Under each week's heading there is a line "Core: … · Depth: …" with a time estimate. The core is the topics of the ★ skills from the skill map, the week's main code and track E: without them you will not pass RS/MTS interviews. The depth can wait: the "optional" blocks, the sections of 05-ГЛУБИНА, track D and the second half of the track C problems. "≈ 11 h core / 21 h total" means that the core, together with track E and one or two track C problems, takes about 11 hours, and the whole week with all five tracks about 21.
Weekly ritual:
- Mon–Thu: theory + code
- Fri: track E, write a transformer from an empty file against the clock. Record the time.
- Sat: LeetCode + math
- Sun: retrospective. Anything you did not understand → goes on the gap list. 1 hour to close one gap.
Module 0
- Module 0. Foundations before you startFor those who need the foundations before week 1: the entry diagnostic and six weeks, F1 to F6.
24 core weeks
Phase 1. Foundations
- Week 1. Neural networks and gradients
- Week 2. Backpropagation
- Week 3. Optimizers and training regime
- Week 4. Information theory and numerical stability
Phase 2. The modern transformer
- Week 5. Tokenization
- Week 6. Architecture, part I
- Week 7. Attention
- Week 8. Assembling the full transformer
- Week 9. Bookkeeping: parameters, FLOPs, memory
- Week 10. Training
Phase 3. Inference, GPUs, scaling
- Week 11. Inference
- Week 12. Sampling strategies
- Week 13. GPUs and FlashAttention
- Week 14. Scaling laws, precision, parallelism
Phase 4. Post-training
Phase 5. Expansion
- Week 19. Other architectures: RNN, SSM, MoE
- Week 20. Multimodality and long context
- Week 21. Production LLMs: RAG, agents, evaluation
- Week 22. Infrastructure and research craft
Phase 6. Interview mode
Progress metrics across the program
Track these every week:
- [ ] Time to write a transformer from scratch (target: ≤25 min)
- [ ] LeetCode problems solved (target: 72, that is, all of Track C)
- [ ] Implementations from scratch (target: 24)
- [ ] Papers read (target: 100+)
- [ ] Anki cards in active review
- [ ] Mock interviews done
What this program leaves out on purpose
- Classical ML (trees, SVMs) beyond the minimum: it rarely comes up in interviews for LLM roles. The minimum (logistic regression, k-means, PCA, a decision tree, metrics) is in Module 0, week F6
- Deep optimization theory: the level of "I understand what is going on" is enough
- Computer vision beyond what multimodality requires
Why do all this, besides the offer
The main argument for this kind of preparation is not that it helps you pass interviews. Systematically closing the gaps in your foundations has an effect that is worth more than the offer itself: your productivity in your current work goes up. Technical ideas appear that were simply out of reach before, and the range of problems you can think about at all gets wider. Where this conclusion comes from is explained in the "Program map".
Twenty-four weeks seems like too long to endure for the sake of a single interview. But if you count the fact that what changes at the end is the level at which you are able to work, the time frame looks different.