Week 18. Evaluation
Learn in the app: tutor, coding problems →
Core: contamination, McNemar, bootstrap, the p-value, selection versus reporting, experiment design · Depth: Pearson and Spearman, bandits versus A/B, red-teaming, test-time compute and CoT faithfulness · ≈ 11 h core / 22 h total
Any improvement from the previous weeks (a prompt, LoRA, DPO) is worth nothing until you show that the difference is larger than the noise. This week is about not fooling yourself: contamination, selecting on the test set, the wrong statistical test. Without it, the experiment design at checkpoint 4 and the research craft of week 22 rest on numbers you cannot trust.
- Benchmarks and their pathologies: contamination (the test leaked into the training data), sensitivity to prompt format, saturation (every model is at the ceiling, and the benchmark no longer tells them apart)
- Evaluating generation: perplexity (defined in week 4), pass@k (the share of tasks where at least
one of
kattempts is correct), LLM-as-judge (a model grades answers instead of a human: position bias, it prefers whichever comes first; verbosity bias, it prefers the longer one; self-preference, it prefers its own) - Human eval, annotation, inter-annotator agreement (how often annotators give the same labels)
pass@1 or pass@k: what do you want from the model? In plain terms. Policy A solves every task with probability 0.6
on each attempt. Policy B knows 60% of the tasks for certain and never solves the other 40%. Both have pass@1 of 0.6.
pass@3: A gets 1 − 0.4³ ≈ 0.94, B still gets 0.6. Same mean accuracy, yet a 34 pp gap on pass@3.
The language for this comes from the theory of choice under risk. By the von Neumann–Morgenstern theorem (1944), preferences over
lotteries (outcomes with randomness), if they are complete, transitive, continuous and independent of irrelevant outcomes,
can be written as maximizing expected utility E[u(x)]. An evaluation metric is exactly such a u of the per-task success rate p.
pass@1 takes u(p) = p: linear, risk-neutral, only the mean matters. pass@k takes u(p) = 1 − (1 − p)^k:
concave, so by Jensen's inequality it dislikes spread of p across tasks. For B, the per-task success rate is
a "1 or 0" lottery, and pass@3 values it the same as a guaranteed p = 1 − 0.4^(1/3) ≈ 0.26 on every task
(the certainty equivalent); the risk premium is 0.6 − 0.26 ≈ 0.34. Takeaway: RL with a "correct or not" reward maximizes
E[r], which is pass@1, and it does not care whether A turns into B. If the model will be sampled k times, measure pass@k at several values of k.
The Pareto front. Besides quality there is cost. One model dominates another if it is no worse on both axes and better on at least one; the non-dominated models form the Pareto front. Example (price per 1M tokens, accuracy): M1 $0.5 and 68%, M2 $2 and 79%, M3 $3 and 76%, M4 $8 and 84%. M3 costs more than M2 and is worse: it drops out. Between M1, M2 and M4, the budget decides, not the table.
In plain terms. Two models, 200 questions. Both are right on 150 and both are wrong on 30. Those 180 questions say nothing about the difference. That leaves 20 disagreements: in 15, A is right and B is not; in 5, the reverse. If the models are equal, disagreements split like a coin flip, about 10 and 10. A skew of 15:5 or more in either direction happens with probability about 0.04. McNemar looks only at these 20 pairs. Two separate 95% intervals, one for A (82.5%) and one for B (77.5%), would overlap and hide the effect.
In plain terms. Out of 1000 ideas you test, 100 actually work. A test with a 0.05 threshold and power 0.8
flags 80 of the working ones as significant, plus 45 of the 900 that do not work (5%). Of the 125 "significant" results, 45 are false, which is 36%.
So p < 0.05 does not mean "there is less than a 5% chance there is no effect".
The p-value is P(data | H₀), and the share of false discoveries also depends on how often ideas work at all.
- Statistics used for the right purpose:
- McNemar: two models on the same test set (the best choice for this standard setup)
- Paired t-test (a t interval over a few measurements already appeared in week 13,
bench.py), bootstrap confidence intervals (resample the test set with replacement and look at the spread of the metric), multiple comparisons (the more checks you run, the more random "significant" results) - Pearson vs Spearman: robustness to outliers, nonlinear monotonic relationships
- The p-value is
P(data | H₀), notP(H₀ | data)
- Safety evaluation, red-teaming (deliberately searching for inputs that break the model), jailbreaks
In plain terms. Two prompt templates, P1 and P2, and a test set of 100 questions. P1 gets 71%, P2 gets 74%. You pick P2 and write 74% in the report. But on 100 questions the random spread of accuracy is about ±4.5 pp (one standard deviation). Suppose both templates truly give 72% and their noise is independent. Then the better of the two will on average show about 2.5 pp more than the truth: you picked it precisely because it got lucky. Out of 20 templates, the "winner" on average brings about +8 pp out of nothing. The honest way: choose the template on one set and measure it on another that you did not touch during selection.
Where to select and where to report. This is the discipline without which all the statistics
above are useless: (a) the checkpoint, prompt, threshold and hyperparameters are chosen on one
dataset, and the result is reported on another; (b) a data slice where the method
"helps especially" is found on one half and confirmed on the other.
Otherwise you find noise and call it an effect. It is the same principle as honest
splitting for estimating heterogeneous effects in causal inference.
Test yourself: "I found a slice where the method gives +8 pp. Why is that not enough, and what should I do?"
In the trainer this is the select_then_report problem: it measures by simulation how inflated the winner is on the sample
where it was selected, and shows that the inflation disappears on a separate test set.
A/B test versus a multi-armed bandit. An A/B test (traffic is split equally and in advance, the decision comes after the test) answers
the question "how much better is B than A". A multi-armed bandit (an algorithm that, as the test runs, gives more traffic to the variant
that is winning right now; for example, Thompson sampling shows a variant with the probability that it is the best given the current data)
answers a different question: "how do we lose as little as possible while we find out".
In plain terms. Two prompts in a product: A has 5% successful dialogs, B has 6%, and the test has 20,000 dialogs. An A/B test gives the worse
variant half of them, 10,000 dialogs, and loses about 10,000 · 0.01 = 100 successes. After a few thousand dialogs a bandit
gives most of the traffic to B and loses several times less. The price is paid in the estimate. A variant that was unlucky at the start
(at 5%, getting 0 successes in the first 20 impressions has probability 0.95²⁰ ≈ 0.36) almost stops getting traffic,
and its underestimate is never corrected. A lucky variant gets more data, and its estimate returns to the truth.
So the per-variant means in a bandit are biased downward, more so for the losers, the difference of means usually overstates the effect,
and the sample size itself depends on the outcomes, so an ordinary t-test or confidence interval does not apply.
The rule: if you need an honest effect estimate for a report or a long-term decision, run an A/B test with a fixed split.
If there are many variants, they live briefly (a promotion, a prompt for one week), and every impression of a loser costs money, use a bandit.
An effect estimate is recovered from a bandit by reweighting with the logged impression probabilities (inverse propensity weighting),
or by keeping a fixed share of traffic on uniform allocation.
Test-time compute: how to spend compute on an answer. A model can be improved not only by training: you can give it more compute per answer. Reasoning models that "think" for thousands of tokens and "generate many, pick one" schemes are built on this. They must be evaluated as strictly as everything else in this week.
In plain terms. A model solves a task with probability p = 0.3 per attempt, and attempts are independent. There is a perfect
verifier (a program that checks the answer, as in the RLVR of week 17). Then among N attempts at least one correct
answer turns up with probability 1 − (1 − p)^N: for N = 1 that is 0.3, for N = 4 already 1 − 0.7⁴ ≈ 0.76, for N = 16
about 0.997. It is the same formula as pass@k above, except that now the verifier picks the answer, not a human.
With an imperfect verifier (a reward model) the growth stops earlier: the larger N, the more likely the attempts
contain a wrong answer that the verifier overrates. This is the reward hacking of week 17, only at inference time.
- Best-of-N (generate
Nanswers and return the one the verifier or reward model scored highest). Quality is limited by the quality of the verifier: with a perfect one it grows as1 − (1 − p)^N, with a noisy one it plateaus or even drops asNgrows - Self-consistency (
Nindependent reasoning chains, their final answers are compared, and the most frequent one wins; no verifier needed). In plain terms. Five chains gave the answers 12, 12, 15, 12, 9: 12 wins with three votes. It helps when the task has a short answer that can be compared (a number, an option) and the errors are scattered over different wrong answers: the model gives the right answer more often than any particular wrong one. It does not help when the model is systematically wrong in the same way (the most frequent answer is wrong, and voting locks in the error), or on open-ended answers (essays, code), where two answers rarely match and there is nothing to vote on - PRM versus ORM. An ORM (outcome reward model) scores only the final answer. A PRM (process reward model) scores every reasoning step. In plain terms. A solution of five steps has an error at step 2, but the answer happens to come out right. The ORM gives a high score, the PRM finds the error at step 2 and gives a low one (the solution's score is often taken as the minimum or the product of the step scores). A PRM gives a denser signal and selects answers better in best-of-N, but it needs step-level labels, which are expensive; an ORM is cheap but rewards "the right answer by the wrong route"
- The "quality versus inference compute" curve. The x axis is FLOPs (or tokens) per task, the y axis is accuracy.
Methods are compared at an equal budget; otherwise any method wins simply because it was given more attempts.
In plain terms. A 1B model spends about
2 · 10⁹ · 500 = 10¹²FLOPs on a 500-token answer (2Nper token at inference, week 9); best-of-16 spends1.6 · 10¹³. A single answer of the same length from a 16B model costs the same. If best-of-16 gives 62% and the 16B model gives 70%, this budget is better invested in model size. Where spending at inference pays off depends on task difficulty: on easy and medium tasks search and verification give more, on the hardest tasks a larger model gives more - Reasoning distillation. A large model generates long reasoning chains, a checker keeps the correct ones, and a small model is SFT-ed on them (week 15). This way the small model learns to reason more cheaply than through RL from scratch (recall the threshold of about 1.5B parameters from week 17). The same trick shortens reasoning: out of several correct chains you keep the shortest, and the model learns to answer more briefly without losing accuracy
- Chain-of-thought faithfulness (CoT faithfulness). The text of the reasoning does not have to reflect the real cause of the answer. In plain terms. In the prompt's examples the right answer is always under the letter "A". On new questions the model picks "A" 30 pp more often than without this skew, and in its reasoning it explains the choice by the content of the task and never mentions the pattern. This is checked by intervention: (1) add or remove the hint and see whether the answer changes and whether the reasoning admits it; (2) cut the reasoning off halfway or insert an error into it: if the answer does not change, the reasoning was decoration, not the cause. Conclusion: a reasoning chain can be read as an explanation only after such a check, and this matters for safety monitoring
In the trainer this is the best_of_n problem: the probability of finding a correct answer among N, majority voting
with tie-breaking, and selection by verifier scores.
Math (Track D): D27: McNemar and the paired bootstrap; D28: multiple comparisons.
Interview question of the week: "Model B is 5 pp better than A on one test set of 200 questions.
Is that significant?" A 3-minute structure: (1) first, clarify the setup: it is the same set, so
the comparison is paired; (2) the conclusion: McNemar on the disagreeing pairs, not two independent intervals; the number: of
20 disagreements, 15:5 in favor of B, p ≈ 0.04, even though intervals for 82.5% and 77.5% would overlap; (3) the scale
of the noise: on 100 questions, one standard deviation of accuracy is about 4.5 pp; (4) where the selection happened: if the prompt
or checkpoint was chosen on this same set, the gain is inflated (out of 20 templates, the "winner" brings about +8 pp
out of nothing), so you need a separate set; (5) contamination and sensitivity to prompt format. Expect the follow-up "what if
you compare 20 slices?" That is multiple comparisons: some "significant" results are random, and a slice is confirmed
on the other half.
Sources: Athey, Imbens, Machine Learning Methods That Economists Should Know About (2019), §6.2–6.3 and the section on multi-armed bandits; Athey, Imbens, PNAS (2016), honest trees. On test-time compute: Cobbe et al., Training Verifiers to Solve Math Word Problems (2021); Wang et al., *Self-Consistency Improves Chain of Thought Reasoning in Language Models* (2022); Lightman et al., Let's Verify Step by Step (2023); Snell et al., Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (2024); Turpin et al., Language Models Don't Always Say What They Think (2023); Lanham et al., Measuring Faithfulness in Chain-of-Thought Reasoning (2023). Links are in 04-РЕСУРСЫ, section "Test-time compute and evaluating reasoning".
Week outcomes
- I can choose a test for comparing two models on the same test set and run McNemar by hand.
- I can build a bootstrap interval for the difference in accuracy between two models.
- I can name three LLM-as-judge biases and a way to control each.
- I can define the p-value correctly in one sentence.
- I can separate selection from reporting: select on validation, report on a separate test set, and estimate how inflated the winner of N variants is.
- I can say when to use a multi-armed bandit instead of an A/B test and why its effect estimate cannot be read like an A/B result.
- I can compute the gain of best-of-N with
1 − (1 − p)^N, compare it with self-consistency and with a larger model at an equal FLOPs budget, and propose an intervention test of whether a reasoning chain is faithful.
Self-check
- Why is McNemar better than two independent intervals for two models on the same test set?
- How do you check a benchmark for contamination?
- When is Spearman preferable to Pearson?
- Two models have the same pass@1. Why can pass@k differ by tens of pp, and which of the metrics is risk-neutral?
- A prompt was chosen out of 20 on a test set of 100 questions, and its accuracy on that same set went into the report. Roughly how inflated is the number, and how do you set up an honest procedure?
- When is a multi-armed bandit better than an A/B test, and why can the difference of means from a bandit not be read as an effect estimate?
- A model solves a task with probability 0.2 per attempt. How many attempts does best-of-N with a perfect verifier need to find a correct answer with probability at least 0.9? Why does the growth stop earlier with a reward model instead of a verifier, and when does self-consistency not help at all?
- How does a PRM differ from an ORM, and why is the "accuracy versus inference compute" curve compared at an equal FLOPs budget? How do you check by intervention that a reasoning chain actually led to the answer?
✅ Checkpoint 4
Design an experiment: "We want to know whether method X helps. How would you test it?" You need the design, baselines, metrics, statistical test, ablations, and what could go wrong. 45 minutes, out loud.