Evaluation
Transcript: cleaned auto-captions with timestamps
CS336 spends eleven lectures making a number go down — loss, latency, dollars per token — and this is the lecture that asks whether the number was ever the right one. It sits at the end of the scaling arc for a reason: scaling laws (L9–L11) let you predict a loss at a scale you cannot afford to run, but a predicted loss is only useful if loss is a proxy for something you care about. Percy's framing is that evaluation is not the mechanical afterthought it looks like — write a script, throw prompts, average — but the mechanism by which the field's objective is chosen. It is also the natural bridge into the data lectures that follow: once you accept that the score is contaminated by what is in the training set, you have to go look at the training set.
Outline, with timestamps
- 00:05 — The numbers you already see: the benchmark tables in model releases, HELM, Artificial Analysis price-quality frontiers, OpenRouter traffic, Chatbot Arena, and X vibes — six incompatible definitions of “good”.
- 03:29 — Karpathy's evaluation crisis: the standard benchmarks are saturated, gamed, or somewhere in between, and nothing has replaced them.
- 05:11 — There is no one true evaluation: four buyers — purchaser, researcher, policymaker, model developer — with four incompatible questions, and evaluation as the field's real objective function.
- 07:57 — The four-question framework: what are the inputs, how do you call the model, how do you score the outputs, how do you interpret the result.
- 14:37 — Perplexity, recalled: what (1/p(D))^(1/|D|) measures, and the in-domain era of Penn Treebank, WikiText and the One Billion Word Benchmark.
- 19:02 — GPT-2 and zero-shot perplexity: train on WebText, test on everyone else's splits: transfer beats in-domain training on small corpora and loses on 1BW.
- 25:17 — Trusting the provider, and perplexity maximalism: why a perplexity leaderboard has to believe your probabilities normalise, and the argument that minimising perplexity is the whole road.
- 29:48 — Cloze cousins: LAMBADA and HellaSwag: likelihood-scored tasks that are perplexity in costume — and HellaSwag's WikiHow contamination wrinkle.
- 32:01 — MMLU, and looking at the instances: 57 subjects, web-collected, knowledge rather than language understanding; clicking into HELM to read the actual five-shot prompts.
- 40:25 — The difficulty ladder: MMLU-Pro, GPQA, HLE: rebuilding the exam three times as it saturates, and the selection bias in soliciting hard questions by open call.
- 52:19 — Instruction following and LLM judges: Chatbot Arena and Goodhart, IFEval's verifiable constraints, AlpacaEval's length bias, WildBench's checklist judge.
- 59:58 — Agents, and pure reasoning with ARC-AGI: SWE-bench, Cybench and MLE-bench force the model-versus-system question; ARC-AGI factors knowledge out and puts cost back in.
- 66:10 — Safety: capability, propensity, dual use: HarmBench, AIR-Bench, transferable jailbreaks, voluntary pre-deployment testing — and why refusal alone is not a score.
- 73:57 — Realism, validity, and the rules of the game: quizzing versus asking, Clio and MedHELM, contamination, label noise, and methods versus models.
05:11 · There is no one true evaluation
The organising claim is that "how good is this model" is not one question. It is four questions wearing the same clothes, and they want four different instruments. A company choosing a vendor for a support bot wants a score on its traffic at its price point. A researcher wants raw capability, unanchored to any use case. A policymaker wants an inventory of benefits and harms. A model developer wants a signal that moves when they change something, so they know to keep the change. Nothing about a benchmark number tells you which of those it answers, and a number built for one is usually invalid for the others.
The second half of the claim is the part practitioners underrate. Evaluations are not passive instruments; they are what the field is hill-climbing. Frontier labs track a suite over time, and how a model gets built bends toward the tracked quantity.
"...if you track something and you're trying to get your number to go up it's going to really influence the way that you develop your model. So that's why evaluation I think is really … a leading indicator of where things are going to go."— Percy Liang, 05:11
Which is why the opening survey matters. Model cards report MMLU, MATH, GPQA, AIME, Codeforces; HELM collates; Artificial Analysis plots a quality index against price per token; OpenRouter ranks by tokens people voluntarily route; the Arena ranks by human preference; X ranks by vibes. Each is a different definition of "good", and a release table picking three of them is making an argument, not reporting a fact.
07:57 · Four questions that decide what a number means
Percy's checklist is the most portable thing in the lecture — four stages, each with failure modes that survive into the final number.
Inputs. Which use cases are covered, is the difficult tail represented, and are the inputs adapted to the model? Multi-turn chat forces adaptation (a static transcript puts the model in a conversational position it would never have walked itself into) and red-teaming demands it, because generic prompts are hopelessly inefficient at finding rare tail failures. But adaptation destroys cross-model comparability. Percy names the trade-off and does not resolve it (13:29).
The call. Zero-shot, few-shot, chain-of-thought, tools, RAG — each is a different measurement of the same weights, and models stay prompt-sensitive enough that the choice is a large share of the variance. It is why MMLU's headline moves from the mid-80s into the 90s on prompting and ensembling alone. It is also where the deepest ambiguity lives: are you evaluating the language model, or the agentic system wrapped around it? The developer wants the former; the user is buying the latter and does not care how many models are inside.
The score. Are the reference answers even correct? pass@1 or pass@10? How is cost folded in, when most leaderboards marginalise it away and the top model may be 10× the runner-up's price? How are asymmetric errors handled, where a hallucination in a medical setting is not one unit of loss like any other? And what do you do with open-ended generation, which has no ground truth at all?
The interpretation. Is 91% deployable? Does a high score reflect generalisation or overlap with the training set? Are you scoring a method, a model, or a system? The rest of the lecture is these four questions applied to one benchmark family at a time.
14:37 · Perplexity: the metric that refuses to die
Perplexity is (1/p(D))^(1/|D|) — the geometric-mean inverse probability assigned to a held-out corpus, the exponentiated cross-entropy you already minimise in pre-training. Through the 2010s that was evaluation: pick Penn Treebank, WikiText or the One Billion Word Benchmark, train on the designated split, test on the designated split. It drove real progress — Jozefowicz et al. took 1BW from 51.3 to 30.0, an enormous move.
GPT-2 broke the frame by refusing to train in-domain at all: train on WebText, evaluate zero-shot on everyone else's test sets. The shape of the result is worth internalising. On small corpora like Penn Treebank, broad web training beat the in-domain state of the art without ever seeing the training split — transfer wins when the target set is too small to learn its own distribution. On 1BW, in-domain training still won comfortably, because at a billion words you can just fit the distribution directly.
Downstream accuracy has since taken over, but perplexity keeps a permanent seat for three reasons. It is smooth — continuous in every token's log-probability, so it fits clean scaling curves where accuracy gives step functions. It is universal — it attends to every token, where task accuracy can be right for the wrong reasons on a gameable set. And with genuinely disjoint splits it is hard to game: no prompt trick makes a model assign higher probability to text it does not model. You can also condition it — perplexity of the gold answer given the prompt — and fit scaling laws directly on the task you care about.
The catch is operational. Task accuracy is black-box safe: take the string, run the scorer. Perplexity requires the provider to hand you probabilities and requires you to believe they normalise. Expose only "probability of this next token" and a buggy — not even malicious — implementation returning 0.8 for everything looks spectacular and is not a distribution. Full logits are verifiable; a scalar is not.
The maximalist position follows: if the true distribution is t and yours is p, the best achievable perplexity is H(t), attained exactly at p = t; a model matching t solves every task; so pushing perplexity down is the road to everything. Percy's caveat is efficiency, not correctness — most probability mass sits on tokens nobody cares about. Two benchmarks are perplexity in costume: LAMBADA (predict the final word of a passage engineered to need the whole discourse) and HellaSwag (pick the likeliest completion), both largely saturated. HellaSwag carries a warning too: it was mined from ActivityNet and WikiHow, and WikiHow is on the open web, so near-duplicates of the eval distribution sit in everyone's pre-training corpus without any verbatim match.
32:01 · The knowledge ladder, and why it keeps getting rebuilt
MMLU (2020) is the canonical exam: 57 subjects, multiple choice, "collected by graduate and undergraduate students from freely available sources online" — both why it exists and why contamination is unavoidable. Percy's quibble is that despite the name it measures knowledge, not language understanding: he is confident in his language understanding and would still fail the foreign-policy questions. It landed right after GPT-3, when one model few-shotting 57 subjects was genuinely startling, and GPT-3 scored about 45%.
More valuable than the scores is what the lecture does next: it opens HELM and clicks into individual instances — the exact five-shot prompt, the candidate letters, the prediction, right or wrong. That is where you discover that few-shot choice, ordering and format all move the number materially (seed a classification task with only positive examples and the model will happily emit only positives).
The subtlest point is what MMLU is for. It was built to probe base models, where a good score means the capability fell out of generic next-token prediction on a broad corpus.
"...if you were able to magically train on a lot of data and be able to do well on MMLU without even trying — this is like … not studying for the exam and doing well on the exam — then you probably have a good amount of quote-unquote intelligence."— Percy Liang, 39:51
Curate multiple-choice questions across those 57 subjects, post-train on them, and you get a great MMLU score that estimates nothing. Same number, opposite meaning, and the difference lives entirely in the training set — which the interpretation stage has no access to.
The rest of the ladder is the field rebuilding the exam as it saturates. MMLU-Pro strips noisy and trivial items, widens 4 choices to 10 (the stated motivation: everyone was scoring ~90 and you cannot give everyone an A) and evaluates with chain-of-thought; accuracies drop 16 to 33 points, which is the headroom developers have since adopted. GPQA engineers expert difficulty: questions written by 61 PhD-level contractors hired on Upwork, expert-validated, revised, re-validated, then attempted by non-experts with 30 minutes and Google. Experts land around 65%, non-experts with search around 34%, GPT-4 got 39% at publication — and by the lecture o3 was near 75%, which is the ladder's story in eighteen months. Humanity's Last Exam pushes further: 2,500 multimodal questions, a $500K prize pool and co-authorship for writers, frontier models used as a filter to reject anything too easy; o3 sat around 20% on a benchmark named for finality. A student objects that an open call for hard questions selects for people already steeped in LLMs, and Percy concedes it (51:46): these sets are hard, and that is the only claim they support.
52:19 · Open-ended: arenas, verifiers, and judges
Everything above is multiple choice or short answer, a shrinking slice of how models are used. Instruction following has no task inventory — the user describes a one-off job — so scoring an open-ended response is the unsolved core problem. Three families exist, each broken differently.
Human pairwise preference. Chatbot Arena serves a real prompt to two anonymised models, records which response the person preferred, and fits Elo. Its virtues are structural: live inputs, so the set cannot go stale, and Elo absorbs new entrants without re-running anything. Its vulnerability is equally structural — once CEOs are tweeting their Arena position, Goodhart applies, and The Leaderboard Illusion documents the protocol asymmetries (privileged access, multiple private submissions) that convert compute into rank (55:02). Plus the unglamorous question of whose preferences these are: whoever happens to be on the site.
Programmatic verifiers. IFEval sidesteps judgment with synthetic, checkable constraints — at most three sentences, no commas, include these words. A script verifies compliance exactly. What it cannot verify is whether the response is any good: write a story about a dog in ten words and IFEval confirms ten words, not that it is a story. Partial by construction, gameable, and the constraints are frankly artificial.
LLM judges. AlpacaEval scores a win rate over 805 instructions against a reference model, judged by GPT-4 — obviously biased, and it broke instructively: small models climbed by emitting longer answers the judge liked, until a length-controlled variant corrected for it. WildBench draws 1,024 examples from a million real human-chatbot conversations and judges with a checklist (chain-of-thought for judging), correlating about 0.95 with the Arena. Note the recursion: the field validates a new benchmark by correlating it with the Arena — the benchmark under active dispute. Sanity check, not ground truth.
59:58 · Agents, reasoning, and safety — the object of evaluation blurs
An agent is a language model plus scaffolding: the program deciding when to call the model, what to feed it, what to do with the output. Evaluate one and the model-versus-system ambiguity stops being philosophical, because the scaffold contributes a large, unreported share of the score. SWE-bench gives 2,294 real issues from 12 Python repositories and grades patches by whether the repo's unit tests pass — an unusually honest metric, since the reference is executable rather than textual. Cybench runs 40 Capture-the-Flag tasks where the agent must actually get a shell and retrieve a key, and adds a good difficulty axis: human first-solve time, up to 24 hours. At lecture time models were clearing tasks that took human teams around 42 minutes, at roughly 20% overall — far more legible than a bare percentage. MLE-bench is 75 Kaggle competitions run end to end, with medal rates still under 20%.
ARC-AGI is the deliberate outlier: grid puzzles introduced by François Chollet in 2019, before the LLM era, with linguistic and world knowledge factored out so only pattern induction remains. GPT-4o scored near zero; o3 scored well — at hundreds of dollars of inference per task. That price is not a footnote, it is the result. A benchmark reporting accuracy without compute lets you buy any score you like, which is exactly why the price-quality frontier opened the lecture.
Safety is where the framework pays off hardest, because the naive instrument measures the wrong thing. HarmBench scores refusal across 510 harmful behaviours. AIR-Bench 2024 grounds its taxonomy in actual regulation and company policy — 314 risk categories, 5,694 prompts — so "safe" becomes checkable rather than vibes. Jailbreaking sits above both: GCG optimises an adversarial suffix against open-weight Llama models and the attack transfers to closed models like GPT-4, so a refusal rate measured on benign phrasing is an upper bound, not a property. Pre-deployment testing by the US and UK AI Safety Institutes gives evaluators pre-release access and produces a report; it is entirely voluntary, with no statutory backing.
Two distinctions are the ones to carry. A refusal-only leaderboard is trivially topped by a model that refuses everything, so a safety score means nothing unpaired with a capability score. And separate capability (can the model do the harmful thing at all) from propensity (will it, as shipped). Behind an API only propensity matters — a model that knows how and reliably declines is fine, modulo jailbreaks. For open weights, capability is what matters, because fine-tuning the propensity away is cheap and well documented. Dual use then entangles them: the safety institutes run Cybench as a safety evaluation, but the same capability is what makes an agent useful for penetration testing. Same measurement; the sign depends on who holds the weights.
73:57 · Realism, validity, and the rules of the game
Standardised exams are far from how models are used, but raw production traffic is not the answer either — much of it is people messing with the system. Percy's distinction is quizzing versus asking: in a quiz the user knows the answer and is testing you; in an ask they do not and want it. Asking is where the value is, and where almost no public benchmark lives. Two responses appear. Clio uses models to privacy-preservingly cluster real conversations and publish what people actually do (coding at the top). MedHELM builds 121 clinical tasks elicited from 29 practising clinicians rather than from licensing exams — note writing, treatment planning, the work rather than the credential. The structural tension: realism and privacy pull against each other, so parts of MedHELM cannot be hosted publicly, and a benchmark you cannot inspect is one you must trust.
Validity is where the lecture gets bleakest and most actionable. The old world had designed splits; the new one trains on the internet and does not disclose the corpus, so contamination is unfalsifiable from outside. Standard practice is 13-gram decontamination, which leaks in three directions Percy names: paraphrases and near-duplicates evade n-gram matching; a problem translated into another language has zero surface overlap and is trivially recoverable by a capable model; and training documents that quote the test set generate false positives that make honest filtering over-aggressive. Two partial routes. Infer it: exchangeability tests exploit the fact that a model trained on a dataset in its canonical order assigns higher likelihood to that order than to a shuffle — a black-box membership test. Normalise it: make stating measured train-test overlap as expected as reporting a confidence interval. Then reference quality, which is not a rounding error — SWE-bench needed a Verified subset after real errors surfaced, and on sets like MATH and GSM8K a substantial share of the residual above 90% is label noise rather than difficulty. Some of the hardest remaining questions are simply wrong.
The closing move reframes everything. Pre-foundation-model ML evaluated methods: fixed data, fixed splits, change the algorithm, and the number attributes the improvement. Today we evaluate models and systems, where anything goes — any data, any scaffold, any inference budget — and the number attributes nothing in particular. Both are legitimate; the confusion is in leaving it unstated. The exceptions prove it can be done: the nanoGPT speedrun fixes data and hardware and races to a target validation loss; DataComp-LM fixes the training pipeline so only data selection varies. Method evaluation drives algorithmic innovation, system evaluation serves users. Either way, state the rules of the game.
What you build with this
No new assignment starts here. Assignment 4 (Data) is in flight — it began at L11 — and this lecture is, unusually, a direct commentary on how that assignment is scored. A4 hands you a frozen staff training implementation you are forbidden to modify and ranks submissions by validation loss on a C4 100-domains split: repo, handout, leaderboard. That design is exactly the "evaluating methods" regime from the final section — the DataComp-LM shape, with the pipeline pinned so the only free variable is your data — and the metric is perplexity precisely for the reasons argued at 25:17: it is smooth enough to rank close submissions and hard to game when the eval split is genuinely disjoint. Read the leaderboard rules with this lecture in hand and you will see why the staff reserve the right to re-run the top five themselves. Looking forward, Assignment 5 (Alignment, from L16) puts you on the other side: you will post-train a model and need to decide what evidence counts as improvement, which is where the capability-versus-propensity split and the LLM-judge caveats become load-bearing rather than academic.
Supporting materials, verified
- Measuring Massive Multitask Language Understanding (MMLU) — Hendrycks et al. (2020) · The canonical 57-subject exam and the reference point for the entire saturation story.
- MMLU-Pro — Wang et al. (2024) · De-noised MMLU with 10 choices and CoT evaluation; the 16–33 point drop that restored headroom.
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark — Rein et al. (2023) · The expert/non-expert calibration (65% vs 34%) is the part worth copying into your own eval design.
- Humanity's Last Exam — Phan et al. (2025) · Prize-pool crowdsourcing plus frontier-model filtering; also the clearest example of the selection bias a student raises in the lecture.
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference — Chiang et al. (2024) · The live pairwise-Elo design Percy walks through.
- The Leaderboard Illusion — Singh et al. (2025) · The Arena critique referenced at 55:02: private multi-submission and access asymmetries.
- Instruction-Following Eval (IFEval) — Zhou et al. (2023) · Verifiable synthetic constraints; the model for "score what a script can check".
- AlpacaEval — Li, Dubois et al., Stanford (2023) · 805 instructions, LLM-judged win rate; the leaderboard whose length bias became a case study.
- WildBench — Lin et al. (2024) · 1,024 real conversations, checklist judging, ~0.95 correlation with the Arena.
- SWE-bench — Jimenez et al. (2023) · 2,294 GitHub issues scored by unit tests; the executable-reference agent benchmark.
- Cybench — Zhang et al. (2024) · 40 CTF tasks with human first-solve time as the difficulty axis.
- MLE-bench — Chan et al. (2024) · 75 Kaggle competitions run end-to-end by an agent.
- ARC-AGI and On the Measure of Intelligence — François Chollet (2019) · Reasoning with world knowledge factored out; the leaderboard reports cost per task alongside accuracy.
- HarmBench — Mazeika et al. (2024) · 510 harmful behaviours; the refusal benchmark shown live on HELM.
- AIR-Bench 2024 — Zeng et al. (2024) · 314 risk categories and 5,694 prompts derived from regulation and company policy.
- Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG) — Zou et al. (2023) · Optimised suffixes that transfer from open weights to closed APIs.
- US & UK AI Safety Institute joint pre-deployment evaluation of OpenAI o1 — NIST (2024) · What the voluntary pre-release testing regime actually produces.
- Clio: Privacy-Preserving Insights into Real-World AI Use — Tamkin et al., Anthropic (2024) · Model-assisted clustering of real traffic; the realism evidence.
- MedHELM — Bedi et al. (2025) · 121 clinical tasks from 29 clinicians. Note: the executable lecture's link for MedHELM points at Clio's arXiv ID by copy-paste; this is the correct paper.
- Proving Test Set Contamination in Black Box Language Models — Oren, Meister, Chatterji, Ladhak, Hashimoto (2023) · The exchangeability test behind "route 1".
- Language model developers should report train-test overlap — Deng et al. (2024) · "Route 2": make disclosure a norm rather than a courtesy.
- Do Large Language Model Benchmarks Test Reliability? (platinum benchmarks) — Vendrow et al. (2025) · Cleaned benchmark variants showing how much of the residual error is label noise.
- Introducing SWE-bench Verified — OpenAI (2024) · The human-validated subset built after errors were found in the original.
- DataComp-LM — Li et al. (2024) · Fixed training pipeline, data as the only free variable; the template for evaluating methods, and for Assignment 4.
- Exploring the Limits of Language Modeling — Jozefowicz et al. (2016) · The 51.3 → 30.0 perplexity result on the One Billion Word Benchmark.
- The LAMBADA dataset — Paperno et al. (2016) · Cloze prediction that requires broad discourse context.
- HellaSwag — Zellers et al. (2019) · Likelihood-scored commonsense completion, mined from ActivityNet and WikiHow.
- Establishing Task Scaling Laws via Compute-Efficient Model Ladders — Bhagia et al. (2024) · Conditional perplexity on downstream tasks as the scaling-law target.
- HELM capabilities leaderboard — Stanford CRFM · The live instance browser Percy uses; the fastest way to actually read predictions.
- Artificial Analysis and OpenRouter rankings · Quality-versus-price frontiers, and revealed preference by token volume.
- Holistic Evaluation of Language Models (HELM) — Liang et al. (2022) · Field map extra. The framework underneath every leaderboard shown in the lecture: multi-metric, scenario-by-scenario, predictions published.
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — Zheng et al. (2023) · Field map extra. The systematic study of judge biases — position, verbosity, self-preference — that AlpacaEval and WildBench inherit.
- Length-Controlled AlpacaEval — Dubois et al. (2024) · Field map extra. The fix for the length-gaming anecdote, and a clean worked example of debiasing an automatic evaluator.
- Lessons from the Trenches on Reproducible Evaluation of Language Models — Biderman et al. (2024) · Field map extra. The practical companion: why two harnesses disagree on the same benchmark, and what to pin so yours reproduces.
- modded-nanogpt (the nanoGPT speedrun) — Keller Jordan et al. · Field map extra. The live example of a strict method evaluation: fixed data and hardware, race to a target validation loss.
Exercises
- Read fifty instances before you trust a number code — Pick one benchmark you already report (MMLU, GSM8K, IFEval, whatever your project uses). (1) Pull 50 random instances with the exact prompt sent to the model. (2) Score them and split into correct/incorrect. (3) Hand-label every incorrect one into: model error, ambiguous question, wrong reference, or formatting/parsing failure. (4) Do the same for 15 correct ones and flag any that are right for the wrong reason. (5) Report the label-noise and parse-failure rates alongside your accuracy. A good answer states what fraction of the residual error is not model error, and names one prompt-format change that would move the score without changing the model.
- Perplexity versus accuracy under a controlled intervention code — Using your Assignment 1/4 training code, train two small models that differ in one deliberate way (e.g. 10% vs 30% of the data filtered out by one heuristic). (1) Measure held-out perplexity on a disjoint split. (2) Measure conditional perplexity of the gold answer on a downstream task. (3) Measure accuracy on that same task. (4) Repeat at two model sizes. A good answer shows which of the three metrics separates the two runs cleanly at small scale and which is still noise, and connects that to why the A4 leaderboard ranks by validation loss rather than benchmark accuracy.
- Build the four-question card for a benchmark you did not design — Take a benchmark from a recent model release table. Fill in Percy's four stages: inputs (coverage, tail, adaptation), the call (prompting, CoT, tools, model-or-system), scoring (reference quality, metric, cost, error asymmetry), interpretation (deployability, contamination, method-or-model). A good answer ends with one sentence naming the single decision this benchmark can legitimately inform, and one naming a decision people routinely use it for that it cannot support.
- Design a refusal eval that cannot be topped by refusing everything — Sketch an evaluation protocol that separates capability from propensity for a specific harm category. Specify: the harmful set, the matched benign set that a hostile-but-legitimate user would send, the pairing that makes over-refusal visible, and how you would report the result as a two-dimensional score rather than one number. A good answer explains what changes if the model ships as open weights instead of behind an API, and why Cybench's dual-use problem does or does not apply to your category.