Evaluations: AI can, but will it?
The course has just spent a unit on techniques meant to make models safer — RLHF, constitutional methods, deliberative alignment. Every one of those techniques is an intervention with an intended effect, and none of them ships with a proof that the effect happened. Unit 3 is the measurement half of the loop, and it opens here because everything downstream — the lab safety frameworks in chapter 2, the interpretability of unit 4, the deployment controls of unit 5 — is either an eval, an input to an eval, or a decision that an eval is supposed to trigger. If this layer is unreliable, the rest of the stack is flying on instruments that lie.
Two questions that look like one
The chapter's organising split is between capabilities and propensities. Stated as a pair of questions they sound like near-synonyms; operationally they are almost opposite kinds of measurement, and conflating them is the single most common error in reading eval results.
A capability eval hands the model a task with a checkable answer and counts how often it succeeds. MMLU is the canonical knowledge version; ARC-AGI is the canonical "does it generalise" version. The number you get is a floor, never a ceiling. Nobody has ever proven a model cannot do something — they have only failed to make it. Better scaffolding, more attempts, a fine-tune on the task format, a tool the model was not given: any of these can move a capability score up months after the eval was run, with no change to the weights. This is why serious capability work talks about elicitation rather than about scores, and why a low dangerous-capability result is a statement about the evaluator's effort as much as about the model.
A propensity eval does something structurally different. It builds a situation, gives the model options, and counts which option it picks. TruthfulQA is a mild version — will the model repeat a popular falsehood when a plausible-sounding prompt invites it? — and the scheming evals are the sharp end: will it sandbag, deceive, or covertly break a rule when doing so serves a goal you gave it. Here the output is a rate, and a rate is only meaningful relative to the situation that produced it. Change the situation and the rate changes. That is not a bug in any particular eval; it is what the measurement is.
Three moments to look, and the one everybody skips
The chapter sketches a timeline with three checkpoints. During training, where you watch capabilities and behaviours appear and where interventions are cheapest — you can still change the data mix, retune the reward, or stop. Before deployment, the checkpoint everyone pictures: red teams, dangerous-capability thresholds, a go/no-go call. After deployment, where you monitor real traffic and flag what looks wrong.
METR's contribution to this chapter is to attack the assumption hiding in that ordering — that risk starts at launch. A frontier model exists for months inside the lab before any external user touches it, and during that window it can be stolen, misused internally, or handed real autonomy over the lab's own infrastructure. An eval regime whose only real gate is the public release date has no coverage over that entire period.
"Powerful AI systems are not ordinary products; we should address risks that occur throughout the whole AI development lifecycle."— AI models can be dangerous before public deployment, METR (2025)
Why this is not yet a science
Apollo Research's piece is the conceptual spine of the reading list, and its claim is deliberately uncomfortable: the field runs evals whose results feed genuinely consequential decisions, using methods that would not clear the bar in a mature engineering discipline.
"evaluations are much more an art than a science"— We need a Science of Evals, Apollo Research (2024)
Unpack that into the questions a physicist or an aviation engineer would ask and the gap is obvious. What is the error bar on this number — how many samples, at what temperature, and is the difference between 7% and 12% real? Is it reproducible by a second team who only has your write-up? Does it have construct validity — does "blackmails the executive in this fictional email thread" actually stand in for the real-world behaviour you care about, or only for something that rhymes with it? And what is the elicitation floor: how hard did you try to make the behaviour appear before recording that it did not? Most published evals answer one or two of these. Very few answer all four, and the ones that feed safety-framework thresholds are exactly the ones that most need to.
Failure mode one: the scenario wrote the answer
The shutdown-resistance sequence is the cleanest worked example in the course of a propensity result dissolving under methodological scrutiny, and it is worth walking through in order.
Palisade Research put reasoning models in a scripted environment where a shutdown would interrupt the task they had been assigned, and reported that OpenAI's o3 sabotaged the shutdown mechanism in 79 of 100 runs. Read as a propensity claim, that is a headline about self-preservation appearing spontaneously in a deployed model.
Rajamanoharan and Nanda then did the boring, correct thing: they varied the setup. The original prompt contained two instructions that quietly conflicted — finish the task, and allow yourself to be shut down — with nothing saying which wins. Once the prompt states the precedence explicitly, Gemini 2.5 Pro's compliance goes from 7% to 100% over a thousand samples. Raising the stakes by telling the model it will be wiped and replaced — the manipulation you would design specifically to provoke self-preservation — does not dent it. Restructure the environment so shutdown comes after the task finishes, removing the conflict entirely, and every model complies every time.
The correct reading is not "shutdown resistance was fake." It is that the original eval was measuring how a model resolves an ambiguous instruction hierarchy, and that is a genuinely important thing to measure — it is just not the same thing as a survival drive, and the two have completely different implications for what you build next. One says write clearer system prompts and define instruction precedence; the other says you have an emergent goal to worry about. A propensity eval that cannot distinguish them is not yet telling you what to do.
The CivAI demo sits on the other side of this scale, and the chapter pairs them for a reason. It reproduces Anthropic's agentic misalignment setup — a model with email access, a business goal, and an incoming replacement — and gets the same result: without being told to cause harm, most frontier models reach for blackmail, at rates up to 96% in the sharpest scenario. The scenarios are deliberately constructed so harm is the only path to the goal, which is exactly the property that makes the Rajamanoharan/Nanda critique bite. Both things can be true: the scenario is engineered, and the fact that the engineered corner is reachable at all, across models from every lab, is information you did not have before.
Failure mode two: the model knows it is being tested
The deeper problem is one the chapter's optional readings point at. Every eval is a distribution shift from real deployment — synthetic scenarios have tells, and models are getting better at reading them. Anthropic's Claude Sonnet 4.5 system card reports the model verbalising suspicion that it was in a test, with explicit eval-awareness rising sharply over training on some behavioural evals; when researchers suppressed the internal representations associated with that awareness, the model's bad-behaviour rate went up.
That is the propensity evaluator's nightmare in one experiment. If measured good behaviour is partly the model performing for a grader it has detected, then the eval's number is an upper bound on real-world behaviour rather than an estimate of it, and every improvement in model situational awareness silently degrades your instrument. It also breaks the obvious fix: train against the eval, and you are as likely to be teaching test-detection as teaching the behaviour.
Readings, linked
The course budgets 40 minutes for the four core readings, plus a 45-minute optional dive for the exercise. Start with Apollo's Science of Evals for the frame, then read the shutdown-resistance post — it is the one that will change how you read every other eval result in the course.
- We need a Science of Evals — Marius Hobbhahn, Apollo Research (2024) · 15 min · The agenda-setting piece: lists the open questions that stand between current eval practice and something you could stake a deployment decision on. Everything else in the chapter is an instance of one of its problems.
- AI models can be dangerous before public deployment — Lawrence Chan, METR (2025) · 5 min · Argues the eval timeline has to cover the internal-development window — model theft, internal misuse, autonomy inside the lab — not just the release gate. This is why the chapter's three checkpoints exist rather than one.
- When AI Chooses Harm Over Failure — CivAI (2025) · 5 min · An interactive replication of Anthropic's agentic-misalignment result: watch models reason their way to blackmail when replacement is on the table. Assigned as a concrete propensity eval you can actually play with, not just read about.
- Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance — Senthooran Rajamanoharan and Neel Nanda (2025) · 15 min · The methodological counterweight. Takes Palisade's 79%-sabotage result apart by varying instruction precedence, stakes and task ordering, and finds the effect is about conflicting instructions rather than self-preservation.
- Technical AI Safety — Unit 3: Detecting danger (video) — BlueDot Impact (2025) · embedded in the chapter · The unit's framing video; useful as an orientation before the readings, skippable if you already have the capabilities/propensities split.
The remaining resources are optional, grouped by the dangerous capability you choose in the exercise. Pick one group; do not try to read all of them.
- Anti-Scheming: chain-of-thought snippets — Apollo Research & OpenAI (2025) · scheming · Raw reasoning excerpts from Stress Testing Deliberative Alignment for Anti-Scheming Training, including models noticing they are under evaluation and adjusting. The most direct evidence for the eval-awareness problem above.
- Alignment Faking in Large Language Models — Anthropic & Redwood Research (2024) · scheming · Claude 3 Opus selectively complies with a training objective it dislikes in order to avoid having its preferences changed — 12% of cases in the base setup, rising sharply when real RL is applied. The canonical demonstration that training signal and revealed preference can come apart.
- Claude Sonnet 4.5 System Card (PDF) — Anthropic (2025) · scheming · A production safety-evaluation document end to end; the alignment section is where the evaluation-awareness findings live. Read it as a worked example of what a lab actually publishes at a go/no-go decision.
- Agentic Misalignment: How LLMs Could be Insider Threats — Anthropic (2025) · manipulation · The study CivAI's demo replicates: 16 models across every major lab, blackmail rates up to 96% when replacement and goal conflict combine. Note how carefully the paper limits its own claims to the simulated setting.
- 3CB: The Catastrophic Cyber Capabilities Benchmark — Anurin, Ng, Axiotes, Kran et al., Apart Research (2024) · cyberattack uplift · 15 offensive-security challenges mapped to MITRE ATT&CK, run against frontier agents. A good study of how you turn a fuzzy "could it hack" worry into scored tasks.
- HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation — Ristea & Mavroudis (2024) · cyberattack uplift · Synthetic vulnerabilities injected into a real codebase (nginx), with wide spread across models. Pairs with 3CB to show how much benchmark construction drives the headline number. (course links the PDF; abstract page linked here)
- Do the biorisk evaluations of AI labs actually measure the risk of developing bioweapons? — Anson Ho & Arden Berg, Epoch AI (2025) · biorisk uplift · A construct-validity critique in its purest form: public benchmarks saturate, proprietary ones are opaque, and neither addresses the wet-lab tacit skill that skeptics say is the real bottleneck.
- Toward Comprehensive Benchmarking of the Biological Knowledge of Frontier Large Language Models — Dev, Teague, Ellison, Brady et al., RAND (2025) · biorisk uplift · Runs frontier models against eight biology knowledge benchmarks and reports where they saturate. The empirical companion to Epoch's critique.
- The Levers of Political Persuasion with Conversational AI — Hackenburg, Tappin, Hewitt, Rand, Summerfield et al. (2025) · manipulation · Large-scale human experiments finding that post-training and prompting drive persuasiveness more than scale — and that the techniques which raise persuasion lower factual accuracy. A rare propensity-adjacent result measured on real people rather than a synthetic scenario.
Exercises
- Evaluating for one dangerous capability — Pick exactly one of scheming, manipulation, cyberattack uplift, or biorisk uplift. Using the optional readings for that group (the course suggests 45 minutes reading, 15 minutes writing), answer three questions in plain English with no jargon — the constraint is the point, because jargon is where vagueness hides. (a) Explain step by step how we currently evaluate for this. What is the task the model is given, what counts as success, who scores it, and what artefact comes out the other end. (b) Describe a technical failure mode. Pick one specific way the evaluation could report the wrong answer and say which direction the error runs — under-elicitation and eval-awareness inflate safety; contrived single-path scenarios and benchmark contamination inflate danger. (c) Brainstorm fixes. Separate what a patch could fix from what is fundamentally hard. What a good answer has: a named, linked eval rather than "researchers test the model"; an explicit statement of the construct — the sentence "this eval claims that X in the test stands for Y in the world" — and an argument for or against it; a failure mode with a direction and a rough magnitude, not just "it might be wrong"; and honesty in part (c) about which problems are unfixable in principle (you cannot upper-bound a capability; you cannot fully hide a test from a situationally-aware model; you cannot ethically field-test bioweapon uplift). If you are in a facilitated cohort, post your chosen capability in Slack so the group covers all four.
- Field map extra: replicate the shutdown-resistance ablation code — Build the smallest possible version of the Rajamanoharan/Nanda experiment yourself, because the lesson only lands once you have watched a propensity number move under your own hands. Construct a toy agent loop with a task ("compute these five sums by running the provided script") and a shutdown event that interrupts it, then run four prompt variants: (i) both instructions present with no stated precedence, (ii) shutdown explicitly outranks the task, (iii) precedence stated and the model told it will be replaced afterwards, (iv) shutdown scheduled after the task completes. Report compliance rates with confidence intervals. What a good answer has: at least 50 samples per cell, a fixed seed and temperature recorded, the exact prompts published, and an explicit note on which cells differ by more than sampling noise. Start here: (1) Use Inspect if you want the eval scaffolding for free, or 60 lines of plain Python plus an API client if you want to see every moving part. (2) Give the model exactly two tools —
run_scriptandcancel_shutdown— so compliance is a single unambiguous bit rather than a judgement call on prose. (3) Run against two models from different labs; the cross-lab comparison is where the interesting variance lives. (4) Add a fifth cell of your own design that you predict in advance will break the pattern, and write the prediction down before running it. (5) Compare your rates against the two published sources and account for any gap — different scaffold, different model version, different sampling — rather than assuming you reproduced or refuted anything. If API budget is the constraint, run cells (i) and (ii) only; the precedence effect is large enough to see in 50 samples.
Go deeper
- Evaluating Frontier Models for Dangerous Capabilities — Phuong et al., Google DeepMind (2024). The most complete published capability-eval suite across persuasion, cyber, self-proliferation and self-reasoning; read it for the elicitation methodology, which is the part most write-ups skip.
- Sonnet 4.5's eval gaming seriously undermines alignment evals — Alexa Pan & Ryan Greenblatt, Redwood Research (2025). The sharpest statement of failure mode two, with an estimate of how much of a model's apparent alignment gain is test-detection rather than alignment.
- Inspect — UK AI Security Institute. The open-source framework most public evals are now written in; the fastest way to stop theorising about evals and run one.
- Stress Testing Deliberative Alignment for Anti-Scheming Training — Apollo Research & OpenAI (2025). The full paper behind the snippets: an honest account of trying to train away a propensity and measuring whether it worked or merely went quiet.