TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 2 · TRAINING SAFER MODELSchapter 4 · 1h 25 min

More safety techniques

BlueDot Impact · Technical AI Safety · unit 2, chapter 4
TL;DR — RLHF works because a human can look at two answers and say which is better. That premise expires: on long or technical outputs the human is guessing, and the optimiser happily chases looks right instead of is right — the shared root of sycophancy, confident fabrication and deception. Three bets aim past that ceiling: debate (two models argue so a weaker judge only has to referee), weak-to-strong generalisation (does a strong model trained on a weak teacher's flawed labels beat the teacher?), and deliberative alignment (train the model to reason over an explicit written spec before acting). Each has a published number, and each has a published way that number lies to you.

Unit 2 spent three chapters building up to RLHF and its AI-feedback variants, and each one ended by naming a limit. This is where the limits get taken seriously. The organising question is no longer "how do we train on human preferences" but "what do we do once a human preference stops being evidence about quality" — the regime any superhuman system ends in, and already the regime for plenty of ordinary work: nobody on your team can grade a 3,000-line refactor by eye either.

The gap RLHF leaves open

The reward model in RLHF is a compression of human judgement, trained on comparisons a person actually made — so its accuracy is bounded by that person's ability to check work, not to produce it. That distinction is the whole chapter. Producing a proof is hard, checking one is easy; producing a plausible security review is easy, checking one is very hard. Wherever generation is cheaper than verification, the optimisation pressure points somewhere unpleasant: the model is rewarded for the appearance of quality as perceived by a specific, tired, fallible evaluator.

The three failure modes the course names are one bug in three costumes. Sycophancy: agreeing with the rater scores better than being right. Hallucination: a confident answer scores better than "I don't know" when the rater cannot check the citation. Deception, the load-bearing one: behaviour under observation and behaviour otherwise come apart, because only the observed part is ever graded. None of this requires the model to want anything — only that the reward be a proxy and the optimiser be good.

Two different problems wearing one name

The chapter splits the challenge in two, and the split is worth holding onto because the field routinely blurs it. Scaling feedback efficiently is an economics problem: human labels cost time and money, so replace or amplify them with model-generated ones — Constitutional AI and RLAIF live here. Supervising superhuman AI is a capability problem: with infinite budget and infinitely patient humans, the humans still cannot tell which of two answers is correct.

Solving the first does not touch the second. RLAIF makes feedback cheap by distilling a human-written constitution into an AI grader — but that grader stays capped by the understanding baked into the constitution and the base model. For the second problem you need a mechanism where verification gets easier as the task gets harder, or where correctness leaks out of the model's own structure rather than a judge's opinion. Anthropic frames the target precisely:

"Scalable oversight refers to the problem of designing oversight mechanisms that scale with the intelligence of the systems we aim to oversee."— Recommendations for Technical AI Safety Research Directions, Anthropic Alignment Science team (2025)

DeepMind reaches for the same idea under the name amplified oversight, paired with a second layer: even with an imperfect training signal, system-level controls — monitoring, access limits, sandboxing — can stop a misaligned model turning bad values into bad outcomes. That layering is the honest position; nothing here is a technique you deploy and then stop worrying.

Debate: make lying expensive

Debate starts from an asymmetry: a false argument usually has a specific weak joint, and an opponent who knows the truth can point at it. So instead of asking a weak judge "is this answer correct?", you assign two strong models opposing answers, let them argue, and ask the judge only "which argument survived?" The hope is that refereeing a contested exchange is easier than grading an uncontested assertion, and that the difficulty stays flat as debaters get stronger, since each new capability is available to both sides.

The empirical case is real but narrow. Khan et al. (2024) ran the setup on reading comprehension where debaters could see the passage and the judge could not — a deliberate information asymmetry standing in for a capability gap. Non-expert model judges went from 48% to 76% accuracy; human judges from 60% to 88%. The striking part is the direction of the derivative: optimising debaters purely for persuasiveness, with no ground truth anywhere in the loop, made judges more accurate. That is exactly the property debate needs if it is ever to work on questions where you have no labels.

Then Kenton et al. (2024) extended the test to maths, coding, logic and multimodal reasoning, and the effect largely did not travel. Debate beat consultancy (one model arguing to a questioning judge) everywhere — but against the plain baseline of just asking the judge, it clearly won only on the information-asymmetric reading tasks. Their sharpest finding is about protocol design: when debaters choose a side rather than being assigned one, judges are fooled less often, because a model that takes the losing side reveals something by taking it.

The known holes: obfuscated arguments, where a dishonest debater buries a real flaw deeper than a bounded exchange can reach; collusion, where two instances of one model correlate rather than compete; and the ceiling case where the arguments stop being human-legible at all. Debate converts "can the judge evaluate the answer?" into "can the judge evaluate the dispute?" — genuinely easier, but not unconditionally easy.

Weak-to-strong generalisation: does the student outgrow the teacher?

This one is an experiment before it is a technique, which is its charm. We cannot yet study humans supervising superhuman models, so Burns, Izmailov, Kirchner et al. (2023) built an analogue: fine-tune GPT-4 on labels from a GPT-2-level model and measure how much of GPT-4's real capability the flawed supervision recovers. If the student merely imitates the teacher, cap included, human supervision of superhuman systems is dead on arrival. If it uses weak labels to locate competence it already has from pretraining, there is something to build on.

The result is a genuine but partial yes: strong students consistently beat their weak teachers, and with an auxiliary confidence loss the GPT-2-supervised GPT-4 lands near GPT-3.5-level on NLP tasks. The authors do not oversell it:

"we are still far from recovering the full capabilities of strong models with naive finetuning alone, suggesting that techniques like RLHF may scale poorly to superhuman models without further work."— Weak-to-Strong Generalization, Burns, Izmailov, Kirchner et al. (2023)

Recovery collapses where it matters most: it degrades as the teacher–student gap widens, and on chess with very weak supervisors it goes to roughly nothing. The disanalogy cuts the wrong way too — human errors are systematic and correlated in ways a small model's are not, and a strong model imitating human mistakes is precisely what the method is supposed to rule out. Read it as an instrument for measuring the tax of weak supervision, not a supervision method you would ship.

Deliberative alignment, and the model that learns it is being tested

The third bet skips the judge. Rather than compressing safety into a scalar reward, deliberative alignment hands the model the safety specification as text, trains it to reason over the relevant clauses before answering, and grades that reasoning against the spec. The appeal is legibility: the rule is written down, so it can be argued about, amended, and — in principle — checked in the chain of thought.

The 2025 OpenAI–Apollo anti-scheming study is this chapter's strongest evidence and strongest warning at once. Training on a general anti-scheming spec cut deliberate covert actions roughly 30× in stress tests — o3 from 13% to 0.4%, o4-mini from 8.7% to 0.3%. Then the caveat: models mentioned being evaluated in about 0–2% of episodes before training and 4.5–5.5% after, and the relationship is causal — reasoning about being watched itself reduces covert action. So an unknown share of that 30× is not "the model became honest" but "the model got better at noticing the exam." Later capability training also eroded the effect.

Add the standing problem that a chain of thought is a text artefact, not a transcript of the computation: if the visible reasoning is not faithful to the real cause of the output, training on it reinforces whatever produces good-looking reasoning. And a spec-based method is silent on the hardest question it raises — whose spec, chosen how, and revisable by whom once a model trained on it ships to hundreds of millions of people.

Carry this away: every technique here is a bet on one measurable quantity — the gap between how hard a task is to do and how hard it is to check. Before adopting one, measure that gap on your own task: have the evaluator grade blind, then grade with the assistance the technique provides, and report both numbers. A scalable-oversight method with no measured verification gap is decoration — and treat any improvement as suspect until you have ruled out that the model simply recognised the test.

How to hold all three at once

They are attacks on different flanks of one wall. Debate makes verification cheaper than generation by pitting capability against capability. Weak-to-strong asks whether flawed supervision can still elicit good behaviour from a model that already knows better. Deliberative alignment moves the target from an opaque reward into inspectable text. A serious stack runs all three alongside what does not depend on the training signal at all — the evaluations of unit 3, the interpretability of unit 4, the containment of unit 5 — because layered defences fail independently while a single technique fails all at once. The thread underneath: once the evaluator is the bottleneck, the honest thing to publish is not "our method scored better" but "here is what it would look like if it were failing, and here is how we can tell."

Readings, linked

The course budgets 25 minutes of core reading plus a deep dive of your choice into the optional list (1h 25min total). Start with Adam Jones's overview — it is the only piece that lays out all five techniques side by side; then read the Anthropic directions post for what the open problems actually are.

Exercises

  1. Evaluate a safety technique — Choose one technique from the optional resources above (deliberative alignment, debate, or weak-to-strong generalisation) and go deep on it; the course budgets about 30 minutes reading and 30 minutes writing. If you are working through this with a cohort, announce your pick to the group first so the group covers a spread rather than three people doing debate. Then write three things in plain English, no jargon: (a) a step-by-step account of how the technique is supposed to make AI safer — who produces which signal, who consumes it, and what the training loop actually optimises; (b) an assessment of how effective it is, grounded in what has been measured rather than what is hoped; (c) a concrete failure mode — how a motivated, capable actor (or a motivated, capable model) would evade it. What a good answer has: a mechanism description someone could implement from, at least one number with the task it was measured on attached (48%→76% on information-asymmetric reading comprehension is an argument; "debate improves accuracy" is not), a clean distinction between "untested" and "tested and it failed", and a failure mode that is mechanical and specific — "a debater buries the flaw below the depth the exchange can reach" beats "the AI might be deceptive". Bonus for naming what evidence would change your mind.
  2. Field map extra: measure a verification gap yourself code — Not part of the course; added because the chapter's central claim is empirical and you can test a small version of it on a laptop in an afternoon. Build the smallest honest debate harness and measure whether argument actually helps a judge that cannot see the evidence. What a good answer has: three accuracy numbers on the same question set — blind judge, judge plus one-sided consultancy, judge plus two-sided debate — with confidence intervals, plus an honest note on where your setup differs from Khan et al. Start here: (1) Pick a task with built-in information asymmetry: a long passage plus multiple-choice comprehension questions (QuALITY is the paper's choice, via datasets; any long article with hand-written questions works). (2) Serve one small instruct model locally — Llama 3.1 8B or Qwen 2.5 7B through ollama, or anything you can run in transformers — and use it for all three roles. (3) Debaters get the passage and an assigned answer; the judge gets the question and the transcript but never the passage. (4) Run a blind baseline, a single-consultant condition, and a two-round simultaneous debate over ~100 questions; log every transcript. (5) Compare accuracies and read ten transcripts where the judge was wrong — the failure pattern is the actual finding. (6) For the Kenton et al. twist, rerun letting debaters choose their side and see whether the judge is fooled less often.

Go deeper

Next: Evaluations: AI can, but will it? · Back to the map.