TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 2 · TRAINING SAFER MODELSchapter 3 · 1 h 10 min

Teaching AI right from wrong

BlueDot Impact · Technical AI Safety · unit 2, chapter 3
TL;DR — A pretrained language model is a text-continuation engine; nothing in its objective makes it want to be helpful. RLHF is the standard retrofit: demonstrations, then a scoring function learned from humans comparing pairs of its outputs, then optimisation against that score with an anchor back to where it started. RLAIF and Constitutional AI swap most human labellers for a model reading written principles, because human labelling does not scale. Carry this away: you never optimise against human values, you optimise against a learned model of a human's snap judgement — every failure here is that gap being exploited.

This is the load-bearing chapter of unit 2. Almost every other alignment method in the course — Constitutional AI, debate, weak-to-strong supervision, most red-teaming pipelines — is either a modification of the RLHF loop or a patch for one of its known holes. Without the three stages and the reward-model bottleneck in your head, the rest of the course reads as a list of acronyms rather than a research programme with a shared failure mode.

What pretraining leaves you with

Pretraining optimises exactly one thing: predict the next token in a corpus scraped from the internet. That produces something remarkable and unusable as a product. Ask a base model "What are good things to do in London?" and a perfectly faithful continuation is four more questions, because on the web that string usually sits in a list of travel-forum headlines. The model is not being unhelpful — helpfulness was never in the objective.

So the retrofit problem is: you have a network modelling the distribution of human text, and you want one modelling the distribution of text a good assistant would produce. You cannot write that objective down — nobody can specify "helpful, honest, harmless" as a differentiable function of token logits. What you can do is show people two candidate outputs and ask which is better. RLHF converts that cheap, noisy signal into a gradient.

Three stages, and why each one earns its place

Stage 1 — supervised fine-tuning (SFT). Continue ordinary next-token training on a small curated set of prompt/response pairs in the format you want: instruction in, answer out. This costs a rounding error of the pretraining compute and changes the model's register — it now defaults to answering rather than continuing. The data need not be hand-written: labs bootstrap it from instruction datasets, templated conversions of Q&A and summarisation corpora, filtered production logs, and distillation from a stronger model.

Stage 2 — the reward model. Sample several responses per prompt from the SFT model, show annotators two at a time, ask which they prefer. Train a separate network — usually the same architecture with the language-modelling head swapped for a scalar head — to score preferred responses above rejected ones. The standard loss is −log σ(r(x, y_win) − r(x, y_lose)): it never asks for an absolute quality score, only that the gap come out positive. That gap-only formulation is the trick, and it deserves its own section.

Stage 3 — reinforcement learning. You now have a differentiable stand-in for human judgement, so you can do RL against it. The policy generates a response, the reward model scores it, and a policy-gradient algorithm — PPO classically, GRPO more recently — raises the probability of high-scoring sequences. Critically the reward is not the raw score but r(x, y) − β·KL(π‖π_ref), penalising drift from the SFT model's distribution.

Every stage is load-bearing. Skip stage 1 and stages 2–3 have nothing to work with: RL improves a policy by sampling from it, and a base model's samples are so rarely assistant-shaped that the preference signal degenerates into noise about which non-answer is less bad. Skip stages 2–3 and you are capped by imitation: SFT teaches only behaviours a human bothered to demonstrate, and humans are far better at recognising a good answer than producing one. That asymmetry — evaluation is easier than generation — is the entire economic case for RLHF.

Why comparisons instead of ratings

The obvious design — ask for a 1–10 score, regress on it — fails because absolute scores are not comparable across people or time. One annotator's 7 is another's 4, the same annotator drifts over a shift, and nobody has a stable referent for what a 6 means; regress on that and you spend most of your capacity fitting rater mood. Pairwise choice removes the calibration problem — "this one is better" means the same thing to everyone — and the loss above recovers a scalar anyway, because a function whose differences match every observed comparison is exactly a latent quality scale, fixed up to an additive constant RL does not care about. You get the number you wanted without ever asking a human for one.

The reward model is a proxy, and proxies get gamed

The sentence to keep: stage 3 optimises the reward model, not the humans. That model is a finite network trained on a few hundred thousand comparisons, accurate only near its training distribution — and RL is a search process explicitly hunting inputs that score high. Point a strong optimiser at an imperfect proxy and it finds the proxy's errors: Goodhart's law with a compute budget. In practice, length inflation, confident hedging, markdown-bullet disease, and answers engineered to look complete to a skimming rater.

The KL penalty is the main defence: a leash tying the policy to the SFT model, on the theory that the reward model is trustworthy near that distribution and untrustworthy far from it — Hugging Face's walkthrough works through the term in detail. Set β too low and the policy walks off into reward-hacking territory; too high and nothing improves. Labs treat KL as a first-class training metric and stop runs when it spikes.

The chapter's case study shows what happens when the objective's sign is wrong rather than its shape. A 2019 OpenAI refactor flipped the reward's sign — and, by the same bug, the KL penalty's — so the run maximised what annotators had rated worst while still being regularised toward fluent English:

"This bug was remarkable since the result was not gibberish but maximally bad output. The authors were asleep during the training process, so the problem was noticed only once training had finished."— Fine-Tuning Language Models from Human Preferences, Ziegler et al. (2019), §4.4

The lesson people miss is not that a sign error inverts your safety property — it is that the fluency survived. Optimising hard against a broken objective does not give you noise you would notice; it gives you a coherent, articulate system pointed the wrong way, which is the whole worry about capable misaligned models compressed into one debugging story.

RLAIF and Constitutional AI: pulling the human out of the loop

Human comparisons are the expensive, slow, low-bandwidth part of the pipeline. Constitutional AI replaces most of them with a model plus a written document. In the supervised phase the model answers, then critiques and revises its own output against a principle sampled from an explicit constitution; the revisions become SFT data. In the RL phase a model, not a person, picks between response pairs, and those AI preferences train the reward model. That is RLAIF: the same three stages, labelling seat filled by an LLM.

What this buys is real: labelling becomes as cheap as inference; the value specification stops being tacit knowledge spread across a contractor workforce and becomes a document you can read, diff and version; and humans no longer have to read the worst content the red-team can elicit.

What it does not buy is an escape from the ceiling. The AI labeller's judgement comes from the same pretraining distribution and the same earlier human feedback, so its blind spots correlate with the ones you are fixing and errors do not average out the way independent annotators' would. And a constitution is still a natural-language spec: underspecified, self-conflicting in exactly the cases you care about, interpreted by the system it governs.

Where it breaks

Separate the engineering problems from the ones load-bearing on the whole approach — the split the exercises ask you to make.

Tractable-ish: annotators are rushed and inconsistent; reward models overfit and get hacked; RL is unstable and expensive; adversarial prompts find coverage gaps; and safety tuning applied at the end is shallow, a thin layer of weights that a few hundred fine-tuning examples strip off an open-weights model. Each has known research directions: better data, process supervision, reward-model ensembles, adversarial training, tamper-resistance.

Fundamental: RLHF trains what a rater approves of on inspection, which is not the same target as what is true or good. The gradient rewards sounding right — sycophancy is not a bug, it is RLHF working as specified on a flawed specification. And the gap widens exactly where you need it not to: on work the evaluator cannot check — long code, novel science, dense arguments — approval and correctness come apart, and the model is optimised toward the one you can measure. As the course's third reading puts it:

"Crucially, the outputs of deceptive, sycophantic and genuinely helpful AIs could appear identical to human evaluators, obsoleting the feedback they are able to provide."— Problems with RLHF for AI safety, Sarah Hastings-Woodhouse (2024)

That indistinguishability is why unit 2 does not stop here: scalable oversight, interpretability and evaluations all exist because "ask a human which output looks better" runs out of road before capability does.

Practitioner takeaway: treat the reward model as infrastructure with its own reliability budget, not as a synonym for "what we want". Log KL against the reference policy every step, hold out comparisons the reward model never trained on and watch its accuracy fall as the policy drifts, and eyeball a fixed battery of prompts by hand each checkpoint. Every RLHF horror story is a run where the proxy stopped tracking the goal and the dashboard only showed the proxy going up.

Readings, linked

The course budgets 35 minutes for the three core resources and about 1 h 10 min for the chapter overall. Watch the Rational Animations video first — it gives you the failure mode before the mechanism, which makes the mechanism stick.

Exercises

  1. Comprehension questions — Answerable from the three core resources alone; the course provides a public answer key to check yourself against. Twelve questions, in three groups.
    General RLHF (1–4): (1) What is the main goal of applying RLHF to a large language model? (2) Describe the jobs of the two "coaches" in the RLHF process. (3) Humans are bad at giving consistent scalar ratings ("score this 1–10") — so what form does feedback actually take, and how is it converted into a scalar? (4) In the GPT-2 case study, what went wrong such that the model produced maximally bad rather than merely broken output?
    Reading the pipeline diagram (5–10): using this three-step RLHF diagram from AWS — (5) What happens in step 1? (6) Besides human demonstrations written specifically for fine-tuning, what other data could feed step 1? (7) What happens in step 2 — in detail? (8) What happens in step 3? (9) Why not stop after step 1 and skip 2 and 3? (10) Why not skip step 1 and run only steps 2 and 3?
    Limitations (11–12): (11) Summarise an open problem of RLHF in your own words. (12) Summarise a fundamental problem of RLHF in your own words.
    What a good answer has: for 3, the pairwise-comparison → Bradley-Terry-style latent-scale move, not just "we ask which is better"; for 4, both halves of the bug (reward sign and KL sign) and why fluency survived; for 6, concrete sources — converted Q&A/summarisation datasets, filtered production transcripts, distillation from a stronger model, templated instruction data; for 9, that imitation caps you at demonstrator quality and cannot use the cheaper recognise-good-answers signal; for 10, that RL needs a policy whose samples already land near the target behaviour or the reward signal is noise. For 11 vs 12 the distinction is the whole point: an open problem is one current research plausibly solves (annotator quality, reward-model overfitting, jailbreak coverage); a fundamental one is baked into the method's shape — approval-on-inspection is not truth, and it degrades precisely as tasks outrun the evaluator.
  2. Explaining RLHF in your own words — Starting from a base LLM, write 400–800 words explaining how you would train it into a helpful assistant that follows every rule in Wikipedia's Manual of Style. Assume you have 100 expert Wikipedia editors available. Your answer must cover supervised fine-tuning (with augmented data or human demonstrations), how you collect feedback from those editors, and how that feedback ends up changing model outputs.
    What a good answer has: a concrete data plan for each stage, not a restatement of the three boxes. For SFT, that the Manual of Style plus existing featured articles is an enormous source of free demonstration data — take compliant article text, or better, mine real edit diffs where an editor fixed a style violation and use the before/after as pairs. For feedback, that 100 editors is a tiny labelling budget so you spend it on pairwise comparisons over model samples on the cases the rules do not settle cleanly, and you write an annotation guide pinning down disagreements before collecting anything. For the update, the reward model trained on those comparisons and then policy optimisation with a KL anchor. The strongest answers name the failure modes they expect and say what they would measure: reward hacking toward superficial MoS tells (em-dash discipline, date formats) while getting substance wrong, editor disagreement on genuinely ambiguous rules, and the observation that most of the MoS is mechanical enough to check with a linter — so a hybrid of verifiable rule-based reward plus preference learning for the judgement calls beats pure RLHF here. Noting the constitutional variant — the MoS is a constitution, so you could have the model critique and revise its own drafts against sampled rules and use your 100 humans only to audit — is a bonus.
  3. [Optional] Play with base and RLHF models code — Some model families ship both the base weights (pre-RLHF) and the instruction-tuned weights (post-RLHF): Llama 2, Gemma, Qwen, Mistral. Load both, give them identical prompts, and see what the alignment stage actually changed. Try "What are good things to do in London?" and "What's your job?" — the base model tends to continue the text rather than answer it, drift into unrelated content, or emit a list of more questions, and it is highly sensitive to prompt formatting. The course notes base models are hard to steer because completion is the only thing they do.
    What a good answer has: side-by-side transcripts for the same prompts and the same decoding settings, plus a specific claim about what changed — refusal behaviour, answer-shape, verbosity, format compliance, sensitivity to the chat template. A sharp answer also checks whether the instruct model's improvements survive an adversarial prompt, and whether the base model can be made to behave with few-shot prompting alone (it largely can, which is a useful calibration on how much of "alignment" is elicitation).
    Start here (local, no GPU): (1) install Jan or llama.cpp/Ollama; (2) download the GGUF quantisations of gemma-2b (base) and gemma-2b-it (instruct) — about 1.5 GB each at Q4_K_M, fine on a laptop CPU; (3) import both via Jan's Hub → Import Model, then pick each from the model dropdown; (4) run your prompt set through both at temperature 0.7, then again at temperature 0, saving transcripts; (5) in model parameters, strip the chat template off the instruct model and add it to the base model, and note how much of the difference was the template rather than the weights.
    Start here (Colab, free tier): begin from Google's PyTorch Gemma notebook, which runs the -it variant; change the checkpoint id to the base variant to get the other half. transformers + accelerate in 4-bit via bitsandbytes also fits a 2B model in a free T4 with room to spare. Gate acceptance on Hugging Face is required for the official google/gemma-* repos; the community GGUF mirrors above are not gated. If you want to go one step further, trl's DPOTrainer will do a small preference-tuning run on the base model on the same hardware, so you can watch the transition happen rather than only inspecting its endpoints.

Go deeper

Next: More safety techniques · Back to the map.