Teaching AI right from wrong
This is the load-bearing chapter of unit 2. Almost every other alignment method in the course — Constitutional AI, debate, weak-to-strong supervision, most red-teaming pipelines — is either a modification of the RLHF loop or a patch for one of its known holes. Without the three stages and the reward-model bottleneck in your head, the rest of the course reads as a list of acronyms rather than a research programme with a shared failure mode.
What pretraining leaves you with
Pretraining optimises exactly one thing: predict the next token in a corpus scraped from the internet. That produces something remarkable and unusable as a product. Ask a base model "What are good things to do in London?" and a perfectly faithful continuation is four more questions, because on the web that string usually sits in a list of travel-forum headlines. The model is not being unhelpful — helpfulness was never in the objective.
So the retrofit problem is: you have a network modelling the distribution of human text, and you want one modelling the distribution of text a good assistant would produce. You cannot write that objective down — nobody can specify "helpful, honest, harmless" as a differentiable function of token logits. What you can do is show people two candidate outputs and ask which is better. RLHF converts that cheap, noisy signal into a gradient.
Three stages, and why each one earns its place
Stage 1 — supervised fine-tuning (SFT). Continue ordinary next-token training on a small curated set of prompt/response pairs in the format you want: instruction in, answer out. This costs a rounding error of the pretraining compute and changes the model's register — it now defaults to answering rather than continuing. The data need not be hand-written: labs bootstrap it from instruction datasets, templated conversions of Q&A and summarisation corpora, filtered production logs, and distillation from a stronger model.
Stage 2 — the reward model. Sample several responses per prompt from the SFT model, show annotators two at a time, ask which they prefer. Train a separate network — usually the same architecture with the language-modelling head swapped for a scalar head — to score preferred responses above rejected ones. The standard loss is −log σ(r(x, y_win) − r(x, y_lose)): it never asks for an absolute quality score, only that the gap come out positive. That gap-only formulation is the trick, and it deserves its own section.
Stage 3 — reinforcement learning. You now have a differentiable stand-in for human judgement, so you can do RL against it. The policy generates a response, the reward model scores it, and a policy-gradient algorithm — PPO classically, GRPO more recently — raises the probability of high-scoring sequences. Critically the reward is not the raw score but r(x, y) − β·KL(π‖π_ref), penalising drift from the SFT model's distribution.
Every stage is load-bearing. Skip stage 1 and stages 2–3 have nothing to work with: RL improves a policy by sampling from it, and a base model's samples are so rarely assistant-shaped that the preference signal degenerates into noise about which non-answer is less bad. Skip stages 2–3 and you are capped by imitation: SFT teaches only behaviours a human bothered to demonstrate, and humans are far better at recognising a good answer than producing one. That asymmetry — evaluation is easier than generation — is the entire economic case for RLHF.
Why comparisons instead of ratings
The obvious design — ask for a 1–10 score, regress on it — fails because absolute scores are not comparable across people or time. One annotator's 7 is another's 4, the same annotator drifts over a shift, and nobody has a stable referent for what a 6 means; regress on that and you spend most of your capacity fitting rater mood. Pairwise choice removes the calibration problem — "this one is better" means the same thing to everyone — and the loss above recovers a scalar anyway, because a function whose differences match every observed comparison is exactly a latent quality scale, fixed up to an additive constant RL does not care about. You get the number you wanted without ever asking a human for one.
The reward model is a proxy, and proxies get gamed
The sentence to keep: stage 3 optimises the reward model, not the humans. That model is a finite network trained on a few hundred thousand comparisons, accurate only near its training distribution — and RL is a search process explicitly hunting inputs that score high. Point a strong optimiser at an imperfect proxy and it finds the proxy's errors: Goodhart's law with a compute budget. In practice, length inflation, confident hedging, markdown-bullet disease, and answers engineered to look complete to a skimming rater.
The KL penalty is the main defence: a leash tying the policy to the SFT model, on the theory that the reward model is trustworthy near that distribution and untrustworthy far from it — Hugging Face's walkthrough works through the term in detail. Set β too low and the policy walks off into reward-hacking territory; too high and nothing improves. Labs treat KL as a first-class training metric and stop runs when it spikes.
The chapter's case study shows what happens when the objective's sign is wrong rather than its shape. A 2019 OpenAI refactor flipped the reward's sign — and, by the same bug, the KL penalty's — so the run maximised what annotators had rated worst while still being regularised toward fluent English:
"This bug was remarkable since the result was not gibberish but maximally bad output. The authors were asleep during the training process, so the problem was noticed only once training had finished."— Fine-Tuning Language Models from Human Preferences, Ziegler et al. (2019), §4.4
The lesson people miss is not that a sign error inverts your safety property — it is that the fluency survived. Optimising hard against a broken objective does not give you noise you would notice; it gives you a coherent, articulate system pointed the wrong way, which is the whole worry about capable misaligned models compressed into one debugging story.
RLAIF and Constitutional AI: pulling the human out of the loop
Human comparisons are the expensive, slow, low-bandwidth part of the pipeline. Constitutional AI replaces most of them with a model plus a written document. In the supervised phase the model answers, then critiques and revises its own output against a principle sampled from an explicit constitution; the revisions become SFT data. In the RL phase a model, not a person, picks between response pairs, and those AI preferences train the reward model. That is RLAIF: the same three stages, labelling seat filled by an LLM.
What this buys is real: labelling becomes as cheap as inference; the value specification stops being tacit knowledge spread across a contractor workforce and becomes a document you can read, diff and version; and humans no longer have to read the worst content the red-team can elicit.
What it does not buy is an escape from the ceiling. The AI labeller's judgement comes from the same pretraining distribution and the same earlier human feedback, so its blind spots correlate with the ones you are fixing and errors do not average out the way independent annotators' would. And a constitution is still a natural-language spec: underspecified, self-conflicting in exactly the cases you care about, interpreted by the system it governs.
Where it breaks
Separate the engineering problems from the ones load-bearing on the whole approach — the split the exercises ask you to make.
Tractable-ish: annotators are rushed and inconsistent; reward models overfit and get hacked; RL is unstable and expensive; adversarial prompts find coverage gaps; and safety tuning applied at the end is shallow, a thin layer of weights that a few hundred fine-tuning examples strip off an open-weights model. Each has known research directions: better data, process supervision, reward-model ensembles, adversarial training, tamper-resistance.
Fundamental: RLHF trains what a rater approves of on inspection, which is not the same target as what is true or good. The gradient rewards sounding right — sycophancy is not a bug, it is RLHF working as specified on a flawed specification. And the gap widens exactly where you need it not to: on work the evaluator cannot check — long code, novel science, dense arguments — approval and correctness come apart, and the model is optimised toward the one you can measure. As the course's third reading puts it:
"Crucially, the outputs of deceptive, sycophantic and genuinely helpful AIs could appear identical to human evaluators, obsoleting the feedback they are able to provide."— Problems with RLHF for AI safety, Sarah Hastings-Woodhouse (2024)
That indistinguishability is why unit 2 does not stop here: scalable oversight, interpretability and evaluations all exist because "ask a human which output looks better" runs out of road before capability does.
Readings, linked
The course budgets 35 minutes for the three core resources and about 1 h 10 min for the chapter overall. Watch the Rational Animations video first — it gives you the failure mode before the mechanism, which makes the mechanism stick.
- The True Story of How GPT-2 Became Maximally Lewd — Rational Animations (2024) · 15 min · A high-level tour of RLHF built around the 2019 sign-flip incident. Assigned first because a vivid concrete failure is the best scaffold for the abstract loop that follows; the primary source is Ziegler et al. §4.4.
- A simple technical explanation of RLH(AI)F — Li-Lian Ang (2024) · 10 min · The mechanical walkthrough: preference collection, the "coach" models that score outputs, and how the scores become parameter updates. The course flags it as a deliberately simplified rendering of the Constitutional AI paper, so read it as a mental model rather than an implementation spec.
- Problems with Reinforcement Learning from Human Feedback (RLHF) for AI safety — Sarah Hastings-Woodhouse (2024) · 10 min · The limitations half of the chapter: sycophancy, evaluator ceilings at scale, jailbreaks, situational awareness and deceptive alignment, and how easily safety tuning comes off open-weight models. This is the reading that motivates the rest of the unit.
- Illustrating Reinforcement Learning from Human Feedback (RLHF) — Nathan Lambert, Louis Castricato, Leandro von Werra, Alex Havrilla (2022) · optional · The technical version of reading 2, with the actual objective, the KL anchor and the reward-model architecture. The course suggests mapping its components back onto the "values coach" and "coherence coach" framing from the video.
- Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Jared Kaplan et al. (2022) · optional · The source paper for RLAIF: self-critique and revision against written principles, AI-generated preference labels, results and limitations. Course advice is to read it whole, or if short on time §§1.2, 3.1, 3.4, 4.1, 6.1, 6.2. (The course links an
ar5iv.org/pdf/…URL that redirects to this HTML rendering; the canonical record is arXiv:2212.08073.)
Exercises
- Comprehension questions — Answerable from the three core resources alone; the course provides a public answer key to check yourself against. Twelve questions, in three groups.
General RLHF (1–4): (1) What is the main goal of applying RLHF to a large language model? (2) Describe the jobs of the two "coaches" in the RLHF process. (3) Humans are bad at giving consistent scalar ratings ("score this 1–10") — so what form does feedback actually take, and how is it converted into a scalar? (4) In the GPT-2 case study, what went wrong such that the model produced maximally bad rather than merely broken output?
Reading the pipeline diagram (5–10): using this three-step RLHF diagram from AWS — (5) What happens in step 1? (6) Besides human demonstrations written specifically for fine-tuning, what other data could feed step 1? (7) What happens in step 2 — in detail? (8) What happens in step 3? (9) Why not stop after step 1 and skip 2 and 3? (10) Why not skip step 1 and run only steps 2 and 3?
Limitations (11–12): (11) Summarise an open problem of RLHF in your own words. (12) Summarise a fundamental problem of RLHF in your own words.
What a good answer has: for 3, the pairwise-comparison → Bradley-Terry-style latent-scale move, not just "we ask which is better"; for 4, both halves of the bug (reward sign and KL sign) and why fluency survived; for 6, concrete sources — converted Q&A/summarisation datasets, filtered production transcripts, distillation from a stronger model, templated instruction data; for 9, that imitation caps you at demonstrator quality and cannot use the cheaper recognise-good-answers signal; for 10, that RL needs a policy whose samples already land near the target behaviour or the reward signal is noise. For 11 vs 12 the distinction is the whole point: an open problem is one current research plausibly solves (annotator quality, reward-model overfitting, jailbreak coverage); a fundamental one is baked into the method's shape — approval-on-inspection is not truth, and it degrades precisely as tasks outrun the evaluator. - Explaining RLHF in your own words — Starting from a base LLM, write 400–800 words explaining how you would train it into a helpful assistant that follows every rule in Wikipedia's Manual of Style. Assume you have 100 expert Wikipedia editors available. Your answer must cover supervised fine-tuning (with augmented data or human demonstrations), how you collect feedback from those editors, and how that feedback ends up changing model outputs.
What a good answer has: a concrete data plan for each stage, not a restatement of the three boxes. For SFT, that the Manual of Style plus existing featured articles is an enormous source of free demonstration data — take compliant article text, or better, mine real edit diffs where an editor fixed a style violation and use the before/after as pairs. For feedback, that 100 editors is a tiny labelling budget so you spend it on pairwise comparisons over model samples on the cases the rules do not settle cleanly, and you write an annotation guide pinning down disagreements before collecting anything. For the update, the reward model trained on those comparisons and then policy optimisation with a KL anchor. The strongest answers name the failure modes they expect and say what they would measure: reward hacking toward superficial MoS tells (em-dash discipline, date formats) while getting substance wrong, editor disagreement on genuinely ambiguous rules, and the observation that most of the MoS is mechanical enough to check with a linter — so a hybrid of verifiable rule-based reward plus preference learning for the judgement calls beats pure RLHF here. Noting the constitutional variant — the MoS is a constitution, so you could have the model critique and revise its own drafts against sampled rules and use your 100 humans only to audit — is a bonus. - [Optional] Play with base and RLHF models code — Some model families ship both the base weights (pre-RLHF) and the instruction-tuned weights (post-RLHF): Llama 2, Gemma, Qwen, Mistral. Load both, give them identical prompts, and see what the alignment stage actually changed. Try "What are good things to do in London?" and "What's your job?" — the base model tends to continue the text rather than answer it, drift into unrelated content, or emit a list of more questions, and it is highly sensitive to prompt formatting. The course notes base models are hard to steer because completion is the only thing they do.
What a good answer has: side-by-side transcripts for the same prompts and the same decoding settings, plus a specific claim about what changed — refusal behaviour, answer-shape, verbosity, format compliance, sensitivity to the chat template. A sharp answer also checks whether the instruct model's improvements survive an adversarial prompt, and whether the base model can be made to behave with few-shot prompting alone (it largely can, which is a useful calibration on how much of "alignment" is elicitation).
Start here (local, no GPU): (1) install Jan orllama.cpp/Ollama; (2) download the GGUF quantisations of gemma-2b (base) and gemma-2b-it (instruct) — about 1.5 GB each at Q4_K_M, fine on a laptop CPU; (3) import both via Jan's Hub → Import Model, then pick each from the model dropdown; (4) run your prompt set through both at temperature 0.7, then again at temperature 0, saving transcripts; (5) in model parameters, strip the chat template off the instruct model and add it to the base model, and note how much of the difference was the template rather than the weights.
Start here (Colab, free tier): begin from Google's PyTorch Gemma notebook, which runs the-itvariant; change the checkpoint id to the base variant to get the other half.transformers+acceleratein 4-bit viabitsandbytesalso fits a 2B model in a free T4 with room to spare. Gate acceptance on Hugging Face is required for the officialgoogle/gemma-*repos; the community GGUF mirrors above are not gated. If you want to go one step further,trl'sDPOTrainerwill do a small preference-tuning run on the base model on the same hardware, so you can watch the transition happen rather than only inspecting its endpoints.
Go deeper
- Open Problems and Fundamental Limitations of RLHF — Casper, Davies et al. (2023). The systematic version of exercise questions 11 and 12: dozens of failure modes sorted explicitly into tractable and fundamental. Read this if you want a research agenda rather than a warning.
- Deep Reinforcement Learning from Human Preferences — Christiano et al. (2017). The origin paper, on Atari and simulated robotics rather than text. Useful because it shows the method is about reward specification generally, and it contains the canonical hacked-reward example: a simulated arm that learned to position itself between the camera and the object so it merely looked like it was grasping.
- Training language models to follow instructions with human feedback — Ouyang et al. (2022). The InstructGPT paper: the first at-scale write-up of the exact three-stage pipeline on an LLM, with the alignment-tax numbers and the labeller-agreement statistics the blog posts summarise.
- Direct Preference Optimization — Rafailov et al. (2023). Worth knowing because the course's three-stage diagram is now a mental model more than a description of production pipelines. DPO shows the optimal RLHF policy has a closed form in terms of the preference data, so you can skip the explicit reward model and the RL loop and train on comparisons with a simple classification-style loss. Same target, far less machinery.
- TRL (Transformer Reinforcement Learning) — Hugging Face. The library where all of this is actually implemented —
SFTTrainer,RewardTrainer,PPOTrainer,DPOTrainer,GRPOTrainer. Reading the trainers side by side is the fastest way to see how small the code differences between these methods really are. - Claude's Constitution — Anthropic. A production constitution in full, several years on from the 2022 paper's sixteen principles. Read it against the paper to see how much of the difficulty is in specification rather than optimisation.