Alignment: RL 1 — reinforcement learning from verifiable rewards
Transcript: cleaned auto-captions with timestamps
This is the second of the two post-training lectures, and it is where the course finally crosses from "how do you get a chat model" to "how do you get a reasoning model." Lecture 15 built RLHF: preference data, a Bradley–Terry reward model, PPO or DPO on top. Lecture 16 opens by finishing that story and then arguing it is a dead end for capability — not because the algorithms are bad, but because the reward is a fitted approximation of something noisy. The escape is to find domains where the reward is not fitted at all, which is the whole content of the phrase "verifiable rewards." Everything after that — GRPO, R1, Kimi, Qwen 3 — is engineering downstream of that one move.
Outline, with timestamps
- 00:37 — DPO recap and the ×PO zoo: SimPO, length-normalized DPO, and why there are dozens of these.
- 05:02 — RL findings are contingent: Ai2's PPO-beats-DPO result, reversed by better SFT in Tülu 3.
- 06:07 — Overoptimization and lost calibration: the two ways RLHF quietly degrades a model.
- 12:07 — The pivot: stop optimizing human approval, go where the reward is checkable.
- 14:18 — PPO in theory: policy gradient → importance-weighted TRPO → ratio clipping.
- 19:11 — PPO in practice: a walk through AlpacaFarm's trainer — rollouts, reward shaping, GAE.
- 28:03 — Why yet another algorithm: the value model costs a second copy of the LM, and DPO is the wrong shape.
- 29:09 — GRPO: advantage = z-score of reward within a group of G answers to one question.
- 35:43 — Is that a legal baseline? Dr. GRPO's two corrections, and where the long CoTs come from.
- 42:50 — Case study 1: DeepSeek-R1 and the controlled R1-Zero setting.
- 54:46 — R1's distillation result, and its two honest negative results: PRMs and MCTS.
- 59:09 — Case study 2: Kimi k1.5 — difficulty filtering, a DPO-flavoured loss, an explicit length reward.
- 68:57 — Why RL infrastructure is hard: rollouts, weight sync into vLLM, wildly uneven batches.
- 73:57 — Case study 3: Qwen 3 — 3,995 RL examples, thinking-mode fusion, a thinking budget.
Finishing RLHF, and finding its ceiling
The first ten minutes close out lecture 15. DPO's trick is worth restating because Kimi reuses it later: assume the policy class is all distributions, solve the KL-regularized RL objective in closed form, notice that the optimal policy pins the reward to β·log(π/π_ref) plus a normalizer, substitute that into the Bradley–Terry likelihood, and you are left with a supervised loss over preference pairs. No reward model, no rollouts. The gradient reads as "push up the chosen response, push down the rejected one," weighted by how wrong the implied reward currently is 02:16.
That simplicity produced an ecosystem of variants — Tatsu's word for the literature is that everyone wanted their own *PO. Two survive contact with practice: SimPO, which normalizes the update by response length and throws away the reference model entirely (giving up the ratio-of-policies derivation in exchange for a cleaner objective), and plain length-normalized DPO. Both were swept extensively in Tülu 3.
Then comes the methodological warning that colours the rest of the lecture. Ai2 published a careful comparison finding PPO beat DPO, plausibly because it is on-policy. In a later paper with a stronger SFT stage, the gap vanished — both methods gained nothing, and only length-normalized DPO helped at all. Neither result is wrong; they are answers to different questions about different base models and preference sets. Treat single RL experiment results as evidence about a setting, not about an algorithm 05:35.
The substantive reason to leave RLHF is overoptimization. Plot proxy reward (how well your RL run maximized the fitted reward model) on x and true win rate on y: the curve rises, turns over, and comes back down. It is a train/test gap wearing an RL costume — the reward model and the true preference agree in expectation but not in finite sample, and RL is very good at finding where they disagree. The lecture's sharpest version of this: the turnover shows up for human preferences and for noisy AI feedback, but not for noiseless AI feedback, which locates the cause in label noise rather than in RL itself 07:13. A second, less discussed cost: an RLHF'd model is a policy, not a probabilistic model. Calibration is not in the reward, so you should not expect it to survive — and across Anthropic's papers, the GPT-4 release and independent work, it does not.
PPO, and the tax it charges
The ladder to PPO is three rungs. Policy gradient: ∇E[R] = E[R(z)∇log p(z)] — correct, and hopelessly high-variance, and strictly on-policy, so every gradient step costs a fresh rollout. Rollouts are the expensive part of RL for LMs, so you want to take several steps per batch of samples. TRPO: sample from an older policy, correct with importance weights, and constrain the new policy to stay near the old one so the correction does not explode. PPO: replace the constraint with a clip on the likelihood ratio — beyond 1±ε there is no additional reward to collect, so the optimizer has no incentive to wander. A typical clip range is 0.2, i.e. ratios confined to [0.8, 1.2].
Conceptually that is the whole thing. The tax is everything around it. PPO needs a value network to compute advantages, usually the same size as the policy, so your GPU memory bill for the LM doubles. And the implementation surface is famous:
"You're in for a very bad time if there's a blog post that says 37 implementation details of PPO."— Tatsunori Hashimoto, 18:38
The walkthrough uses AlpacaFarm's PPO trainer, and it is worth watching for the shape rather than the details, because your GRPO code will have the same skeleton. The outer loop is unremarkable — collect rollouts, compute a loss, backward, clip grad norm, step. The subtleties live in three places. Reward shaping: RL for language models is really a contextual bandit — prompt in, response out, reward immediately, no state transitions — so the per-token KL penalty is applied per token while the actual task reward is a single scalar attached to the last token. The KL estimator gets clamped at zero when the log-ratio goes negative, which is numerically stable and no longer a KL 24:09. Gradient bookkeeping: you build a reward-weighted surrogate loss and hand it to autograd; there is an implicit stop-gradient on the weight, and you must not differentiate through it. Generalized advantage estimation: tune γ and λ to trade bias against variance — except that in the bandit setting γ = λ = 1 works fine, which collapses GAE to a plain baselined policy gradient and quietly tells you the machinery was not earning its keep.
GRPO: delete the critic, keep the baseline
So: PPO is heavy and fiddly, and the value model is the heaviest part. DPO is light but the wrong shape — math problems do not come as preference pairs, and DPO is natively offline. GRPO, from DeepSeekMath, takes the third door. Keep PPO's clipped ratio objective, delete GAE and the value network, and define the advantage as a z-score within a group:
A_i = (r_i − mean(r_1..r_G)) / std(r_1..r_G)
A "group" is the natural object here: one question, G sampled answers. The mean reward over your siblings is a genuinely good baseline because it absorbs question difficulty — the thing you want subtracted out. Baselining across different questions would be pointless; the interaction between questions is only that they share a gradient step 31:50. This is close to REINFORCE with a leave-one-out baseline, which had been studied independently.
Two details reward attention. The KL term in the GRPO paper is not the naive average log-ratio; it is the π_ref/π_θ − log(π_ref/π_θ) − 1 form, whose extra terms cancel in expectation. It is a control variate — a lower-variance sample estimator of the same KL, and a reusable trick any time you need KL from samples 32:23. And in the purely online case — one gradient step per batch of rollouts — the clipping never binds and GRPO degenerates to a policy gradient with group-normalized rewards. That is why nano-GRPO implementations fit on a slide: compute rewards, normalize per group (with a 1e-4 epsilon in the denominator so a group of all-correct answers does not divide by zero), compute the KL term, step.
The z-score is not a legal baseline
Here the lecture does something better than presenting the algorithm: it audits it. The policy gradient theorem lets you subtract any baseline that does not depend on the sampled trajectory — that is what makes the group mean legitimate. Nothing in that theorem lets you divide by the group standard deviation, and nothing licenses dividing by response length either. Both are in vanilla GRPO. The Dr. GRPO analysis removes both and reports equal or better reward with dramatically shorter outputs.
What does each term actually do? Dividing by the group std up-weights the extremes: std is small exactly when a question is trivially easy (all rewards 1) or hopeless (all 0), so those groups get amplified gradients. That is backwards relative to the folk theorem that you want to train on problems at the edge of the model's competence — a curriculum effect pointed the wrong way 42:15.
The length normalization is nastier, because it is an incentive bug rather than an efficiency bug. Divide the advantage by response length, and a wrong answer — negative advantage — has its penalty diluted by being long, while a right answer maximizes its positive advantage by being short:
"What this does is it actually produces a model that kind of BSes as aggressively as possible. If the model thinks it can't get the answer right, it just produces the longest possible response."— Tatsunori Hashimoto, 40:36
Fix it and reward is unchanged while output length stabilizes instead of climbing. This matters far beyond a loss-function footnote, because R1's headline observation was that chain-of-thought length rises steadily through RL training, read in the paper as the model learning to think harder on harder problems. Dr. GRPO's counter-reading is that a biased objective is paying for length directly. The same paper deflates the "aha moment": run DeepSeek-V3-Base on math problems and it already says things like "aha, I can do this" — RL amplified a behaviour that was already in the base model rather than inventing one. Tatsu finds both arguments credible, while being clear that R1 is nonetheless an excellent math model 48:47. Hold both thoughts: the recipe works, the folklore around it is mostly unearned.
Three recipes, read side by side
DeepSeek-R1 comes in two objects. R1-Zero is the controlled experiment: take DeepSeek-V3 base, no instruction tuning, and run GRPO on math-style tasks with two rewards — binary accuracy, plus a format reward for putting reasoning inside thinking tags. That gets close to o1. The format reward looks cosmetic and reportedly is not; it is the scaffolding that makes long CoT usable at all. R1 is the shipped model, and it adds exactly what you would expect: an SFT initialization on long CoTs (the provenance of that data is left vague in the paper), a language-consistency reward because unconstrained RL makes models code-switch mid-CoT, and a standard SFT + RLHF stage afterwards — 600K reasoning items judged by V3 on non-verifiable tasks like "write a proof of X", plus 200K non-reasoning items from V3's SFT set, with GRPO reused as the RLHF optimizer too.
Two findings from R1 outlive the model. First, distillation: generate ~800K CoT traces from R1, fine-tune Qwen 2.5, and small models jump enormously on competition math. And the SFT-priming effect is even cheaper than that — the s1 work fine-tuned Qwen 2.5 on 1,000 long CoTs from Gemini 2.0 Flash Thinking and got strong math accuracy, which argues the base model already has the capability and you are eliciting rather than installing it 52:03. Second, the negative results, which the lecture rates as scientifically the most valuable part of the report: process reward models and MCTS were both tried at scale and neither helped. Note the reversal — DeepSeekMath's own plots had process supervision as the best line; by R1 the team had concluded outcome rewards win. Rich intermediate feedback is genuinely more informative; it is just too hard to obtain reliably.
Kimi k1.5 landed simultaneously, matched o1, and is the better paper on the parts R1 waves at. Its data curation is explicit: balance across math domains, drop multiple-choice and true/false because they are guessable, and keep only problems the SFT model fails at best-of-8 — difficulty filtering as a first-class step. Its optimizer arrives from the DPO direction: same nonparametric assumption, solve for the implied reward, but with no pairwise preferences to plug into Bradley–Terry, they impose the resulting equality as a squared loss. Take the gradient and you get something you already recognize — a baselined policy gradient (batch mean, no std division) with an explicit regularizer where GRPO clips. Convergent evolution: baseline plus regularization plus policy gradient is the load-bearing structure, and the specific parameterization is taste.
Kimi's most forward-looking choice is treating CoT length as a cost rather than a trophy. Their length reward computes λ from where a rollout sits in its group's length range: correct answers are pushed toward the shortest end, incorrect ones toward the middle. Turn it on from step zero and RL stalls in a local optimum where the model, correct at nothing, minimizes length; so they enable it only later in training 67:18. Add an easy-to-hard curriculum with sampling proportional to (1 − success rate), generated test cases for code, and — surprisingly for a "verifiable rewards" pipeline — an 800K-sample learned CoT reward model for math answer equivalence, which is really advanced string matching.
Kimi is also the only one of the three to discuss infrastructure, and the constraints transfer directly to your assignment. RL is harder to keep GPUs busy on than pre-training because on-policy training means inference in the loop; because you are shuttling between a training framework and an inference server (weights out, trajectories back); and because long CoTs make batches ragged. The current state of the art is unglamorous: run vLLM with dummy weights, push real weights in through partly undocumented paths, and tear the server down each iteration to reclaim memory. Tatsu notes the NCCL-collective weight-sync API was considered for the assignment and rejected as too immature 72:17.
Qwen 3 is the same playbook with two new results. The efficiency one: reasoning RL on 3,995 examples, after best-of-N difficulty filtering, validation decontamination, and manual filtering of CoTs that got the right answer by guessing. Small-data RL works — though as the lecture is careful to say, that it works small is not evidence it does not also scale. The mechanism one is thinking-mode fusion: fine-tune the post-RL model on data tagged think and no-think so one parameter set serves both modes, then exploit that at inference by interrupting a running CoT with a fixed string ("considering the limited time by the user, I have to give the solution by thinking directly now") and a close-think tag. That converts the thinking budget into a dial and yields clean test-time scaling curves from a single model. The ablation across stages carries the sting: general RLHF improves instruction following and general tasks but costs math/STEM in thinking mode. The alignment tax did not disappear; it moved.
Numbers worth carrying
| PPO clip range | 0.2 → likelihood ratio confined to [0.8, 1.2] |
| GAE in the LM bandit setting | γ = λ = 1 works; collapses to a baselined policy gradient |
| GRPO advantage epsilon | 1e-4 in the std denominator (you will need this in A5) |
| R1 distillation corpus | ~800K CoT traces from R1 → Qwen 2.5 |
| R1 post-RL SFT | 600K reasoning (V3-as-judge) + 200K non-reasoning |
| s1 long-CoT priming | 1,000 examples |
| Kimi difficulty filter | keep only problems that fail best-of-8 |
| Kimi math verifier | CoT reward model trained on 800K samples |
| Kimi length reward λ | [−0.5, 0.5]; enabled only later in training |
| Qwen 3 reasoning RL set | 3,995 examples |
| Assignment 5 budget | 2 H100s per run (1 vLLM + 1 policy); 4 h wall-clock for the leaderboard |
What you build with this
This lecture is the direct spec for Assignment 5: Alignment (2025 handout PDF), the last assignment of the course. You take Qwen 2.5 Math 1.5B Base, evaluate it zero-shot on MATH with the R1-Zero prompt and a regex-plus-format reward function, then climb the same ladder the lecture climbs: SFT on correct traces, expert iteration (SFT on your own correct rollouts — the "RFT" baseline from the DeepSeekMath plot), and finally GRPO.
The GRPO half maps one-to-one onto the second half of the lecture. compute_group_normalized_rewards is equation 3 with the 1e-4 epsilon. compute_naive_policy_gradient_loss and compute_grpo_clip_loss are the two branches — the naive one is the pure-online case where clipping is inert; the clipped one is only meaningful once you go off-policy with multiple epochs per rollout batch (cache the old log-probs, and do not differentiate through them). The handout notes its version is a special case of DeepSeekMath's GRPO with a verified reward, no KL term and no reference-model updates — the KL term made no difference to performance while costing a resident reference model. Then the experiments are the lecture's audit, run by hand: grpo_baselines, grpo_length_normalization and grpo_group_standard_deviation ask you to ablate exactly the two terms Dr. GRPO removes, and think_about_length_normalization asks you to reason it out before you measure. The infrastructure Kimi describes appears in miniature: one GPU serving vLLM, one holding the policy, and weights pushed across between steps. Scoring ends at a leaderboard — best MATH validation accuracy within 4 hours on 2 H100s, temperature 1.0, max 1024 tokens, all 5K validation examples. An optional supplement covers safety alignment, instruction tuning and RLHF for anyone who wants the lecture-15 material as code too.
Supporting materials, verified
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Shao et al. (2024) · where GRPO is introduced; equations 2 and 3 on the slides are from here, along with the outcome-vs-process supervision plot the lecture revisits.
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (2025) · case study 1: R1-Zero, the SFT-initialized R1 pipeline, distillation into Qwen, and the unsuccessful-attempts section on PRMs and MCTS.
- Understanding R1-Zero-Like Training: A Critical Perspective (Dr. GRPO) — Liu et al. (2025) · the audit: the std division and length normalization are not valid baselines; growing CoT length and the "aha moment" are largely artifacts.
- Kimi k1.5: Scaling Reinforcement Learning with LLMs — Kimi Team (2025) · case study 2: difficulty filtering, the DPO-derived squared-loss objective, the length reward, the curriculum, and the only serious RL-infrastructure discussion of the three.
- Qwen3 Technical Report — Qwen Team (2025) · case study 3: 3,995-example reasoning RL, thinking-mode fusion, the thinking budget, and the stage-by-stage ablation showing general RLHF costing math performance.
- Proximal Policy Optimization Algorithms — Schulman et al. (2017) · the clipped surrogate objective GRPO inherits.
- Trust Region Policy Optimization — Schulman et al. (2015) · rung two of the ladder: importance weighting plus a trust region, before clipping replaced the constraint.
- High-Dimensional Continuous Control Using Generalized Advantage Estimation — Schulman et al. (2015) · the GAE box in the PPO diagram; the γ, λ knobs that turn out not to matter in the bandit setting.
- OpenAI Spinning Up — PPO — OpenAI · the clean statement of PPO shown on the slide, if you want the conceptual version before the implementation reality.
- The 37 Implementation Details of Proximal Policy Optimization — Huang et al. (2022) · the blog post named on the slide; read it to understand why people wanted an alternative to PPO.
- Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO — Engstrom et al. (2020) · the paper showing that PPO's implementation choices, not its objective, carry much of its measured benefit.
- AlpacaFarm (code) and AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback — Dubois et al. (2023) · the PPO trainer walked through on screen; the RLHF-vs-DPO controlled comparison on the slides is run here.
- Secrets of RLHF in Large Language Models Part I: PPO — Zheng et al. (2023) · source of the PPO-for-language-models diagram (policy, reward model, value model, GAE) on the slide.
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafailov et al. (2023) · the derivation recapped in the first ten minutes, and the one Kimi reuses without the Bradley–Terry step.
- Learning to Summarize from Human Feedback — Stiennon et al. (2020) · the pairwise reward-modeling objective DPO substitutes the implied reward into.
- SimPO: Simple Preference Optimization with a Reference-Free Reward — Meng et al. (2024) · length normalization plus dropping the reference model; one of the two DPO variants the lecture keeps.
- Tülu 3: Pushing Frontiers in Open Language Model Post-Training — Lambert et al. (2024) · where SimPO and length-normalized DPO were swept, and the source of the reversed PPO-vs-DPO conclusion.
- Scaling Laws for Reward Model Overoptimization — Gao et al. (2022) · the canonical quantitative account of the turnover curve that motivates the whole lecture.
- Let's Verify Step by Step (PRM800K) — Lightman et al. (2023) · the process-reward-model line of work that R1 tried and abandoned; read it to see what was being given up.
- s1: Simple Test-Time Scaling — Muennighoff et al. (2025) · the 1,000-example long-CoT SFT result mentioned as joint work with Percy Liang's group; also the budget-forcing precursor to Qwen 3's thinking budget.
- nano-aha-moment (code) — McGill NLP · the compact GRPO implementation shown on the slides; the clearest way to see that the advantage computation really is four lines.
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs — Ahmadian et al. (2024) · field map extra: the leave-one-out baseline GRPO nearly is, argued for on its own terms before GRPO made it famous.
- Reinforcement Learning: An Introduction (2nd ed.) — Sutton & Barto · field map extra: the baselining result cited on the slide when the lecture checks whether GRPO's advantage is legal. Chapter 13 is the relevant one.
- verl — Volcano Engine Reinforcement Learning for LLMs — ByteDance · field map extra: a modern production RL-for-LLM stack, the kind of implementation Tatsu gestures at as better than the AlpacaFarm code being walked through.
- Assignment 5 supplement: safety alignment, instruction tuning and RLHF (PDF) — CS336 staff · field map extra: the optional companion assignment; the 2025-named file has been replaced upstream by this 2026 revision.
Exercises
- Group-normalized advantages, from the equation code — In the assignment repo, implement compute_group_normalized_rewards and compute_naive_policy_gradient_loss, then grpo_microbatch_train_step. Steps: (1) uv sync --no-install-package flash-attn && uv sync; (2) wire the three functions into tests/adapters.py; (3) uv run pytest tests/test_grpo.py until green; (4) write a five-line script that feeds one group of eight fake rewards (seven zeros, one one) through your advantage function and prints the result. A good answer notices that the epsilon is load-bearing for the all-zero and all-one groups, and that the advantage is constant across tokens within a response so broadcasting must be deliberate, not accidental.
- Reproduce the Dr. GRPO audit yourself code — Run three short GRPO trainings on MATH from the assignment (the grpo_length_normalization and grpo_group_standard_deviation problems): vanilla, no-std-division, and no-length-normalization. Steps: (1) parameterize the two terms with flags rather than forking the loss; (2) log validation reward and mean response length every 5–10 steps on ≥1024 validation examples; (3) stop a run early if it clearly diverges; (4) plot reward-vs-step and length-vs-step for all three on shared axes. A good answer reports that removing length normalization flattens the length curve at equal or better reward, and says explicitly which of the two terms your data can and cannot distinguish given the noise.
- Prove the illegality — Write the policy gradient theorem with a baseline b(s), show why E[b(s)∇log π] = 0, and then show concretely why the same argument fails once you divide by the group standard deviation (hint: std depends on r_i, the very sample you are weighting). Then write down the leave-one-out variant — mean and std computed over the other G−1 rollouts — and state whether it restores unbiasedness, and at what cost in variance. A good answer names exactly which independence assumption breaks and what the resulting gradient is biased towards, connecting back to the easy/hard up-weighting.
- Port Kimi's length reward onto GRPO code — Add Kimi k1.5's length reward to your GRPO loop: for each group compute λ from the rollout's position in the group's [min, max] length range, add λ to the reward for correct answers and min(0, λ) for incorrect ones, and gate the whole term behind a step threshold. Steps: (1) implement it as a reward-shaping function, not a loss change; (2) run with the term on from step 0 and confirm you can reproduce the stall Kimi warns about; (3) re-run enabling it after N steps; (4) compare final accuracy and mean CoT length against unmodified GRPO. A good answer explains why the early-enable run collapses (a model correct at nothing can only optimize length) and quantifies the accuracy paid per token saved.