Alignment: RL 2
Transcript: cleaned auto-captions with timestamps
This is the last lecture Percy or Tatsu give in the course, and it is deliberately not new material. Lecture 15 built the alignment stack up to RLHF and DPO; lecture 16 introduced RL from verifiable rewards and named PPO and GRPO. Here Percy re-derives the same objects slowly and then runs them, so that the assignment-5 GRPO implementation stops being a formula copied out of a paper and becomes a loop you can reason about. It closes the course's argument the way the course opened it — by refusing to let an abstraction stay abstract until you have watched the tensors go through it.
Outline, with timestamps
- 0:05 — Where this lecture sits: the last Percy/Tatsu lecture, and a deep dive rather than new material.
- 1:13 — RL, restated for language models: state, action, outcome and verifiable rewards, and transitions that are just concatenation.
- 5:30 — Policy gradient: the log-derivative trick, and why naive PG is SFT weighted by reward.
- 9:27 — Sparse rewards: a policy that never succeeds takes a zero gradient and stays stuck.
- 15:32 — Baselines: subtract any b(s), keep the optimum, and watch the variance fall from 5.323 to 1.155.
- 24:41 — The optimal baseline, the E[R|s] heuristic, and how it becomes the advantage function.
- 30:48 — Why GRPO is a language-model-shaped algorithm: the group is the free baseline PPO needs a critic for.
- 33:37 — The sorting task, and two reward functions — one too sparse, one with an exploitable loophole.
- 37:29 — A deliberately tiny non-autoregressive model, and the tensor shapes of a rollout.
- 43:14 — Deltas: raw, centered, normalized, max-only — four ways to turn rewards into updates.
- 48:12 — Log probs, the naive loss, and the freezing interlude that kills a whole class of silent bug.
- 53:17 — The clipped GRPO loss and the KL penalty, with its low-variance estimator.
- 59:31 — The full loop: three models, two nested loops, and which of them you actually have to store.
- 67:33 — The runs: raw vs centered vs normalized, why the loss curve lies, and the closing argument for RL.
RL where the dynamics are a +
Percy starts by re-stating the MDP for language models, and the interesting part is how much of standard RL falls away (1:13). The state is the prompt plus whatever the model has generated so far; an action is one token; the reward is a function of the finished response, computed by deterministic code rather than asked of a human. Because the reward arrives once, at the end, discounting and bootstrapping mostly stop being relevant — there is nothing in the middle to bootstrap from.
Two consequences are worth holding onto. First, the transition function is s' = s + a: you have a perfect, free, exact model of the environment. A roboticist would give a great deal for that; it is what makes planning, search and test-time compute available to language models at all. Second, the state space is invented rather than physical.
"The notion of state is really made up."— Percy Liang, 3:21
In control, reachability is the hard part — there are joint configurations you simply cannot get to. Here every state is reachable by definition: the model writes whatever scratchpad it wants. The difficulty moves entirely into whether those freely-chosen tokens land on a correct answer. Same theory, inverted intuitions — which is exactly why importing RL folklore wholesale into LM training goes wrong.
Policy gradient is SFT with a scalar in front
The derivation is three lines (5:30). Write the objective as E[R] = ∫ p(s) π(a|s) R(s,a), take the gradient, and note that only π depends on the parameters. Then apply the identity ∇π = π · ∇ log π to put a π back in front so the whole thing is an expectation again:
∇ E[R] = E[ ∇ log π(a|s) · R(s,a) ]
Sample a prompt, sample a response, step. What you are stepping on is the supervised fine-tuning gradient — maximize the log-probability of this response — scaled by the reward. That reframing does more work than it looks like. If rewards are binary, naive policy gradient is rejection-sampled SFT: correct rollouts get a unit-weight SFT update, incorrect rollouts get multiplied by zero and vanish. The only structural difference from ordinary supervised learning is that the dataset is regenerated from the current policy every iteration, so it improves (or collapses) as training proceeds.
That single difference is where all the pain comes from. With sparse rewards — hard math problems, a weak initial policy — almost every rollout scores zero, and a batch of zeros produces a gradient that is exactly zero (9:27). Supervised learning always moves somewhere, even if somewhere bad; RL can sit perfectly still. The bootstrap out of that hole is generalization: get the easy problems right, update, and hope the improved policy starts landing occasional wins on the harder ones. A student asks the obvious follow-up at 12:45 — why not set failures to −1 so they push the model away? Percy defers it, because the answer is not a hand-set constant. It is the baseline.
Baselines: the entire game is variance
Optimize E[R − b(s)] instead of E[R], for any b that depends on the state but not the action (15:32). Because ∫ π(a|s) da = 1, the term E[b(s)] contains no π at all — it is a constant offset, the argmax is untouched, and every estimator built this way stays unbiased. What changes is the spread of the per-sample gradient, and RL lives on per-sample.
The toy example makes the point in four numbers. Two prompts: from s1, action a1 pays 11 and a2 pays 9; from s2, a1 pays 0 and a2 pays 2. The correct policy is a1 in the easy state and a2 in the hard one. But naive policy gradient compares raw magnitudes across states: the wrong action in the easy state carries a weight of 9, while the right action in the hard state carries 2. Locally you push hardest on the mistake. In expectation it still converges — but "in expectation" is doing a lot of work when a run has a fixed compute budget and can lock into a local optimum first (19:02).
| gradient weights, no baseline | {11, 9, 0, 2} | std = 5.323 |
| with b(s1)=10, b(s2)=1 | {1, −1, −1, 1} | std = 1.155 |
Same optimum, roughly a 4.6× tighter estimator, for the cost of one subtraction per state. Statisticians will recognise this as a control variate; Percy makes the connection explicitly at 30:48.
There is a variance-optimal baseline — for a one-parameter model, b*(s) = E[(∇π)² R | s] / E[(∇π)² | s] — but in high dimensions it turns into covariance terms nobody wants to compute (24:41). Drop the gradient weights and you are left with the heuristic everything in practice uses: b(s) = E[R | s], the mean reward you expect from this prompt. Still an expectation, still only estimable — but now it is a target with an obvious estimator.
One subtlety the students push on and Percy answers cleanly (22:27): "doesn't depend on the policy" is a statement about differentiation, not about provenance. You may absolutely build b out of a frozen copy of a previous policy. What you may not do is let gradients flow through it. Freezing is what makes the theory true in code — a theme the lecture returns to with teeth in a few minutes.
From baseline to advantage to the δ slot
Define V(s) = E[R | s] and Q(s,a) = E[R | s,a]; in this outcome-reward setting with a denoting the whole response, Q and R coincide. The advantage is A(s,a) = Q(s,a) − V(s) — how much better this response was than what this prompt usually yields. Substitute the heuristic baseline and the baselined reward is the advantage (26:21). Two derivations, one object.
Percy then does the useful abstraction. Every method in this family has the shape ∇ log π(a|s) · δ, where δ is some reward-derived scalar: the raw reward, the reward minus an estimated mean, that difference divided by a standard deviation, or a learned critic's advantage. PPO fills the slot with a value network. GRPO fills it with the group mean. The slot is the invariant; the filling is fashion.
"The exact form of GRPO isn't that important, because next year there's going to be GRPO2 or whatever — but all this policy gradient stuff, I'll be giving the same lecture next year."— Percy Liang, 29:39 (captions read "GPR"; resolved from the lecture script)
Which explains why GRPO is a 2024 algorithm and PPO is a 2017 one (30:48). Language modeling hands you a group for free: sample G responses to the same prompt and their mean reward estimates V(s) directly — no critic to train, no second network to keep in memory, no second set of hyperparameters to get wrong. A walking robot has no equivalent; no two rollouts share a state that cleanly, so you are forced back to a value function that aggregates over everything. GRPO is less an improvement on PPO than an exploitation of structure PPO could not assume.
The flip side is worth internalising before you debug a stalled run: if all G responses to a prompt earn the same reward, the centered δ is exactly zero and that prompt contributes nothing. That is correct behaviour — there is no relative preference to express — but on a task where the policy has saturated or bottomed out uniformly, whole batches go dead.
The executable lecture: sorting three numbers, badly
The second half is lecture_17.py, about 560 lines you can run on a laptop. The task is sorting: the prompt is n numbers, the response is n numbers that should come out sorted (33:37). Two candidate rewards, and the comparison between them is the most transferable part of the lecture.
| prompt [3,1,0,2] → response | [0,1,2,3] | [7,2,2,5] | [0,3,1,2] |
| positions matching sorted(prompt) | 4 | 1 | 1 |
| inclusion + adjacent-ordering | 7 | 3 | 6 |
The first reward — count positions equal to the ground truth — is what you actually care about, and it is useless as a training signal: a garbage response and a nearly-correct one both score 1, purely by luck of one aligned digit. The second gives a point for each prompt token that appears anywhere in the response plus a point for each adjacent non-decreasing pair, and it ranks those same three responses the way a person would (7, 3, 6). That is the one the runs use.
Percy flags, and leaves as an exercise, that the second reward has a loophole. It does: the ordering term tests x ≤ y, so any constant response collects every adjacency point without sorting anything, and the two terms are gameable independently. Partial credit bought a denser gradient and sold a shortcut — and in the training logs the model takes it.
The model is deliberately not a transformer (37:29): fixed prompt and response length, a per-position encode matrix and a per-position decode matrix, input and output embeddings tied, and each response position decoded independently — explicitly not autoregressive, because autoregressive generation is the part that makes toy RL code hairy. Three einsums and you have logits. The rollout shapes are the thing to actually memorise, because they are the same in a real system: prompts [batch pos] → logits [batch pos vocab] → torch.multinomial for G samples → responses [batch trial pos] → rewards [batch trial] → deltas [batch trial], broadcast back over positions at loss time because the reward is outcome-level and has no position index. Swap in process rewards and that last broadcast becomes a real per-position tensor. The generation call is where production code hands off to vLLM.
Three details that bite
The δ menu. One function, four modes (43:14): rewards passes them through; centered_rewards subtracts the per-prompt mean over the group; normalized_rewards also divides by the per-prompt standard deviation plus 1e-5 — that is GRPO; and max_rewards, which is Percy's own and not in the paper, zeroes everything below the group max to stop the policy settling for partial credit. Centering is the answer to the student's earlier question about negative rewards: with one success and nine failures the mean is 0.1, so the nine failures get −0.1 and finally push down. Normalizing buys scale-invariance — multiply every reward by 100 and nothing changes.
The freeze. The ratio π(a|s) / π_old(a|s) is only meaningful if the denominator is a constant. Build both from the same live parameter and the ratio is identically 1, whose gradient is 0 — the toy is literally w = 2., p = σ(w), p_old = σ(w), ratio.backward(), w.grad == 0 (50:28). Wrap p_old in torch.no_grad() and the gradient comes back. Percy notes on camera that he did not wrap it in his own compute_loss. Take the class of bug seriously: it raises nothing, produces no NaN, and simply makes your updates quietly wrong.
The clip and the KL. The GRPO loss computes r = exp(logπ − logπ_old), forms both r·δ and clamp(r, 1−ε, 1+ε)·δ, takes the elementwise minimum and negates (53:17). Inside the trust region it reduces to the naive update rescaled by a constant; outside, the magnitude is bounded. The script uses ε = 0.01, considerably tighter than PPO's customary 0.2. The optional KL penalty uses the low-variance estimator E_p[q/p − log(q/p) − 1] — unbiased, because E_p[q/p] = 1 makes the two extra terms cancel, and non-negative sample-by-sample, which the plain −log(q/p) is not. Note the two regularizers point at different anchors: the clip restrains you relative to π_old (the checkpoint that generated this batch), the KL relative to π_ref (a slower frozen model, refreshed every ten epochs here). Asked why not KL against π_old, Percy is honest that they are two knobs on one continuum, with the reference model's job being to hold the objective still long enough to be worth optimizing (65:17).
Which leaves the accounting: three models in play — the live policy, π_old, and π_ref. Only two of them cost memory. Because the inner loop reuses one batch of generated responses, π_old can be cached as a tensor of log-probs rather than a copy of the weights; π_ref cannot, and doubles your parameter footprint (64:07). The outer loop generates, the inner loop takes several gradient steps on that generation, because inference is the expensive half.
What the runs showed, and why the loss curve lies
The experiments are 100 epochs of 10 inner steps, 10 responses per prompt, three fixed prompts ([1,0,2], [3,2,4], [1,2,3]), Adam at 1e-3, seed 5 — about twenty seconds of laptop time (67:33). With raw rewards, mean reward climbs and sorting does not really happen: the model learns to emit prompt-ish tokens in vaguely non-decreasing order, which is the loophole, collected. With centered rewards it is better — failures get pushed down, tied groups abstain instead of reinforcing mediocrity — and still ends in local optima, with batches where every δ is zero and only fresh sampling can restart progress. Full normalization changes little here, and the lecture script points at Dr. GRPO's finding that dividing by the standard deviation introduces a length bias worth avoiding — inapplicable in this toy, where every response is the same length.
The most useful five minutes are the last ones. The loss curve looks bad while the reward curve goes up, and that is not a bug to chase:
"Minimizing the loss is a little bit of a lie here … it's not like we were minimizing one loss function that you can measure over time, because the set of responses is changing over time."— Percy Liang, 72:00
The loss is computed against a dataset the policy regenerates each epoch, and against its own samples, so it has no fixed reference point — it is circular by construction. Reward, on a held-out set if you have one, is the only meter. Anyone who has stared at an RL training dashboard trying to read tea leaves in the loss should stop.
Percy's close is a two-part argument. The optimistic half: RL is how models get past imitation, because supervised data can only teach you to reproduce what is in it — if you can measure it, you can optimize it. The pessimistic half is bigger than the optimization problem: designing rewards that are not hackable, in environments that generalize, is unsolved, and the sorting demo is a two-minute proof of the failure mode. And a systems observation the course ran out of quarters to teach — RL infrastructure is substantially harder than pretraining infrastructure, because you are running inference workloads inside your training loop, shipping weights to inference workers, spinning up environments, and holding several models simultaneously, all distributed.
What you build with this
This lecture is the mechanics behind the GRPO half of Assignment 5: Alignment (handout), which starts with lecture 16. A5 walks from SFT through expert iteration to GRPO on math reasoning with a verifiable answer-checking reward. Map the pieces directly: the assignment's reward function is sort_inclusion_ordering_reward's real-world cousin — parse the model's final answer, compare against the key, return a scalar; its advantage computation is compute_deltas with the group mean and standard deviation; its loss is compute_loss in clipped mode; and the outer-generate / inner-step structure is the same two loops. If you get an assignment run that trains but does not improve, the ordered suspects from this lecture are: rewards all tied within each group (dead δ), gradients flowing through π_old (dead update), and a reward with an unintended shortcut (training fine, learning the wrong thing). Watch reward, not loss.
Supporting materials, verified
- Executable lecture 17 (trace) — Percy Liang (2025) · the lecture itself; source at lecture_17.py, which is where the exact reward functions, compute_deltas modes, ε = 0.01 and the KL estimator live.
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Shao et al., DeepSeek (2024) · introduces GRPO; the whole second half of the lecture is a transcription of its algorithm box into runnable code.
- Proximal Policy Optimization Algorithms — Schulman et al. (2017) · the source of the clipped surrogate objective and the ratio; GRPO is this minus the critic.
- Understanding R1-Zero-Like Training: A Critical Perspective — Liu et al. (2025) · cited in the lecture script as Dr. GRPO: dropping the standard-deviation normalization removes an optimization bias that inflates the length of incorrect responses.
- CS224R: Deep Reinforcement Learning — Chelsea Finn, Stanford · Percy points here for the policy-gradient derivations this lecture compresses; the lecture script's direct link to the 2025 policy-gradient notes now 404s, so start from the course page.
- Qwen3 Technical Report — Qwen Team, Alibaba (2025) · the script's pointer for the following guest lecture (Junyang Lin), which is not in the public playlist.
- The Llama 3 Herd of Models — Grattafiori et al., Meta (2024) · the script's pointer for the final guest lecture (Mike Lewis), likewise not public.
- Approximating KL Divergence — John Schulman (2020) · field map extra the derivation of the k3 estimator q/p − log(q/p) − 1 that the lecture writes down without naming; explains why it is unbiased, non-negative and lower-variance.
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (2025) · field map extra what this machinery looks like at scale with a purely verifiable reward, and the reference point for the "R1-Zero-like training" the Dr. GRPO paper critiques.
- Reinforcement Learning: An Introduction (2nd ed.) — Sutton & Barto (2018) · field map extra chapter 13 is the canonical treatment of policy gradient, baselines and the variance argument this lecture reproduces in LM clothing.
Exercises
- Break the reward, then fix it code — Percy leaves the loophole in sort_inclusion_ordering_reward as an exercise. (1) By hand, find a response for prompt [1,0,2] that scores at least 3 without being sorted. (2) Brute-force all 5³ responses over the training vocabulary for each of the three training prompts and tabulate reward against "is actually the correct sort". (3) Report how many distinct responses tie or beat the correct answer. (4) Propose a repaired reward — e.g. require strict < in the ordering term, or make inclusion a multiset match — and re-run the brute force. A good answer shows the correct sort is now the unique argmax, and says what that cost you in gradient density near the initial policy.
- Sweep the δ menu code — Run run_policy_gradient for all four deltas_mode values at three seeds each, logging mean reward per epoch and the fraction of prompt-groups whose δ is all-zero. (1) Plot mean reward with a seed band. (2) Plot the dead-group fraction over training. (3) Explain why max_rewards behaves differently from centered_rewards on this task specifically. A good answer notes that the two curves answer different questions — reward measures progress, dead-group fraction measures whether progress is still possible — and that a single seed here is not evidence of anything.
- Instrument the clip code — Switch loss_mode to clipped and record, per inner step, the fraction of ratios that fall outside [1−ε, 1+ε], for ε ∈ {0.01, 0.2} and num_steps_per_epoch ∈ {1, 10, 50}. Separately, wrap old_log_probs in torch.no_grad() and compare gradient norms with and without. A good answer shows clipping is near-inactive at step 1 and bites as the inner loop runs on, and quantifies how much of the update the missing freeze was silently changing.
- Prove the two identities — On paper: (a) show that E[∇ log π(a|s) · b(s)] = 0 for any b not depending on a, and therefore that any baseline leaves the gradient unbiased; (b) show that E_p[q/p − log(q/p) − 1] = KL(p ‖ q), and argue informally why it has lower variance than the single-sample −log(q/p). A good answer for (a) is three lines ending in ∇ ∫ π = ∇ 1 = 0, and for (b) notes the estimator is non-negative pointwise while the naive one is not.