CS336 // FIELD MAP
← deep dive
TRANSCRIPT · LECTURE 16Tatsunori Hashimoto · 80 min

Alignment: RL 1 — reinforcement learning from verifiable rewards

Cleaned auto-captions · timestamps open the video at that moment · caption errors corrected where the slides settle them, otherwise left as heard

DPO recap and the ×PO zoo

00:05 Okay. We'll get started. This is the second of the post training lectures and we're going to talk about well we're going to finish off the lecture from Tuesday first. Whatever remaining content we have. And then we'll talk about reinforcement learning from verifiable rewards kind of the new exciting thing that's been happening over maybe the last six months or so. And really powering all of these reasoning models that you kind of see dominating the landscape. So, before I get into the verifiable rewards part, the exciting, math RL content, I'm going to go back and I'm going to finish off RLHF, right? So, this is, a very quick recap of the last few slides from Tuesday's lecture, right? So, we want to do reinforcement learning from human feedback, which is the setting where we observe pairwise preference data, like which of the two responses are better, and we'd like to get a language model policy that sort of maximizes some sort of underlying reward. For this preference data, right? And if you remember, DPO is this algorithm that allows to allows us to optimize this reinforcement learning objective. And this objective is hard because our language model, the policy, so to speak, is on the bottom of this expectation, right? You're sampling from your language model. You're not just maximizing the likelihood under the language model. And the way we're going to do this, if you remember the derivation is that we're going to, make the nonparametric assumption that our policy class is the set of all functions. We're going to rewrite the reward as a ratio of the policies. And then plug that in to the Bradley-Terry objective. And

01:42 So what we're going to do is we're going to find a policy such that sort of the implied reward maximizes the probability of seeing sort of the references sorry, the pairwise preferences that we see, right? So now this is great because it's basically supervised learning on a kind of alternatively parameterized objective. Which is nice. DPO updates have this following form. The reason why DPO has kind of took the world by storm for a bit a year ago or so is that it can be written in this very nice form where what you're going to do is you're going to take these gradient steps which are going to be multiplied by beta. That's kind of your regularization. You're going to have higher weight when your reward estimates are wrong. So you're going to update more when your sort of implied rewards are not correct. And then it you're going to increase the likelihood of the good example and you're going to decrease the likelihood of the bad example. Right? As I was kind of saying, if you sort of remember what I was saying last lecture, really RL algorithms kind of boil down often to just upweight the good stuff, downweight the bad stuff. And the subtlety really is in deciding, h what the good stuff is and how much to upweight, right? And so this is one particular choice of doing that. And we'll see, other things happen. And so DPO works pretty well. For a while essentially all the open model releases use some variant of a DPO to do their post training. And really it's just very easy to get working compared to PPO. And so for a while it was really the dominant approach. In the time

03:23 Since then there was a huge flood of papers that was like asterisk-PO. Basically everyone wanted to come up with a DPO variant. There was dozens and dozens of papers. There's so many that I don't think it's it's worth it to go over the big ocean of such papers. I'll mention two, not because necessarily that I think they're like the right ones, but because I think they're ones that have been used recently, by people that are sort of really pushing the limits of this kind of DPO style post training. One is SimPO. And SimPO just makes a very simple modification or two simple modifications. The first one is to really normalize the update size by the length of the responses. We'll see kind of this theme appearing later. And the other thing they do is they just get rid of the reference. So now we've lost the DPO sort of mathematical argument that what we're doing is looking at ratios of policies but well this is more purely looking at something looks like something where we upweight the good stuff and downweight the bad stuff under that sort of underlying motivation SimPO totally fine there's other variants where you don't do necessarily this like reference policy removal where you just sort of normalize by length called length normalized dpo that kind of people have done. And this, these two forms, DPO with length normalization and SimPO were the things that were tried, pretty extensively by, Ai2 folks when they did Tülu 3. And I think one thing that I'll point out, so I'm going to pause for a moment because I feel like this is an important conceptual, not conceptual, just an important empirical point, which is that in RL a lot of the findings are very contingent on the specific setting, right? So

RL findings are contingent: the Ai2 PPO-vs-DPO reversal

05:02 Depending on the environment in which you run it, depending on the base model that you have, depending on the post-training preferences that you're running on, you will find pretty different conclusions. And so one example of this, right? So, a bunch of the Ai2 folks have been doing a lot of really good post-training empirical work. And they had one work where they were sort of comparing DPO and PPO and they had found that PPO was better than DPO because of maybe it's on policiness and they show this, this jump exactly is the DPO to PPO gap. And then in a later work in Tülu 3, they find actually if you do your SFT in a sort of nicer, better way, actually that just eats up all the gains of both PPO and DPO. So neither of these really get any gains. And the only thing that does better is maybe DPO with normalization, right? Like pretty different conclusions. Of course, a lot is different between this left and right paper. But it's not that one is wrong and the other is wrong. Really that you should be careful about reading too many like generalized conclusions about any of this stuff on the basis of one paper right so that's an important note even as I go and talk about PPO and I go and talk about GRPO later in this you

Overoptimization and lost calibration

06:07 Shouldn't really take any single experimental result necessarily as gospel okay so the last thing I want to end with for RLHF is sort of two important things and one of them will really just motivate I think the entirety of today's lecture so I think that's important so the first is overoptimization. So in some sense this is just overfitting in fancier naming but it is a term that I think is very important because what it's saying is essentially as you optimize your policy more and more so think about this x-axis as how much RLI did initially your reward kind of goes up and up and up but eventually the reward that reward models that you fitted on human preferences are diverging from real human preferences and the more you optimize you eventually kind of diverge out right like you're not doing any better you're just optimiz optimizing but not actually improving your rewards, right? So this overoptimization thing appears basically everywhere in RLHF. It is a very big problem. And so that is kind of a concern. And overoptimization is a phenomenon that really happens in many ways because of the noisiness of human preferences and the complexities of human preferences. So so some of my students had a study where they basically did RLHF on human preferences. They did RLHF on noisy versions of, AI feedback and then they did RLHF on non-noisy versions of human feedback. And you see clear over optimization phenomena both for humans and for noisy AI feedback, but not really for sort of clean noiseless AI feedback. And so really, if you're going to be post-raining, you should expect to see curves that look like things on the

07:44 Left. As you train your model to do better and better on your sort of proxy rewards that you've measured, you're not necessarily going to get better and better models in terms of sort of human preference win rates. Okay? And then the other thing which is an important side note is that when you do RL, right, remember that I said that we're no longer in probabilistic world, right? So when you're doing supervised fine-tuning or when you're doing pre-training, it's quite clear that what you're doing is you're trying to do probabilistic modeling on some distribution, right? You're doing distribution matching for something, right? But RLHF is all it's all just a policy. There's no underlying necessarily distribution. And how this manifests is that often you'll get much less calibrated models, right? A number of results across many different papers have shown that, RLHF models often, especially at temperature one, show much more overconfident behavior. So, this one was from one of the anthropic papers. This one was from, I think, the GPT4 release. This one was from a paper that one of my posttos did. In all these cases essentially the RLHF models are much less calibrated and maybe that's fine right because calibration isn't part of the reward that you're putting in. So so that's as designed but you should be very careful thinking of these models as sort of calibrated probabilistic models which you might be sort of tempted to do if you're kind of coming into this from a generative modeling background. So, okay. So up until now, we've been kind of thinking about RLHF and all these things. Actually, I'll stop for a moment. If anyone has any final questions for RLHF, I'll answer them before I go on because we're going to totally switch topics to RL from verified rewards in a moment. Yes. Going back to the human

09:24 Preferences graph, what's what are the two axes that we're looking at? This one? Yeah. Yeah. So, the x-axis is proxy reward. So, here we're fitting like a reward model. We're not we're not doing DPO here, right? So let's say we're doing something like PPO where we fit a reward classifier and the X-axis is saying how well did my RL algorithm optimize that reward classifier. Right? So this is success at RL on is on the X- axis. Y-axis is this is kind of like true win rates either measured by AI feedback here or human feedback. So this is like real human votes on the y- axis here. And so this is saying you might think that succeeding at RL will make you succeed at the actual task that you're trying to optimize for. That's not the case. It will overfit at some point for at least human preferences. Yes. So question why is it written in that not like beta log a minus log b or just like log of a b like are you just asking like why don't they like pull the fractions out in the log or because I thought is this like a numeric computation thing that like when you take these two ratios the taking the log is easier to compute and that's why we want to do no I don't think that I think taking ratios is actually add numerically. So I think this is really just like a save the space the ratios are more compact and they're also intuitive objects right so the moment you put simple it becomes a different form because you no longer have the same beta concept yes that's right yeah SimPO is just like a very different motivation for why you would do this thing right there's no ratios it's a very fundamentally different object

11:03 Yes in the graph where you showed like a more sort of hacking thing I was wondering like how you like how the measure like the ground truth of like the reward score, right? Cuz presumably the reward function is a model like the human preference, right? But it's kind of like a train test gap, right? The x-axis here, you have a train set, you fit a reward model on it and you're measuring that fitted reward model, right? And the y-axis is fresh samples from the true reward oracle, right? And so they're measuring in expectation the same thing, but they're not measuring in finite sample the same thing. Like the same this is the same as like train test gaps in machine learning right same conceptual object they are in expectation measuring the same thing for any finite sample that's not the case okay good so I think all this questions about reinforcement learning from human feedback I think brings up a good point which is that human feedback is very difficult to optimize it's very difficult to scale and so you might wonder are there other kinds of more effective RL that we can bring to bear

The pivot: go where the reward is checkable

12:07 On post training right And that's going to get us here, right? So, up until now, we've been kind of talking about ChatGPT / GPT-3.5 and that era of models. And now I think we want to talk about o1 and the the new set of kind of reasoning style models, right? And the way we're going to get there is kind of this thinking of saying, okay, we have this very powerful tool now. We have this reinforcement learning tool that we can use for post- training. And initially, the instinct was to say, there's one true objective, and that objective is whether people like the bot, right? And so, if we just optimize for the true objective we care about, that's all we need. Turns out to be very difficult, right? Human human approval is easy to hack. It's hard to collect at scale. It's hard to RL at scale. For all those reasons, you see, overoptimization and all these kind of issues. On the other hand, we could instead turn to where RL has really dominated, right? Things like alpha go or alpha fold. And then looking at those domains you might say what we really need is kind of a domain which we know the true reward and for which we can evaluate such true rewards very quickly and efficiently and at scale right and if we could do that then we can bring to bear all of the successes of reinforcement learning from the past into language modeling. Right? So that's the overall mindset that we're we're we're adopting now. Right? We have the tools that we've built. Now let's apply them into a very kind of different way of using the same tools, right? Taking inspiration from successes in RL. Okay. And so today I'm going to talk in two parts, right? So the first part is just going to be an algorithmic part,

13:44 Right? So this part on the very top I'm going to talk about different algorithms. First I'm going to talk about PPO in more detail because I kind of cut that from last lecture for length reasons. And then I'm going to transform PPO into GRPO. A simpler version in some sense and then we'll talk through those objectives and various implementation details of GRPO to get us to multiple like canonical variations of GRPO and having done that we will then walk through the three I think big reasoning model sort of training things that have been

PPO in theory: policy gradient, TRPO, clipping

14:18 Done as case studies since after part one you'll have all the tools that you need to understand everything that is happening in part two right and after you do part two you'll understand at least how these three Chinese open LMS were made. Okay, so now we are going to return to PPO. It is the thing that I didn't want to do but we have to do in order for you to understand why GRPO exists, right? So what is PPO? At least in my sort of mental model of what PPO is, we start from the simplest possible thing and then we move slowly but surely downwards to PPO, right? So the simplest possible thing I mentioned this in last lecture. This is the policy gradient, right? So on the left, I would like to optimize the expected reward under my policy P of theta, right? And we're going to optimize that through gradient descent. So I take the gradient of that object and I get the right hand side of that equation, which is the policy gradient, the expectation under my current policy of the reward and I'm going to take gradient steps that sort of either increase or decrease the probability according to the sign of the reward Z. Right? Hopefully straightforward. If this is not familiar to you, you can sort of brush up on how that policy gradient derivation happens. Now there are two things that we might want to think about here that are kind of inefficient, right? So the first thing is that this is what you might call purely on policy, right? So the way that the reinforced gradients work, I have to sample from P theta and then I will immediately take a step on those sampled examples, right? And so every time that I want to take a gradient step, I have to compute a reward, right? And I have to do a roll out. You will sort of understand this in your in your

15:54 Assignment five, but the expensive part of RL is the rollouts, right? We have to actually run the language model and get samples. And we know that's slow, right? From Percy's inference lecture, we know that's very complicated. It's very tricky. We would much rather sample less frequently, right? We we do a roll out once, we sample once, and then we take multiple updates on those rollouts. So that motivates something called TRPO, right? Where instead of taking updates from P theta, I would like to allow my sort of updates to go stale. So what I'm going to do is I'm going to sample from P theta old, this sort of base distribution down here but still get sort of valid policy gradients. Well, how can I do that? Well, I can do what's called an important sampling correction and then I can keep, my policies close to my old one so that I don't get too far from my old one and get sort of crazy reward estimates. Right? Very simple idea. Just do a little bit of off to on policy correction. You get something called TRPO. This a of T is a lower variance version of R of Z. I'm not going to talk about advantage estimation in much more detail. PPO is a is a simple extra step beyond this. It says, instead of doing KL divergences, I'm just going to clip the advantages. And this is naturally going to force my policy to remain close because if I go too far, then I'm not going to be able to get higher and higher rewards. These rewards will just get clipped off at sort of a one minus epsilon or one plus epsilon. There's an upper bound to how much reward I can collect. So there's no incentive for the RL algorithm to make the policy really different from the current one, right? So it's a soft version you can say think of kind of this idea. Okay, so PPO very

17:31 Successful RL algorithm used in many different places used in sort of very toy RL environments it was used in OpenAI sort of Dota bot works very well in such actual RL tasks oh yeah there was a video I forgot about that yeah look you can run around yeah okay and at a conceptual level it's not that complicated right if you go look at the openai documentation for PPO it looks fairly simple right you see that equation that I listed above, it's this PPO clip objective. And really the only thing that's maybe slightly complicated is that this a of t is actually calculated by using something called a value function. So you need a second neural network to essentially compute kind of the expected rewards. And that would be used to in some sense lower the variance of my gradients, right? I'm not going to go into the details once again of sort of all the RL details but the one implementation difference here that matters is that we need this value function and that will become important in a moment here. But PPO in practice is a very different beast from PPO in theory. and you're in for a very bad time if there's a blog post that says 37 implementation details of PPO. I don't want to know all 37 of those guys. And it's a there's a huge long list of sort of different PPO variants in this blog post and all of them have the different scores on the RL benchmarks and there's a whole paper written on why implementation details matter in PPO and how if you really mess it up, you're not even computing the policy gradient correctly anymore, but it actually works better, right? So it's kind of a

PPO in practice: a walk through the AlpacaFarm trainer

19:11 Crazy place. If you look at the implementation details of PPO and so we actually do have to maybe look at a live implementation detail. I'll actually go over it quickly because PPO isn't the point of this lecture. But what I want to get across to you is, you all haven't been necessarily working on this post- training space. Like why do people care so much about alternatives to PPO? It is because of things like this, right? PPO works well, but it is kind of a beast. And this is one example I think from the early days of RLHF re-implementations. There was this nice diagram describing, this is how you would do a pretty standard implementation of PPO. This diagram is pretty gnarly, right? We've got a reward model for RLHF. We've got a value model that's supposed to keep track of kind of the expected reward. We've got generalized advantage estimation. I haven't even explained what that is. And that's going to go into the policy language model. And we're going to do, our policy gradient updates. So there's there's all this sort of machinery that has to go into place for PPO to work. So we're going to look at an implementation. A few years back now actually some of my students and I did reimplementations of algorithms like PPO and other people have used it. So this is an implementation that's been somewhat tested. Not to say that this is a good implementation. I'm sure something like more modern implementations from Volcano Engine (veRL) and other things probably work a little bit better but I just want to walk you through all the different components. Also because as you write GRPO in your assignments your structure will probably mirror some aspects of this right all RL algorithms have similar kind of outer loop. So if we look at

20:52 This right so a PPO step looks very similar to a standard gradient update step in a language model. So this is the outer sort of loop. What we're going to do is we're going to get a bunch of rollouts. And once we get a bunch of rollouts what we're going to do is we're going to compute some losses and then we're going to just take backwards and then gradient steps with clip norm and that's it right very unintimidating. The outer loop of RL is usually fine, right? Loss computation is also basically the same thing. So, this is probably small for those of you in the back. But if you look at this, what you have is one block that's computing the loss for the value function. Remember, PPO has both a value function and a policy. So, we have to update both of those at once. So, we're computing how close is the value function to kind of the actual sort of returns, the rewards that we saw. And we want to keep those close because the value function is supposed to provide variance reduction. And then we've got the actual sort of rewards that we're going to update on for PPO. And then we have sort of the clipping constants 1 - epsilon 1 plus epsilon. This code is literally just copying this back over. And just to give you sort of some, intuition for hyperparameters, a typical clip range might be something like 02 here, right? So, so you would be allowed to go essentially from 0.8 to 1.2 two in likelihood ratio between your old and new policy. The rollouts is where things start to get gnarly, but this is still essentially just calling inference. The left side is really just sampling, rollouts from your current policy and then you evaluate the log problems of your samples and you move on to the right and what do you

22:30 See? Well, the only subtlety here is that your value and your reward and your policy may have different tokenizers. So you got to do retokenization. But otherwise you just feed your rollouts into the model. So okay, at this point you're like looking at this code and you say okay that's fine like this is not so horrible. But the places where things start to get a little bit tricky. Is first we have to do some reward shaping right? So what does this mean? So stepping slightly back. One thing that's weird about reinforcement learning for language models is if you think about it from the RL perspective technically it's kind of like a contextual bandit. Well, what is a contextual bandit? A contextual bandit is something where you get an input and you have a bunch of possible, actions to take and you immediately get a reward, right? In language modeling, you get a prompt, you can give an output, and you immediately get a reward. There's no like state transitions. There's no environment exploration. There's none of this like complexity, right? Reward shaping here is essentially constructing per token losses in order to be able to give you something more easy to learn for the RL algorithm. And so what happens in practice for PPO and for GRPO which you'll be implementing is that the KL terms essentially the regularization that we're applying is actually computed per token whereas the actual true reward like the did you complete the task or not kind of thing that's computed at the very last token. Right? So you see that there's essentially a per token reward for the regularization and there's a single terminal reward at the end for success or not. Yes. What does it mean? What does not work? Oh, right, right, right. This is

24:09 This is a Right, right. This is kind of funny. This basically we're only adding per token KL penalty, but once again, this is one of those funny PPO implementation things. This second line here, this is the actual true KL calculation. But when this number goes to negative, it becomes numerically unstable and you can actually hit that pretty often. And so this gets clamped off at zero, which all of if you clip off a log likelihood ratio at zero that's not a KL anymore, but it is some approximation to a KL. Yes. What are so we've talked about like boss function. What are our update steps especially like as it relates to the reward function? Is that the same as other Yeah. Yeah. Good good point. So the updates as I was saying in this outer loop is actually just taking gradient steps. So in terms of your code it will look no different than just taking normal gradient steps. What is happening here though is in some sense let's go back to the policy gradient equation here. Right if you write down this thing this is r of z * p theta z. This is kind of like just taking gradients with respect to theta of rzp log p theta z. It's a weighted loss that you then take gradients of. And what you're not doing is technically if you're taking a real gradient, you should also take a gradient into P theta, but you don't do that. There's kind of a stop implicit stop gradient. And so all you do is you compute this inner loss, you feed it to the autograd, and you'll get the right steps for the rein for the policy gradient. Yeah, you'll have to do that for the assignment as well. And we have like a little tutorial that explains kind of what's going on there in case that

25:51 Quick explanation was not clear. Okay. And finally, the part that I think is the gnarliest is there's a the box that I did not explain called generalized advantage estimation, right? So what you need to do whenever you have a policy gradient, the variances of the gradient are often very high, right? And so you want to do as much variance reduction as you can. And instead of multiplying the gradients with the rewards directly, you can instead you can show that the following a of t quantity which is essentially a discounted advantage estimate. That's a sort of RL term is an appropriate substitute. And so one of the things that PPO does that's very different is to use this advantage estimate where you can tweak lambda and gamma to trade off bias and variance of your gradients. One of the funny things though is despite kind of the annoyance of this implementation complexity of like maintaining this value function and estimating all these quantities you can just pick gamma equals lambda equals 1 which reduces this whole thing basically to a baselined policy gradient. So your guts taking R minus the implied value. And that works well too. And so the point of making you go through this was a lot of the implementations details for PPO are both simple in that it is just the outer loop RL. You do just take gradients and also kind of annoying, right? You have to do all these sort of clipping things. You have to think about what am I going to do with the generalized advantage estimate. What am I going to do to train the value estimate? But you expect to see, overall increasing rewards. Including reward models and sort of negative KL rewards kind of go down. This is a once again a contextual bandit. So you expect to see pretty reasonable training curves,

27:31 Not like crazy RL ones. Okay, so I went through PPO. That was kind of a whirlwind tour, but I think hopefully you get the context of a what PPO is and b that it is sometimes a little bit tricky to get working right. And so, this has motivated a lot of people to try to find alternatives to PPO. What we're going to want to do is to apply, essentially these RL algorithms to settings where PPO applies. And PPO applies very well to general RL settings where you have sort of

Why yet another algorithm

28:03 Rewards. But we don't want the complicated implementation. And maybe more importantly, we want to get rid of the value model. That is actually if you try to implement PPO really, really annoying, right? Because a value model is usually as big as your policy. So now in terms of GPU memory, you're paying twice the cost of your language model. Right? Now you might say, okay, why can't we use DPO? Well, DPO is well suited for like pairwise comparisons like Bradley-Terry comparisons. It's not so good if what you want to do is let's say do reinforcement learning on math questions and check whether or not the answers are correct, right? There's no pairwise structure so to speak inherently there. So maybe DPO is not great, right? DPO also, originally is kind of an offline algorithm in a sense in that it has a whole bunch of sort of pairs that you initially collect and you just update your model on those. You could make it online by iterating, but that's not usually the way in which people apply DPO. So then now this brings us to the new hotness which I think is GRPO, right? So GRPO is actually very simple both in motivation and in actual

GRPO: the group z-score advantage

29:09 Implementation. So where you start with is conceptually you start at PPO, right? You start with very similar pieces. You think about kind of this clipping thing. You think about policy updates in very similar ways. But what you do is you remove the really complicated generalized advantage estimation. You just get rid of it completely. And you replace it with something that is much simpler. Right? So what is the much simpler thing? Well, we are going to replace the advantage which used to be this like GA thing. It was the sum o over returns and had the value function in there. Instead, it's going to be this equation three at the bottom of this slide here. And what is this? The advantage of response I is equal to the reward that response I receives minus the mean of the responses within my group. And I'll define what a group is in a moment. And then divided by the standard deviation of the rewards within the group. So this is a zcore if what that is of the rewards within the group. Right? Now what is a group? A group is a very in some sense natural object for language model RL. You have an input question. Let's say like solve this math problem. Right? That is a group. And I have many different candidate responses. I have capital G different responses. And those are all the responses within my single group. Right? So the nice thing is if you think about it right maybe problems are harder or easier right some math problems are much harder than others and because of that the average reward that my sort of other samples receive is a natural baseline for myself and

30:46 People have explored exactly these kinds of algorithms if you look up like reinforcement learning with leave one out or sorry policy gradients with leave one out this is the leave one out baseline in action minus the standard deviation piece right the yes I was wondering talking about here like so you said it's a it's like they have a batch of questions right and then there are answers to it which are reported so like is each like answer only corresponding to one unique question or are we doing like multiple like kind of like trajectories of the same question for multiple questions right you do multiple answers per question and that's how you get kind of variance reduction in like multiple different questions and also like kind of like multiple answers for each of the questions. Is that a really important example? No. No. So for each question Q, GRPO samples a group of outputs, right? So so I mean this is not batched like you normally you would if you were doing this for real, you would have multiple questions and for each question you'd have G responses and those would get baseline together. Across questions you would have no baselining or any interaction really other than the policies get updated together. Okay. So the baseline is only doing within the same question. And that's why the baseline makes sense, right? Because a question is like hard or easy and the mean of that reward isn't sometimes capturing the question difficulty and you're subtracting that guy out. Okay, good. And the other thing, this is a fun note. I'm going to mention it because I just kind of think this is fun and cool. This DKL, if you've seen lots of KL divergence computations, this is actually a little non-standard, right? Because the

32:23 Natural kale divergence estimate is actually just you take a bunch of samples and then you compute the average of the log ratio which is kind of this inner term over here right but this one from GRPO actually has these two extra terms it has the log or has the ratio of pi ref over pi theta and has this minus one and you can kind of convince yourself that if you take the expectation of this with respect to pi theta this is just going to be cancelling with this one over here. So this is a control variant scheme that reduces the variance of this KL divergence estimate which is cool because maybe you all need to estimate KL divergence from samples. This equation two is just a slightly nicer way of estimating that exact same thing. Okay. And GRPO is really nice. Oh one last note. If you're only taking one step which is doing the pure online case, right? All this clipping stuff just kind of disappears and all you're doing is policy gradients, right? Policy gradient is just upweing good stuff and downweing bad stuff where the rewards that I'm multiplying the policy gradients by is this avi. So it's just like an incredibly simple algorithm for the truly online case where you're not doing multiple steps on a single example. So there's multiple kind of different repositories for GRPO implementations. I can point you to several including this one that I sort of copied and put onto this slide here. But basically I can just say that it's it's exactly the way you think it would go. Sort of in the outer loop you would compute the reward for each nor roll out you normalize for each group the mean and the variance of

34:03 The rewards. You'd compute the KL term in this case per sequence which isn't quite right for the more heavyweight implementations. And then you do gradient updates on the loss. And this is one example of the loss computation that you also have here. And the advantage computation unlike in the PPO case is just really simple. This is almost exactly line for line the equations that I showed you with just one minor difference that's not shown in the equation which is that as you do in almost everything because we're dividing by the standard deviation you add a tiny fudge factor of 1 eg4 to make sure it doesn't numerically blow up on you right and you'll have to do this in the assignment as well. You'll have to add a little epsilon to your GRPO setup. So, how well does this work? GRPO works pretty well. This is from the original DeepSeekMath paper. And I'll get back to this later because it is interesting to look at this plot and this result in light of later R1 results. They show essentially the two fine-tuning based methods RF and online RFT. This is basically a fairly weak baseline I would say. where what you're doing is you're only looking at examples that sort of get the right answer. They're doing math in this case. So you only get examples where you get the right answer and you fine-tune on your own outputs that got the right answer. Right? Reinforcing correct answers with fine-tuning. GRPO with outcome level rewards where you only get a correct or not answer is the yellow. The blue one is process level rewards where it's kind of you got a system that looks at each step of your reasoning and sort of gives you a grade for that. And they're arguing, maybe process rewards are better. We'll talk a little bit more about that later. But in either case, you see that GRPO works and

Is that a legal baseline? Dr. GRPO's two fixes

35:43 It works pretty well. Okay, any questions about the basic GRPO piece before we move on to kind of details about GRPO and thinking kind of deeply about what's actually happening in the algorithm? Good. Okay. So now let's think about the difference between GRPO and PPO and what we've done and what's different. Okay. So really there's, as I was saying, only one difference, although it's a really important difference. It's replacing the advantage estimator with this thing, right? The the mean or the zcore, let's say, of the rewards. And so now I'm going to kind of go back to the policy gradient theorem or the policy gradient result, and I'm going to think through this result with you, right? So when we when we take a policy gradient update, right, what can we do? So I'll just go back a couple slides, right? On this slide, we've got right here at the very top policy gradients. This is the most basic RL algorithm that you can do, right? You take gradients where you multiply the log probability gradients with the reward. Right? This is something that I'm always allowed to do. This is a mathematical equivalence. Right? Now, one other thing I can do is called baselining. I can take this reward Z and I can subtract any constant or any in fact any random variable that doesn't depend on Z itself. And this would still be a valid policy gradient. Right? So so this baselining thing is really important because what you're going to try to do is you're going to try to subtract constants that give you lower variances on this expectation. Right? So that's

37:22 Called baselining. And if we go here, that's kind of a classic result that you can look up in Sutton and Barto. They say, okay, look, we've got the policy gradient. You can subtract out any baseline B of S because B of S, when we sum it up across the policy is going to be zero, right? So this is fine, right? We can always baseline, but let's look at this A of A of I. Is this a baseline? Well, we're subtracting the mean and that's kind of a baseline because all the other rewards are not dependent on RO of I. So maybe that's okay. I mean, technically this notation includes RO of II, but if I remove that, that's a valid baseline. But one thing that's really weird is I'm dividing by the standard deviation here, right? That that's not something that really seems allowed according to this derivation in Sutton and Barto. And that turns out to be a problem. Some folks that have gone and reanalyzed GRPO and its behaviors basically argue that GRPO has two things that at least mathematically are a little bit off. And the first thing is this division by standard deviation right as I was kind of talking you through just now this breaks that sort of contract that a baseline just needs to be subtracting a zero mean variable that's independent of my draw. And the other thing that GRPO does, which I did it kind of, sorry, I glossed over when I previously presented it, is that it's actually dividing, kind of the rewards by the length of the output. And that's going to have that's also a little bit weird according to the policy gradient theorem. This is not something

39:00 That would naturally show up. And so these authors who did like this pretty interesting study of GRPO algorithms kind of argue that maybe we should just get rid of these two things. And if you do, then you'll actually have much shorter output length and higher reward without having much longer responses. And so let's talk about these results carefully for each one of these two fixes. And hopefully by talking through these you will gain an intuition about how the RL algorithm works. Right? So first I want to talk about the standard deviation. This one's maybe somewhat obvious what it's doing. Okay let me let me go back because I think it's easier to talk about when I'm highlighting the equation here. So I'm dividing the advantage by the standard deviation. So what does that mean? When the standard deviation is small, right, the reward is going to be amplified. It's going to be more important for me to optimize that group when the standard deviations are small, right? And when is the standard deviation small? Well, it's when the problem is too easy or it's too hard, right? Because that's when the rewards are either all zero or they're all ones. And so there's a bias in the standard deviation term that upweights problems that are too easy or too hard. The authors argue this slows convergence. Maybe true. At the very least, it certainly breaks the validity of the policy gradient. The second thing which is subtle but also interesting is the length normalization. Now let's kind of look at what's happening here. So we have this length normalization before the GRPO reward. Now what does that do? If my model got a question wrong, right? Then the best thing to do is I'm receiving negative

40:36 Reward in here. So the best thing to do is to make the response really long. And if I get the answer right, the best thing to do is to make the answer short so I can sort of maximize my positive reward. Right? So what this does is it actually produces a model that kind of BSes as aggressively as possible. If the model thinks it can't get the answer right, it just produces the longest possible response, which is a very bad incentive for you to give the model. And so if you fix this, what happens is, on various sort of toy tasks like GSM8K and so on, you can get a reward that's just as good. The red one is sort of the modified version, but the output length doesn't keep growing and growing and growing. It sort of like stabilizes at a certain point, right? And so there is actually some interesting observation that like maybe some of the really long CoTs that people are seeing in things like GRPO are a result of these like actual implementation details and choices rather than inherently long CoTs being a necessary part of the performance of these models. And I think that's like a very interesting although not fully proven out class of hypothesis. Cool. Okay. So that's the GRPO algorithm. Hopefully now you're all familiar with it and I think you now have the background to go through all three of these papers now. DeepSeek-R1, Kimi k1.5 and Qwen 3. Any questions? Yes. So a question about why like operating is bad just to confirm like my understanding. I guess it's bad because like in those in either the very easy or the very hard cases we don't want to actually like aggressively update the model. Yeah. I think this gets into sort of wishy-washy folk theorem

42:15 Territory, but actually some of these papers will talk about this folk theory and so I'll mention it which is what you really want the RL algorithm to do is to get problems that it can like do somewhat well on like it can get some reward on but is not so easy that it can like already solve them right so there's like kind of a curriculum effect where you want to feed the models the right level of difficulty and if you're maximizing the that standard deviation that's kind of the wrong direction Right? You're really maxing out the stuff at the extremes that you either already all know or are just way too hard for you to solve. Cool. Okay. So, we're going to talk

Case study 1: DeepSeek-R1 and R1-Zero

42:50 About all three of these papers today. R1, Kimi k1.5, and Qwen 3. R1 and k1.5, I think, are pretty interesting because they came out at roughly the same time. Sadly, R1 was the only one to get like a gigantic social what's it called? Reception. But both of these actually show how to do RL based reasoning on like math and other things with LLMs, right? And they because they're contemporaneous you can kind of see almost two parallel ways to tackle the same problem like which things are similar, which things are different and so on which is great. Qwen 3 is the newest of the releases. And they do some fairly interesting variations of ideas in R1. And also they have some new kind of tricks that R1 doesn't have, which I think are pretty interesting things to look at, especially if you're interested in things like inference efficiency, of reasoning models. So, I'll start with R1. I think R1 is, kind of amazing in the sense that it's a it's a, archive paper that launched a whole social phenomenon. Never let your advisers tell you that, your archive papers will never matter. This one lost like what almost half a billion dollars of Nvidia valuation. You two can one day maybe cause that kind of a wave. R1, I think, is quite remarkable because it in many ways replicates all of the qualitative properties of the o1 recipe in a way that is extremely simple. So the key properties that I want to talk through and make sure you all

44:26 Understand. The first thing right is that it hits the performance targets that OpenAI o1 set. Everyone was really excited about reasoning model. So this is very exciting. The second thing is that it opens up a RL recipe that is not only just like a replicable one but I think more importantly one that is extremely simple right it doesn't have any search it doesn't have any process reward models. I think lots of people at the time thought maybe we need all these complicated pieces to get reasoning models. R1 really shows you don't need any of that, right? And then finally, there's lots of interesting insights about the interaction between supervised fine-tuning and RL that I think continue to be really important. Okay, so the starting point of R1 is they build on DeepSeekMath. And actually some of the equations I showed you of GRPO are from DeepSeekMath where they originally proposed GRPO as an alternative simpler or system more systems efficient variant of PPO, right? To them actually the most important piece was they wanted to get rid of the value model just because it's really annoying to have around. But one thing that's really interesting is they actually go for this yellow line the outcome supervision which is not actually the best performing model in DeepSeekMath. Talk about that again at the very end of this section here. So I'll walk through all the different pieces of R1. So I'll start with R1-Zero which I think of as the controlled setting, right? So R1-Zero is a very pure form of RL learning. It basically takes essentially the model that is

46:01 Pre-trained plus mid-trained before doing any RLHF or instruction tuning and then throws it into the math RL loop and then they try to find out how well does that do right so the details here how do they do the reinforcement learning well they have a bunch of sort of mathish tasks the data is not public they take DeepSeek-V3 as their kind of base model and then their rewards there's two forms of rewards that they use. One of them is an accuracy reward. So like did it get the math question correct? Right? It's a correct or not reward. It's binary. they have a format reward that basically forces the model to put its CoT within like thinking tags like thinking start thinking end of thinking tags, right? And that's important if you want your model to have these like long CoTs being used. Now, the format reward feels like something that doesn't matter, but from many papers and from having talked to many people, apparently it is a pretty critical part of actually getting this whole like reasoning RL thing to work. Once you do this right, all they're doing is doing RL on top of the base model. Nothing very fancy, but the results are pretty striking. They get performance that is getting pretty close to OpenAI o1 by just doing some RL on top of the model that they already had, right? Without any like, CoT fine-tuning or anything like this. And there's kind of two things that they note in their paper as being really interesting about R1-Zero. And I want to talk about this because I think it's important to carefully examine what is happening in R1-Zero. So the first thing they say is, "Oh, it's very cool that if you just let the model do RL on this

47:41 This verifiable rewards, the length of the CoT just kind of increases like pretty predictably." and in commentary that I don't know if I necessarily agree with, in the paper they're like, "Oh, it's learning to solve harder and harder problems by thinking harder and harder." It's like, well, maybe. They also, point out it's kind of cool that, they learn, phenomena like backtracking. They call this the aha moment. I think much has been made about this in sort of public discourse that wow it's cool that RL training can give models these kind of emergent insights. I'll kind of refer you back to the Dr. GRPO paper the one that was talking about the corrections to GRPO. And I think they have honestly pretty good and interesting arguments that both of these are not particularly interesting phenomena. Like first of all they argue that like the length just goes up because of the biased objective not because it's an inherently interesting object. And second they argue well if you just run DeepSeek-V3 on a bunch of math questions it'll also sometimes output things like aha I can do this or that which is maybe not like a deeply new phenomenon that arises from RL. Both of these seem like given more recent evidence kind of credible things that like maybe there's nothing like emergent and special about R1-Zero but it is actually working very well right that is a good math model. So R1-Zero you can think of as kind of a research setting, right? They're taking a controlled model, they're doing something very controlled on top of it, which is math RL, and they get a good model out, right? But if you're trying to build a really strong model that you're going to ship to the world, this is not what you're going to do, right? You're going to like basically do everything that you can to get the best model that you can, right? So what would

49:19 You do in that sort of more unrestricted setting? Well, you're going to maybe insert some supervised fine-tuning. You're going to take, CoTs from, some undisclosed source and you're going to fine-tune your DeepSeek model on that before you do your RL. And after you do that, right, you don't want your model to be kind of like this math savant that can't do anything else. So, you're going to, apply your usual post-training pipeline on top of that to make sure that it can do all the other tasks that, people normally want to use these models for, right? So, this is kind of the the pipeline differences. And so the key differences both within the pipeline and within the RL is they do SFT initialization to try to get the model to know how to do long CoTs without starting with RL. They add a language consistency reward in order to make sure that the chains of thought remain in a single language and then they do a secondary RLHF stage kind of at the end. So I think this makes a lot of sense, right? Every time you want to do something advanced like reinforcement learning, you're probably going to start by doing a little bit of supervised fine-tuning. and so even in sort of reasoning models or long CoT models like DeepSeek-R1, this is the case, right? You start with long CoT supervised fine-tuning data and then you're going to, do RL. I will point out that the description of where they get this data and what this data is remains very vague. Like they don't tell us like what was the CoT data derived from how did they filter it I don't really have any idea based on reading the R1 paper the claimed benefit of this is that

50:57 If you CoT the model on sorry if you sft the model on long let's say English CoTs then this gives you an interpretability benefit right as you do RL you're not going to get like weird gibberish it's going to kind of keep the model closer to these like more interpretable CoTs that you started out with So and that would be kind of good for users, right? Like as you're, using a math model, it would be nice to see its reasoning as it goes. And an additional thing, right? So so when they do SFT initialization, they use a ton of data. But one really interesting thing is that for a lot of models, even a tiny amount of SFT on these kinds of like long CoT data can be good. One thing that some of my students in collaboration with Percy, we did was, basically take a bunch of long CoTs from Gemini 2.0 flash thinking and, fine-tune Qwen 2.5. And maybe surprisingly, with just a thousand examples, you get really, really high, math benchmark accuracy, with just a little bit of long CoT fine-tuning. So I think both of these are really pointing to the fact that the base model already has a lot of kind of like thinking capabilities that you're just like kind of priming and extracting from the model. And after that of course you're going to do RL, right? So after you've gotten the model set up with SFT, as with kind of the instruction tuning and RLHF pipeline, so you start with SFT and then you do RL to basically get the model to actually optimize the rewards you're looking for. The RL part is basically the same as R1-Zero. Not not huge differences, but with a minor difference that you're going to add a

52:34 Language consistency loss. And they have I think I think this note is pretty interesting. This is a minor note, but I'll describe it anyway where they basically say they add this like language consistency reward because, during the training process, if they just let the model RL, they find that actually the CoT will language mix like it will switch between languages. And if many of you have kind of seen people playing with reasoning models, I've seen people on the internet post things like, oh, it's kind of weird that like Grok 3 like suddenly switches to Chinese and the CoT. And this is kind of consistent with those kinds of things that if you aggressively RL a model, like actually, there's natural tendencies for models to language mix rather than staying in a single language. So it actually requires an additional reward to keep it in the single language. And then finally after you've done RL on like math and other verifiable domains you basically layer on the usual post training. So you do instruction tuning and then you do sort of the pair wise preference tuning afterwards right so they do an SFT spe where they combine both reasoning data on non-verifiable tasks like write a proof of something these are not verifiable and then so they use their own model as a judge for whether or not they got the answer correct. They have non-reasoning data like sort of write a nice essay and they use the same SFT data set as what they used for DeepSeek-V3 and then finally for RLHF they actually still use GRPO for RLHF which is kind of cool they use the same RL algorithm for everything and then they basically just follow the V3 RLHF pipeline like there's not really anything different for this

54:13 Post-training part right how well does it work works very well I think Many of you probably experienced this as well, right? Like R1 was in many ways a shock because it matched the o1 performance, really kind of across the board on a very simple recipe, right? Like as I describe this, I don't think any of you found any of it particularly surprising. But the outcomes kind of speak for themselves. You know, you've got sort of on the English tasks basically, tied or matching o1 on across the board. Really slightly worse on code

Distillation, and the two things that did not work

54:46 Models but very close really across all these different tasks. The final thing that the R1 paper showed is that you can take these big models and you can distill them into other models. So you can take your big DeepSeek-R1 and you can take those chains of thought like in their case they take almost a million chains of thought and then they fine-tune Qwen with those chains of thought and they actually get big boosts in sort of math performance relative to the base model which is only getting something like 50% performance on AIME for the 32B model. So they get 25 plus% boost on this task which is which is pretty surprising. Cool. And then finally there's two I think interesting and good observations from R1 and I think scientifically maybe this was the biggest contribution of R1 in a way. So I think R1 maybe had three contributions scientifically, right? One of it was it showed that outcome based rewards with GRPO works, right? It's like the positive proof and then R1 also had two other kind of like negative results kind of contributions and they are sort of contained in the very last part of the R1 report and they basically say okay like we tried two things like pretty extensively. We tried PRM and we also tried MCTS. And neither of those really helped us at all on like replicating something like o1, right? And so to get into a little bit of detail, right, PRM are basically process reward models. Those are kind of like systems that can give you intermediate rewards on a proof, right? So when your model is giving a chain of thought, a PRM would be able to say, "Oh, you went wrong at this step in the middle, right?" And obviously that is

56:26 Much richer and very powerful form of feedback. An RL algorithm can make really good use of a PRM. But unfortunately, it's also very difficult to get a PRM in the first place. And so, R1 and the DeepSeekMath people like they had gone down this road of doing PRM for a while and they kind of concluded this doesn't work quite as well as outcome based rewards. And thus far, I think outcomebased rewards remain the way that in which you would build these models. The second thing that I think hasn't really panned out is searchbased methods. I think lots of people were interested in in search based approaches to reasoning. Thus far at least it hasn't really panned out in the same way that RL and outcome based rewards has that remains I think kind of the strongest kind of baseline and system in this universe. Okay. So any questions about R1 like kind of their setup or any of the other findings? Yes. PRM GRPO and PRM are kind of two different Oh yes. Oh yeah. Yeah. That's right. Yes. Okay. Good. Sorry. I said I was going to mention it, but I didn't. So that is totally on me. That's right. Exactly. And so I think especially for the PRM, I think it's really interesting and telling that in DeepSeekMath, they were very much convinced by sort of the strength of the PRM, which is this blue line with PS. And then in R1, they, concluded that actually this approach that had worked for DeepSeekMath was not really going to work for R1. And they had gone with outcome based rewards. Yeah, thank you for reminding me. I was about to forget my promise there. Yes

58:04 Consistency is that for I guess ease of understanding or is do they actually know how it protects performance? Yeah. So so yeah in this note here they basically say if you ablate away the language consistency experiment it results in degradation of the model's performance. But they're gonna put it in anyway because they prefer to have CoTs that are more readable to humans. It's an interesting trade-off. There have been lots of research about whether CoTs are faithful and they're not truly faithful. We kind of know that. But maybe it's better to have a slightly more faithful CoT than to have the extra, I don't know, half percentage point performance on AIME. Yes. Translated out the end, but if you if what you care about is that someone can read it, like so I mean I'm sure they can do like translations or other kinds of like post-processing to make the CoTs better. And in some ways you might think of OpenAI and these other vendors efforts to like summarize the CoTs as being very similar, right? Like because

Case study 2: Kimi k1.5

59:09 The raw CoTs are probably much messier and then they probably like sort of rationalize it away. That is one way and I think that is an effective way to get interpretability. But I do think if you're interested in somehow if you aesthetically believe that the raw CoT is very important for monitoring and closer interpretability then I do think you do want something like this. Right. Cool. Okay. So now we're going to move on to Kimi k1.5. And why do we study this one? If you look at the timestamps for when R1 and k1.5 were released, it's it's kind of, contemporaneous. And it achieves very similar results. It does so using outcome based rewards in RL. It doesn't use the same algorithm. It has kind of different details and different interesting insights. And so we can learn about kind of what's the same, what's different, and maybe, what parts are maybe important in this process. So to just show you the headline result before we start just to get you to believe that this is a paper that's worth being discussed here right Kimi k1.5 is the the dark blue bars here you can see OpenAI's o1 as the next highest bar so they're beating or matching o1 across the board on a bunch of important tasks and they do basically things similar to R1. So they do SFT, they do RL, they have a different RL algorithm, but they do RL. They also describe their data set construction in a little bit more detail. And this actually gets used in Qwen 3 later, so it's worth discussing. So let me talk through data. As Percy said

60:49 Earlier, maybe data is the most important thing in your whole pipeline. And so we should always pay attention when people talk about their kind of data curation, strategies in like a large scale training paper. And so Kimi k1.5 does several things to try to curate their data set. So what they first do is they try to balance across different domains like they have an automated I'm guessing LM based tagging system to categorize basically math questions by different domains and disciplines and then they kind of balance across these to try to get diversity across different domains. They exclude even though these are verifiable multiple choice and true false questions because they argue that these are too easy to hack or to randomly guess. So they're only looking for verifiable answers that are kind of short and can be evaluated by things like regex or LM. And this is maybe the most interesting piece of the curation here. What they do is they take the model that doesn't do any reasoning their SFT model and they have this model generate 10 answers and the pass rate is used in order to determine whether to include that example or not. And I think the exact thing they use for their latest selection strategy is they only select examples that fail best of eight. So if it if they can get any one out of a correct, then it's excluded for being too easy. SFT data similar to R1, very little description. Who knows where they got that from? They

62:27 Just say they do some prompt engineering. So clearly it was distilled from something else. But we don't really know what they distilled off of. So I'll talk about the RL algorithm. The Kimi one's kind of interesting. It's a different variation. It's actually maybe closer to DPO in a way. But you end up with a algorithm that I think you'll very much recognize. So you can kind of think of this in a very interesting way as convergent evolution of RL algorithms. So you start once again at the very top. This is our classic goal. You know we're sampling from a data set. We're sampling from our policy. We want to maximize rewards. We don't want to be too far from our base policy. So this is the KL regularizer. If you remember our DPO derivation you make a nonparametric assumption you say pi star is the optimal policy is an arbitrary function that means that the reward can be written as the log normalizer plus the ratio of policies right this is the exact same thing we did in dpo and now in dpo what we did was we took these rewards and we plugged them into the kind of the Bradley-Terry preference function we don't have this here right we're not doing pair wise preferences so we don't actually take that step. Instead, we write down this equation and we say we know that for the optimal policy, we're going to have equality here. And so, all we're going to do is we're going to make this a difference and we're going to add a squared loss on top, right? We're going to try to drive the left and the right sides close by just adding a squared loss. It's a reasonable thing to do. People have done things like this before. And this gives us our loss. It's basically trying to drive the right side and the left side of

64:05 Something that should be an equality for the optimal policy close together. And well, this is a little bit of an exotic looking object or maybe exotic looking initially, but if you take the gradient, it just looks a lot like GRPO. You've got your gradients of your policy, right? This is your policy gradient stuff. And you've got a baseline reward. And actually, what's the baseline reward? I'm just going to average the R within my batch, right? So here this is actually doing something different. This is the normalizing constant I think over the batch. But we're doing essentially similar kinds of baselining as GRPO and we've got a slightly different like square of the log loss regularization to keep my policy close rather than doing clipping. Right? So zooming out, what's happening here? Very similar to GRPO in that this first part is a baseline loss, but there's no standard deviation thing happening. This second part is analogous to the clipping that happens in GRPO, but instead of doing clipping, we're explicitly regularizing the policy, right? So we've got the same ingredients just in slightly different form. And so hopefully you can see, that as long as you kind of have this policy gradient thing and the right baselining and something that looks like regularization, you can get a working RL algorithm. The other thing that the Kimi folks do, and in some ways I think this is more forward-looking or they got this more right than the R1 folks, which is that they realize that, if you're shipping a reasoning model, what you really care about is inference cost. And if you care about inference cost, you had better try to control the length of your CoTs, right? If you have really long thinking chains, that's going to cost you a ton of money or cost your

65:42 Users a ton of money. And so instead of celebrating the really long CoTs, the Kimi folks say, we want to really compress the CoTs as much as possible while keeping performance high. And so they have this length reward thing here. Where what they're doing is for kind of each batch, they're looking at the maximum and the minimum lengths. And what you have is lambda where lambda is roughly like, where are you in the range of lengths within your batch, right? So if you're at plus.5 you're really short and negative.5 you are really long right and the reward is basically going to be lambda whenever you get the answer to be correct. So if your answer is correct you're going to incentivize yourself to be really short at the very shortest end of this range. Whereas if your answers are incorrect, you're in this bottom part of this length reward, which means that you're incentivizing the CoT lengths to be roughly shorter than the center of the range of the rollouts. This is a somewhat funky loss to me to be honest, but you can kind of understand the dynamics of this loss, which is you're incentivizing sort of correct answers to be as short as possible and incorrect answers are incentivized to be kind of averageish, right? And so you don't have as strong of an optimization pressure to push down incorrect answers to be short. And one final note on this length stuff is the Kimi folks realize that if you add this re reward early on in training, it stalls RL because, it basically forces the model to say I'm in a local minimum. I don't get

67:18 Any of my answers correct. The best I can do is to have my CoTs really short and then you just can't get out of that sort of local minimum. And so they actually only turn this on later on during training. So they kind of initially do a bit of unconstrained RL and then they add this length reward in afterwards. And they also have additional cool details. Who knows how much of this stuff is necessary or important, but they actually have a whole curriculum set up. They basically have assigned difficulty labels. So the data set just kind of like top down like they just manually or via LMS kind of annotate the difficulty labels. And then they go from easy to hard in that order. And then they also as they're going on sample problems proportional to one minus success rate. So if you're 100% succeeding, you just never sample that question ever again. And for the rewards, they basically for code they take problems with ground truth solutions and they generate a bunch of test cases. And for math they basically use actually a reward model where that's used to compare like ground truth human written answers to the LM output. So instead of using something like a regex or using SymPy which is what other people have done the Kimi folks actually use a model to do equivalence checking. It's kind of surprising to do this in a verifiable reward case but they don't seem to have a problem. The reward models are very accurate because it's really all it's doing is advanced string matching. One of the things that's really cool about the Kimi paper is that they also talk about kind of the infra issues that arise doing RL I think I don't think I've seen any of the other RL reasoning papers actually talk

Why RL infrastructure is hard

68:57 About systems almost at all. So, it's nice to see that they're kind of talking about this, what their structure is, what their layout of this is. You'll have to deal with this in A5, like I think a very mini version of this as you implement things like rollouts in RL. But one thing I'll note, right, is why is RL so hard to make efficient? And in many ways, it's harder than normal pre-training to have your GPUs like fully utilized during RL. And the reason I think is because there's rollouts involved, right? So you have to be generating sequences and whenever you're generating sequences not only are you kind of slow in the sense of inference is slow but also you have this other issue which is that you have to switch from RL to inference and back and you have to be passing sort of data back to the RL worker and the RL worker has to pass model weights to the inference server and vice versa. So you've got all this sort of message passing that can happen. And finally this is unique to long CoT models, but if you have really long CoTs, batches can become very uneven. So you have to handle that cleverly somehow. And so the Kimi folks do this, relatively nice but fairly standard thing of having different workers that are assigned to do the RL updates, having different workers assigned to do inference. And they have basically kind of message passing to be able to pass the weights to the inference worker and the inference worker can basically make data sets for the RL worker. And they have almost the same kinds of setups that you will have. They use vLLM for inference. And they also do very advanced stuff where they

70:37 Have to have vLLM with dummy weights and they have to kill it because of the complexities of doing kind of this passing of weights from one worker to another. Finally, I think one thing that I think is interesting that the Kimi people show in their results is they have like per iteration results where they show like as RL proceeds like how does performance scale and they show really nice scaling which is in the blue and they also show the growth of the length of the responses and much like in Dr. GRPO because I don't think their RL algorithm is inherently sort of balanced or sorry biased towards longer responses. Most of the sort of responses actually sort of plateau out at a sort of target length like we saw with Dr. GRPO rather than grow unboundedly like we did with the vanilla GRPO. Cool. Any questions on Kimi before we finally conclude with Qwen 3? So back to like the RL setup in detail of Kimi. Yeah I was wondering like so I guess like when you're like doing the rollouts during like training like the rollouts the inference is done by the app right? Yes. And as I update the parameter that will mean that like as they update the model they also need to sync that to the to the vLLM inference through worker and that's what this is for. Yeah. So I think the most annoying part at least with current libraries of this process is essentially that step of taking RL

72:17 Weights and putting it into vLLM. There is an experimental API that is supposed to allow you to use NCCL collective calls to shove a set of weights into vLLM. But at least we were thinking about having you use that for the assignment actually. But it has too many undocumented parameters for it to be a little bit mature. And so maybe next year's iteration this will actually be fairly mature technology. But right now I think a lot of people do things like start a vLLM with dummy weights. The weights then get loaded somehow into memory with sort of hacks on top of vLLM. And each iteration they often tear down vLLM to make sure they can free the GPU memory fully. Yeah, there's a lot of like I think the RL for LMS I think is fairly new still. So I think the infrastructural support remains a little bit immature but I think in maybe a year things will actually just get much nicer. Yeah, question back there. You have a lot of like accuracy rewards and like rewards. How are you supposed to combine them together? Yeah. So you how do you combine the different rewards? I think this is one of the RL not quite black magic but really like RL magic things of like you just tune weights like in all cases all the rewards are just added together but with weights and how are the weights determined empirically in order to maximize the downstream performance right especially for things like format rewards you can almost think of them as like shaping or like surrogate rewards like you don't necessarily really care about the formatting reward it's more of a means to an end to get a good long CoT within the tags that will get you the answer Cool.

Case study 3: Qwen 3, thinking-mode fusion, thinking budget

73:57 Okay, so the final one I want to talk about is Qwen 3. And thankfully Qwen 3 released their report before the end of the class, so I get to include it. And I think this is the most recent and modern of the RL for reasoning kind of models to have come out. And so we can kind of see how they've like built upon the previous works like where they've changed things. And they have actually pretty interesting scaling and data results that are are new and unique. So the overall picture is very similar to what R1 and Kimi have done, right? So Qwen basically take their base models. They'll do a long CoT sft stage. So that's this first stage over here. They'll do their reasoning RL. They'll do something funky that I'll I'll talk about later called thinking mode fusion. And then they'll do RLHF RL and then that's the model that they ship. Of course they then distill that in various ways, but we can kind of forget about that for the moment, right? So we've already seen this in R1, right? RLHF comes after reasoning and then distillation comes yet after that. And we already know a lot of the playbook. So I can actually go through this fairly fast. Much like Kimi, they basically curate the data by filtering for difficulty using best of end. So if your base model that hasn't been RL can already answer it, if you sample like end times, then you can just get rid of it, right? You they also do some like decontamination things where they remove things that are too similar to validation data. And then they also manually filter essentially their SFT set. So like for their initial SFT data for long CoTs they like manually filter it for whether they are

75:33 Guessing or whether they actually got it right. The one thing that empirically I think is really interesting about the Qwen 3 RL results is that they are actually doing this RL only on 3,995 examples which is a very small number of examples to be doing this on and they get pretty good gains out of the RL process. And so you can view this as RL on verified rewards as being very efficient. You could also think of this as being analogous to many sample efficiency results in the past like people have shown that you can instruction tune a model with very few samples or that you can distill a long CoT with very few samples but that doesn't necessarily mean that it doesn't continue to scale right we don't really know why. What this does show is that even with very few examples, you can sometimes do RL which is surprising and cool. So what is the two new sort of Qwen specific stuff that they do? The thing that they do is this thing called thinking mode fusion which I think is kind of interesting and where I think the field or the various trends are going is in controlling inference right so what they do is they want to have both a thinking model and a non-thinking model in the same single set of parameters so what do they do well after they train a model with RL they have a model that can do thinking and now they're going to fine-tune it again to do one of two things they're going to fine-tune the model with some data that has a think tag and then it's going to do the normal CoT thing and you can get this data from yourself right like the original thinking model can generate this or you can have a no think tag in which

77:11 Case it should immediately kind of emit an answer and in this case they're going to have to sort of supervise fine-tune the model to know what no think means and to immediately try to emit the answer and one kind of interesting side effect of this is that they found that if they do this training where they train the model to have a think and a no think tag. Then what they can do is they can sort of if the model continues to think and you want to terminate the thinking process, you can kind of actually like terminate the thinking process, add a special string that's like considering the limited time of the user, I have to give a solution by thinking directly now and then end think tag and then it sort of accurately gives the answer. So this gives them a control knob by which they can sort of more precisely control a maximum number of thinking tokens. And this gives them pretty clean test time scaling from a single model. And so they can set like a maximum thinking budget. And of course the sort of very maximum out to infinity is just the original thinking model but they can kind of early terminate to the left and they can get graceful degradations in model performance. That's not too bad at the very sort of beginning. You can like have the thinking tokens and still get some pretty good performance. And so you can look at Qwen 3 also does a nice ablation where they give you the performance at the reasoning RL stage at the thinking mode fusion stage and at the general RL stage. And one thing that's very interesting to me here is that if you look at the first two sets of rows the general task and the instruction

78:48 Following ones reasoning RL helps thinking mode fusion helps and of course RLHF continues to help here as well right like kind of everything helps in this regime but if you look at kind of math or stem performance in the thinking case general RL hurts performance and in the non-thinking case it helps performance so there really is seems to be at least somewhat of a trade-off in, do I optimize for general purpose instruction following up here or do I optimize for like math encoding kind of down here, right? And so those are kind of interesting sort of properties that are emerging and it'll be cool to see sort of more future models kind of sidestep these tradeoffs somehow. Okay, cool. So to put that all together our sort of initial motivation was to say RL is very powerful. We kind of figured out that we can do RL with language models in the RLHF domain. But you can't just like hill climb on noisy pair wise preferences forever, right? So one solution is to pick domains in which you can't do reward hacking and then just go for it, right? RL and narrow domains is one good solution. And GRPO is one very simple algorithm and hopefully you all kind of have a sense of like okay I can just do policy gradients with some good baselines and that will enable RL on all of these kinds of verifiable reward domains. And then finally there's lots of successful recipes in the wild and you hopefully now you've seen what's in common, what's different, what implementation tricks matter. Cool. Thanks a lot. And I will see you all next week.