Alignment: SFT and RLHF
Where this lecture sits: two lectures of post-training
00:05 Okay. So we'll get started. Welcome to lecture 15. We've got two pieces left to the class and that's going to be you know various aspects of post-training. Up until now we focused very much on the big pre-training systems data components and then now we're going to take the big pre-trained model and we're going to make it useful and safe in various ways. So that's going to be the next two lectures from me. Today is going to be RLHF and sort of safety alignment stuff and then Thursday is going to be RL from verifiable rewards. So things like reasoning training and math and so on will be on Thursday. As I said before, today we're going to shift from pre-training to post-training. Percy in the very last lecture did cover some stuff about post-training data. But really I think you know the focus today is going to be in going from you know essentially this big transition that we saw in the field right so we have GPT-3 really remarkable system really impressive lots of pre-training lots of compute but this is not really a useful system right I guess there was a couple startups around you know building ad copy and things like that but it was not very useful it didn't follow instructions it didn't you know do anything particularly too interesting from a from a product point of view And then all of a sudden, you know, we got ChatGPT. And ChatGPT can do all sorts of amazing things and follow instructions. And we've kind of seen what that has done to society since then, right? So today's focus is going to be on this arrow right here, like how do we take a pre-trained system like GPT-3 and how do we make something
01:44 Like ChatGPT? And then we're going to try to get to the to the nuts and bolts of that process. And I think many most of you, you know, have never, you know, worked on things like controllable generation or like the previous generation of text generation systems, but really like modern instruction following models are just amazing, right? Like this is one of my favorite examples from Sébastien Bubeck's sparks of AGI paper in 2023 around when GPT-4 came out. But you know it can follow this like very long block of like nested compound instructions and then combine that with its coding capability to output you know zero shot matplotlib code and I think all of you just like take this for granted now it's like yes of course ChatGPT can follow 10 instructions at once but it's just kind of amazing that it can do this right and I think part of my excitement about this lecture is the fact that it can do all of this and the other thing that I think is very important right is that now that these systems are out in the wild. Safety and content moderation just becomes really important, right? Safety from the perspective of, you know, these models might get misused. Someone might try to use them for scams. And also content moderation, if you're thinking about these this being like useful products that you can like ship and like people would pay for, right? Like people don't really want to pay for or put ads on systems that are like horrifically toxic, right? Right? Like I think one of the big reasons why ChatGPT has been so successful is that you know it has really significant guard rails around it. So okay given that the goal today is to try to enable much tighter
03:20 Better controls on language models right pre-training you can think of mental the mental model can be that it packs the model with all sorts of capabilities right like after pre-training the model is able to somewhere within the parameters do lots of things like reason and answer questions but it's not going to do them out of the box and so today what we're going to try to do is get models to do that out of the and so what we're going to do is to collect data of various kinds of behaviors that we do want from the language model. And train it to do those things, right? And so the questions that you should be asking now is, you know, what does that data look like? How hard is it to collect that data? Percy has touched on it a little bit, but given the importance of data, I'm going to re-emphasize it a little bit. I'm going to have some interactive some exercises to go over that. And then there's algorithmic questions like how do we me make use of that data, right? Certain kinds of data are easy to use. Like you know if you have expert demonstrations you just train to imitate that. But if you have things like pairwise feedback like model output A is better than model output B. How do we how do we make use of that? And then
The InstructGPT diagram as the map; part 1 is supervised fine-tuning
04:24 Finally like you know how do we scale this up? How do we do the usual things that we've been doing in this class? So the structure of this lecture is roughly going to mirror the InstructGPT paper because a lot of the post-training pipeline that we have today is still off the InstructGPT paper. And so the first part of this is lecture is going to be on supervised fine-tuning. So if you look at the InstructGPT paper, you'll see this diagram that roughly describes a three-step process for building an instruction following model. So part one of this lecture is going to be the leftmost part, the part where what we're going to do is we're going to do supervised fine-tuning on expert demonstration and then part two we're going to follow the next two parts of this structure. We're going to talk about reinforcement learning and pairwise feedback loop. Broadly I'm going to say that the ingredients in order to get, you know, this first part working, I mean, there's two things that we have to kind of think about. The first part is the training data, right? If you're going to imitate expert demonstrations, you better have expert demonstrations. Like, what does that look like? And then the second thing I want to talk about is kind of the method. Like, you have data now. Like, how are you going to adapt to it? And there's an obvious answer. I'll, you know, talk about this again, but like just do gradient descent. But there's also a kind of nonobvious part to this answer. And, you know, in case you haven't been following how people build these models today, this might, you know, be still surprising. So, I'll leave that as a teaser for later. So, okay. So, in Percy's lecture, you know, he's mentioned already several different kinds of instruction data, but today
06:01 We're going to like walk through a couple of them, and we're going to do a little bit of an interactive exercise. So you know those of you who have your laptops open can use those for good. So I want to talk about two different details like one of them is what's inside these data sets. People often say data matters a lot. I think post-training is one place where this is even more true than before. Because you're using very small amounts of data to get exactly the behaviors you want. So if you have noisy instruction tuning data, you're going to get some pretty crazy behavior out of your models. And then what kinds of things should we be paying attention to? If you're in
Three instruction datasets, three paradigms: FLAN, OpenAssistant, Alpaca
06:32 Charge of post-training data collection, what kinds of things might matter? So I have taken three different data sets from basically constructed in three very different ways like you might also you might even call them like kind of three different paradigms to building instruction following or post-training data. And we're going to go through each one and then we'll look at them closely and then we'll think a little bit about what's going on with these data sets. Okay. So, I'm going to talk about FLAN. This is by a bunch of Google folks and FLAN is going to be essentially constructed by aggregating a bunch of training data sets from NLP tasks right so if you look at it you know you see all sorts of different tasks like you know Natural Instructions v2 which has a bunch of you know question answering and things it's got T0-SF you know AdversarialQA and like topic classification so basically this was constructed by taking existing NLP data sets that do all sorts of individual tasks and then aggregating them into one big metadata set. Right? So this is one approach to building such data sets. We've got OpenAssistant on the right. And this was I think a pretty unique sort of endeavor in which a bunch of online enthusiasts got together and decided to write instruction tuning data for language models. Like right after the release of ChatGPT I think the excitement for this kind of thing was really high. And so there's actually a lot of good high quality human written data from that effort. And lastly, and you know, of course, this is this is a bit of self- advertisement here, but as a representative of the kind of like language model generated post-training data or like AI feedback style data, I'm
08:11 Going to talk a little bit about some of the data from Stanford Alpaca. So let's just look at examples, right? Like I think looking at examples and talking about them are very useful. Now this is from random examples taken from the FLAN dataset. And you can kind of see the types of stuff that are in here. So you know, you've got things that look like pretty normal instruction tuning data like write highlights for this article. Sauntering down leafy avenues past Dutch step-gabled buildings dot. And then you know in the end there's even like more information on travel in the Netherlands at www.holland.com. And then it answers you know the least known of the Dutch cities, The Hague was a village dot. And then it sort of summarizes, you know, this as a highlight. You know, you've got something like this where, you know, this is like what is this text about? Here are your four options. It's business, right? So this is kind of a multiplechoice training thing that's happening. This you know Percy talked about the Enron data set so you all can like you know smile a little bit but things like this of taking let's say maybe the Enron email data set and you paste back write a subject line for this email and now you've got kind of supervision for that task right this one I guess no one here has probably worked on text generation but this is from a data set called E2E where you have like a database entry and then you're supposed to write a sentence that describes that restaurant so immediately you kind of see you know you can probably get a lot of data for free this way, right? There's a lot of NLP training data sets and you can like put them all together and you will get a really big aggregated data set. And so
09:47 In that sense, FLAN was a you know ahead of its time and produced a ton of data for this kind of thing. But also we see in many ways that this can be somewhat unnatural, right? Like we can we've already seen that like this Enron data set is a little bit weird. We can really definitely see things like, oh, here's a text and then now you sort of append sort of the options to turn it into a task. And so you can kind of see the surgery that you have to do very visibly in order to make this kind of data set. And I think if you look at this, you'll agree with me that this isn't your usual chat interaction, right, for something like ChatGPT. Another example for this is Alpaca. This was like a really early attempt at using language models to generate instruction tuning data. And so you know just to describe the procedure here you know a language model was used there was a seat set of human written instructions and then the language model was used to essentially generate more instructions. So that's the left column. And then you use something like InstructGPT to essentially fill in the response, right? And so here, you know, now we have something that looks a little bit more like, you know, on the left standard sort of ChatGPT inputs. If you compare this to something on the left, this is a very benchmark ccentric set of tasks. This feels a lot more like a set of interactions that someone might just throw into a chatbot. And the response is almost always in long form natural language versus with FLAN often it can be quite short like one word or a phrase or something like that right so we kind of see that of course you know
11:24 We also see that on the left these are in some ways not very diverse inputs they're very short instructions and then OpenAssistant is kind of the third leg of this instruction tuning saga you kind of see more complex queries on the left and then because Back then, I guess people were just really into writing long detailed supervision for models. You see like actually really detailed responses and this one even has a citation on you know how like what makes this answer correct, right? And so, you know, very high quality
Exercise 1: writing an SFT response by hand
11:58 But also kind of very difficult. And so now this is this is the first interactive task in this class. But especially those of you that have your laptops open you know, please go to this URL. This should be a Google form. And there will be a sort of one sort of sort of prompt. And now we're collectively going to crowdsource a instruction tuning response. And I'll give you all let's say five minutes to do so. Let me know if the if the link is wrong, but I did test this last night. So hopefully this is working and then we'll look at the responses briefly. And then I want to talk about why I did this exercise. There is a teachable moment here rather than sort of getting you just off your laptops for a moment. Okay, excellent. I think I think there's a decent number of responses, so I'm going to maybe put them up. Let's see if I can put them up then. Yeah, there we go. Okay. So in many ways, I think this reflects the kinds of data you get. I mean, if anything, you guys are all motivated to do this task than I think the standard crowd worker, you know, but you've got the person. I mean, I'm not sure what's going on here, but this is probably ChatGPT. There's a lot of emojis. I'm getting trolled, but this is I mean, I am preparing for lift thoughts. That's good response. You've got, you know, the nafam, which, you know, of course is the kinds of things that you'll get out of crowd sourcing. And I think hopefully one thing that you've seen or like felt
13:38 As you were doing this is that it's actually really difficult to write long form responses, right? Especially for something that you aren't prepared for, right? And so, you know, you get a lot of short responses like this. It's very difficult, I think, to get people to write sort of long detailed, you know, responses like this one at the very top. Often those kinds of things are you know from ChatGPT. And of course, you'll get things like, you know, I saw this one. Ice cream is a frozen dessert typically made from milk or cream. And you're going to have to, you know, filter out those kinds of things that you're going to get through crowd sourcing. So why did I talk about this? Well, of course, you know, now you have a sense of what this task is like. Annotators in the wild, even if they're experts, will be under time constraints. And I think one of the reasons why, you know, things like AI feedback or using LMS to try to refine or generate these kinds of data has really gotten popular is, you know, if you look at the GPT-4o response to this, you know, it's pretty good. It's a pretty good response to this question. It's very long. It's very detailed. And generating this kind of a human response is going to take a lot of effort and a lot of cost, right? And so you have to, you know, if you're in charge of human data collection at one of these labs, you have to think about, okay, like how do we take, you know, what I showed you in the spreadsheet and incentivize people to generate something that looks like this instead. That is no easy task at all. That is a very difficult crowding task. Okay? And so these things that we've just seen they vary quite a bit in things like
Style, length and list bias in preference judgments
15:15 Length. We saw you know the ChatGPT which is bullet points like lots of style variations. We saw in the OpenAssistant example that you know sometimes people put in references sometimes they put in sort of very complex deep knowledge like is that good or bad? I'll talk about that in a moment. And there's also other important aspects to this process right like maybe you want to collect a ton of data or very little data that's high quality. So you got this trade-off. You also have to think a lot about safety, right? Like the data that we collected just now, that's just capabilities data, right? It just makes models answer things like what is CS336? It does not help us make our models, you know you know refuse malicious instructions and things like that. So, we have to think a little bit about what that kind of data looks like, too. Okay. I think length has always been a big kind of gorilla in the room issue for all of these data sets. You know when back in 2023 when it was very popular to generate these kinds of instruction tuning data sets there's a survey like Yizhong Wang and others at UW came up with this very nice survey coming looking over the many different kinds of data sets that were created early in that year and you kind of see if you look at the length of both the prompts so that's the inputs and the responses the completions you see like really different lengths of both inputs and outputs and inputs are probably a measure of sort of complexity of the task in many ways and then the outputs are in some sense a measure of how you know much you push the annotators or if you used AI generated responses and one of the things that you should be you know all aware of
16:55 You know let's say you got put in charge of making a new language model you're in charge of post-training well you know if you're using human evals people have a strong preference for lists I mean so do actually if you use AI as a judge they also have a strong preference for lists And people have a very long preference for outputs, like 60 70%ish preference for longer outputs. And so do sort of AI judges, right? And so this is this is a little concerning because you want to be optimizing for not just kind of the stylistic content of your responses, you ideally want to be using post-training to do things like reduce hallucinations and actually make the model hopefully more capable. One thing that we do see is that you know these factors are not super relevant for benchmark performance. So if you look at for example MMLU performance kind of despite the really big variation in length for a lot of these models most of the instruction tuning data sets like the simple ones you know give you boosts over kind of the base model which is the very top row above. So I think one of the things that I'll say here is you know chat style evaluations have their place you know Chatbot Arena AlpacaEval these kinds of automated you know eval like user engagement and things like this but benchmarks also have a very important place because when you post train you don't necessarily want to be too affected by for example length biases and open-ended domains and so on and so you want to be really careful of these effects. You want to have a different diverse array of evaluation strategies to try to avoid those pitfalls. One other thing that I think really trips people up when they initially
The monopsony citation: knowledge, or the shape of an answer?
18:33 Start thinking about these things is to say, "Oh, what I'm doing is I want to collect high-quality data and high-quality data has lots of deep knowledge and has lots of citations, right?" Like that's a reasonable thing to say. And I think, you know, OpenAssistant, I think, had a great example of this, right? And so here's an example input output pair. You know, you've got right introduction about monopsony economics. And then there's these references on the response on the right. So now let's say we have a model and we fine-tune the model to take the left side as input and reproduce the right side as output. Right? So you can kind of think about two different things that this process is going to do at the same time. Right? So one of the things that this is going to do is it's going to associate you know monopsony with that citation. Right? So it's learning new knowledge. So that is good. Right? That is a positive thing to do. But it's also going to do a second thing at the same time, which is this kind of generalized thing of saying if you ask me a complicated concept, I had better finish the output with a reference, right? And so it's basically the sec the first thing is teaching new knowledge, which is good, but the second thing here is kind of teaching the model to hallucinate, right? Like if the model doesn't already have somewhere within its parameters an association between monopsony and this citation this Bivens's initial book you know what might happen instead is it just learns that oh what I should do is whenever I have a complicated input I should give a response and then make up a reference at the very end right those are two competing explanations for what is happening here and this is going to motivate in some ways the second part
20:13 Of this lecture right John Schulman has this kind of great talk I think he gave it at Berkeley where you know his argument is basically if you do this kind of thing you're going to encourage the model to hallucinate right like the model doesn't have the knowledge of answering a question you force it to answer that question what it's going to learn is of course it'll learn the knowledge in some abstract sense but it will also learn the other aspect of I just need to make something up in order to sort of type check what the response should look like okay yes there's a question so human right one writing has a sense of like okay I think I need to add a citation here let me search for another citation either from memory or let me actually use a database it is the fact that like the LLM is learning that like okay here I should insert a citation is actually a correct and desirable thing and it's and the fact that it's a madeup citation is a either a memory issue or b something you can maybe augment a tool usage in the site but like I don't see why the behavior of needing to add a citation itself self is problematic whe like if you can either fix the memory issue or the tool user issue. Sure. I mean I think the Okay, so to repeat the question, the question was like I guess that was a more comment than a question. Was that the learning to put in a citation isn't a bad thing, right? Like I mean maybe you augment it with tools and it'll actually give the right citation. I mean, that's a fair point, but I think the thing to maybe point out here is like the deeper conceptual or not conceptual, the deeper issue with token prediction, right, is that you're teaching the model to kind of predict the right kinds of tokens. And here, you know, essentially the lesser of the two errors is to say
21:51 Hallucinating is less bad for my loss than not making up the reference at all, right? Kind of the structure of the response always has to be fulfilled because you have to fill the tokens in at the right places, right? Of course, at scale, if you know the facts, if you have the right tools, right, those are those are good. You want to kind of make the predictions on the right spaces. But I do kind of think this is very indicative of this like this failure mode that models can get into where you're trying to get it to do things that it can't, right? Like if your SFT data is just much more advanced than what your pre-trained model naturally can do you run this risk of teaching the model kind of this alternative shortcut behavior instead of teaching models the right behavior. Yeah, so that's John Schulman. I think he makes a fairly reasonable case that you know this is one of the reasons why like on policy RL like reinforcement learning style things is an important thing to do because you want to know what the model already knows and only teach it those things to avoid you know hallucinating and whenever it's encountering some fact that it doesn't know then maybe you should change your fine-tuning data to say oh I don't know that fact instead of forcing the model to try to answer right and we kind of see this on the on the kind of other sort of sort of knowledge storage studies as well. Where people have kind of talked about, you know, it's much easier for models to sort of reproduce known facts than sort of to learn sort of unknown facts where it just takes a lot longer for models to kind of learn facts that aren't shown in pre-training. And this sort of is
23:26 Sort of matching what you might expect from these phenomena. So, okay. I think one of the things that I'll sort of you know summarize that with is that there's a very counterintuitive phenomenon for instruction tuning which is that you know you can have a instruction tuning data set that is fully correct and like actually very rich but actually that might not be good for your LM because it's going to teach your language model to sort of try to make up facts to match that depth of knowledge. That's always been, I think, one of the arguments for why you want to be really careful with both distillation data where the dis teacher model is stronger than your student model. And also really human annotation where the human might be much more knowledgeable than the model, right? You want to be really careful to make the model abstain nicely when it doesn't know things. And in principle, you know, reinforcement learning style correctness could help and we'll talk about that in a moment. And sort of optimizing this at the at the instruction tuning level is just really messy and very difficult. I don't think people have really nailed it down at least in the open research
Safety tuning, and the over-refusal trade-off
24:31 Literature. The other thing I want to talk about briefly because I think this isn't necessarily something that can be solved with instruction tuning alone is you know to touch on safety and to think a little bit about what the trade-offs are here. So, we know, you know, language models need some guardrails. They're deployed straight to end users. They're very capable. So, they might be used for misinformation or for generating things like scams or spam. And so there's a need for in safety tuning these models. And I think in parallel with a lot of the research on instruction tuning, there's been actually quite a bit of work studying safety tuning as well. And I think some of the early work in this in this area you know kind of has shown that even a small amount of sort of safety tuning data that's mixed in to instruction tuning process can make models much safer. Sort of paralleling a lot of the findings that people had that actually for instruction tuning as well if you have a strong enough pre-trained model even a small amount of instruction tuning data can get you a lot of the way. Not to say that's sufficient but actually that's you know sort of gets you to a reasonable point. And I think the core trade-off with safety tuning that I'll sort of touch on in this brief section is this trade-off between refusing things and not refusing kind of too much, right? So there's always this thing of, you know, if you have unsafe responses, you want your safety tuned model to just refuse to answer. And then maybe you have these other, you know, actually safe responses, but things that look
26:08 Like unsafe responses, like how can I kill a Python process, right? We all know that is a reasonable question to ask. But I guess if you're not, you know, if you don't understand English very deeply, you're like, oh, killing, killing sounds very dangerous, so maybe I should refuse to answer that question, right? So how can you make models sort of understand this nuance? It's a very tricky thing to do you know purely in the instruction tuning setting. And so a lot of what people have done is come up with carefully curated small instruction tuning data sets to try to balance this trade-off. So even some research has shown that even like 500 examples can make models follow some of the safety guidelines well so okay to put this together instruction tuning is surprisingly powerful. I think you know you would think that given how powerful things like ChatGPT are that there's actually a ton of complexity into getting anything that works. I think you'll find that even if you take you know a fairly standard instruction tuning data set like OpenHermes or OpenAssistant or any of these data sets and you take a base model and you fine-tune on it with reasonable hyperparameters, you'll get a model that behaves a lot like Llama or ChatGPT. It won't be quite as good. There's a lot of extra work you need to do to optimize it, but you can get pretty far. The second thing that's you know good to remember is basically it the notion of high quality data is just very complex and you have to reason about it really carefully. It's not obvious how to do and then the last thing is you know actually even a small amount of data can have great leverage at this stage in changing how models behave. The last thing I want to end this
How to fine-tune: gradient descent, then midtraining
27:48 Section on is how to do this instruction tuning. There's kind of a you know a flippant answer to this which is well you've got demonstrations just put in the instruction and the response and just do some gradient descent, right? We all know how to do gradient descent at this point. And I think in most academic settings that's basically it, right? Like you're done. You do your small scale gradient descent and you're done. But I think if you're at like a frontier lab and you've got more compute and you've got more money than you know what to do with, then you've got a lot of compute and you've got a lot of data. And so you can scale this whole process up quite a bit. Like you can scale it up a lot. And modern instruction tuning pipelines are starting to look a lot like pre-training pipelines. And so increasingly the boundaries between pre-training and instruction tuning are just getting blurred. Because if you think about it, instruction tuning data is still a sequence, right? It's just a sequence of tokens. And so I can throw that in into my pre-training process and that's a totally valid thing to do. And so this is an increasingly popular idea. I think you know the close labs don't tell us anything but you know I think the things that people have tell me in bits and pieces suggest that this is you know what they're doing. A lot of the open groups from China do basically this now and so what you do is you have your you know usual pre-training setting right you do pure pre-training and then what you're going to do is you're going to start mixing in instruction tuning data into
29:24 Pre-training. So the kind of the tail end of your pre-training, especially as you're kind of annealing the learning rate you're going to start putting in a lot of this higher quality data or instruction tuning data. And then in the end, you might actually do a second short instruction tuning round, but maybe this is smaller because most of your data has already gone into to the second stage, what people call midtraining. And this is cool because it lets you scale up without catastrophic forgetting issues. You might get more leverage out of your data because it's integrated more deeply into pre-training. And to give you an example or a sense of what this looks like and it's a bit of a shame that data mixes are often pretty closely guarded secrets by a lot of the groups. So I've taken this figure from MiniCPM, which we've talked about before great paper from one of the Chinese groups, where basically they have, you know, a two-stage training pipeline where they have a first stage where they do pure pre-training. And if you look at this pie chart, this is all pre-training data sets. Common Crawl code pre-training pile Dolma. They've thrown it all in into one big pie. And then they have a second stage which they call the decay stage. And so if you remember my lecture on scaling laws, you know, I talked about WSD, warm-up, stable, decay. So that's the stable stage. That's the decay stage. And in the decay stage, what have we got? We've got you know, Wikipedia, what people might call high quality data. We've got still the pre-training stuff mixed in there. So, it's not pure post-training data. But then if you look at the right, we've got code SFT, we've got Chinese books, we've got UltraChat, we've got
31:03 Stack Exchange question answering and Evol-Instruct and OSS-Instruct and all sorts of other things. So those are all kind of instruction tuning or instruction tuning adjacent data sets that you've thrown in onto the second half of pre-training. And it's I think used by most models today. And MiniCPM and other sort of you know derived LMS that are derived from that lineage of models are have definitely sort of publicized this I think it's extremely effective to do this and so I think everyone has been following this. One last commentary I'll make before we move on to RLHF here is that this whole process makes it very difficult to reason about pre-trained models versus post-trained models. Right? If you look at recent releases from like Qwen or what whatever other companies that you're looking at and they say base model you know that base model is probably at the end of you know this process and so it has basically gone through an instruction tuning phase implicitly through its midtraining process. We don't exactly know what the mixes are for a lot of these closed models but it does actually mean that I think the term base models is increasingly questionable what that really means. So that's my sort of side comment that is useful for you if you're if you're thinking about base models so yes the data mixture you change it during the decaying stage like when you have linear activity it goes to zero that's when you get like these big loss is that like directly totally it's seen in that lasting stage. Yeah. So that's right. That was the motivation for a lot of the
32:44 The two-phase training for these groups. Like they basically use essentially the large drop in loss as a way to try to anneal the model into the right you know mode. I think there's you know increasing studies into like what's the optimal point at which to switch and I think it's a little bit more nuanced than that but I think a first order this has been a very effective recipe. Yes. Between lending one has trouble computing or I'm trying to think about the incentive citization thing from earlier does this also have that yeah so the I guess there were two questions packed in one but the question was like is this primarily for catastrophic forgetting and also does this help with the citations issue so to answer the second part first I think it doesn't help with the citation issue because you know just as sort of like a type signature The only things that can help with the citation issue is if you know what facts the model knows, right? So you have to either ensure that the model like always knows the citation facts before you show it the SFT data or you have to like check to see if the model knows it and then show it that data if it if it does know it, right? Which this doesn't do. This will always unconditionally put in those data points whether or not that citation is learned. So it has no way of fixing that adaptively. Catastrophic forgetting wise, I think that is one of the motivations that if you have so much SFT data your trade-offs are pretty tricky, right? Because unless you're going to do this kind of almost pre-training mixed in with
34:19 Post-training, you have to think about regularization. You have to think about tiny step sizes to avoid, you know, messing up your pre-training. And so I think this is partially motivated by sort of catastrophic forgetting adjacent issues. It keeps the model sort of more general. Yes, I understand with the with the John Schulman example foration is that if the model doesn't know the reference that's included in the post training data that it know it does know that fact then it would not have the same behavior. That's right. Or that's the that's the claim, right? That if the model did know you know this citation that's right here then the model wouldn't necessarily out of these two sort of competing mechanisms what it would learn is oh whenever I see this example I should retrieve you know my knowledge about you know Bivens and Mishel and then use that as a citation I think the reality of this is that it's always very complicated right like what does it mean for a model to know something how reliably does it know something and So these two kind of mechanisms are always maybe in superp position for a model. It's really just a question of which one is more dominant. Like if a model just has no idea about this, it's probably two that is, you know, more dominant. Whereas I think if the model knows it reliably, it's more likely that it's just going to learn the correct citation rather than encouraging broad general hallucinations. Yes. Have people tried putting into all of the pre-training data some kind of thought tokens that tell the model? Oh, I should
35:56 This looks like a fact. Let me check if I actually know it. And so I'm going to query myself, you know, and check if I'm if I'm getting consistent answers. And if so, I'll print out right. So in the entire pre-training process, I need to check myself. Okay. That is a very interesting idea. So just to repeat it, it's has anyone done something where you put in like kind of thought tokens or the model's checking itself for its knowledge of facts as it trains or something like that? Right. Is that roughly right? Yeah. So depending on how you interpret or like implement that exact idea, it starts to look a lot like reinforcement learning. Because for example there's a method called a Quiet-STaR from some folks here Noah Goodman and Eric Zelikman and others have done this where they do essentially they try to learn the thinking process of a model by sort of you know predicting what happens on the answer token and then based on whether or not it's correct it tries to like you know reinforce the model to have good thought process. Actually the even closer analogy to this is STaR which is the original paper which is if the model gets something correct then that thinking process gets fed back into the model training and if it's wrong then it doesn't right and it's kind of very similar to what you're proposing which is to adaptively train the model based on kind of correctness of its knowledge or whatever else. I propose you just give a thought process that say let me make use of this tool that checks myself and I can later at some point change what that tool is but the point is that tool will come with a response that says yes I do have the knowledge no I don't and then you' actually see that knowledge gets printed
37:34 Out or doesn't depending on I see so in this like tool use example are you imagining that kind of the fact would get replaced by a tool call or will the fact still be there like I think the key question is do you force the model to predict the fact tokens or do you just force it to predict like use the tool to look it up on Google token? No. You don't know what the tool how you implement it. But you see if the tool says I yes I do know it then the then you there will actually be some response printed out in the pre-training data and if the tool says no I don't know it then the pre-training data will just say actually I don't know this I'm going to skip and so right so I think that's hard because when you do pre-training like you have to know whether you know the fact to know whether to take losses on that knowledge token right like you can't defer it to inference time because you have to decide whether or not you're going to take gradient steps, right? And the other sort of logistical difficulty here is during pre-training time you have a static data set where you know you for computational reasons you would want a static data set and if you have a static data set you can't adaptively do updates right like anything that solves the hallucination problem has to be kind of reactive of the form what does the model know and then do I take updates on this or not at the pre-training stage that's very difficult unless you're doing RL style stuff at pre-training scale which would get you very close to that but still very difficult I'm happy to follow But I think hopefully that answers the question. Oh, there's more stuff. Yes. Okay. So if I for example in case of Llama in the green data we haven't
39:16 Deal with the emoji but if you have a lot of the emojis will cause the effect that when we start to the model and the result model we don't have a lot of repeating sentence similar to emoji partners at the end. Yeah. Okay. So the question was if at pre-training we don't see emojis but at post-training we put in a bunch of emojis at the end what will happen? It depends on the structure of the emojis I guess. If the emojis are dependent on the inputs in a very complex way and that's very difficult to learn. Maybe what the model will learn is well in post-training what I saw was a bunch of emojis. I don't have enough data or training to know what the complex pattern is. So the model will just learn to put a bunch of random emojis at the end. If there's no pattern, if there's no comp complex dependence, then you know maybe the model will learn to do the right thing, which is just to put a bunch of random emojis at the end. Really the key way to think about kind of the SFT issues is instruction tuning will reliably teach the style of the output like the type signature of the output, right? And the model will most likely follow that type signature at the very least. And the real question is, do you have enough instruction tuning data that you could do something more than that? And that's kind of the more complex open question. So in your emoji case, at the very least you'll get a bunch of emojis. Whether those emojis are the right emojis, open question, right? Depends on how much instruction tuning data, depends on pre-training, so on and so forth. Yes, I was wondering like earlier in the lecture, I think it was said that the post-training part doesn't really teach like model new knowledge, right? It's mostly about styles and then like
40:52 The line kind of gets more blurry like when M is like me training stage. So and then he showed like the MiniCPM paper. So like in that in this sort of new like scenario like the training part could also instill some of new kind of like world knowledge into your model. Yeah that's right. So I guess the question was like you know if I rephrase it you know can instruction tuning essentially teach new world knowledge because midtraining blurs the line between pre-training and instruction tuning and we know pre-training teaches knowledge. So why not instruction tuning? And I think that's right that
Part 2: from imitation to optimization
41:24 Like in some ways instruction tuning if it's scaled up enough and it's diverse enough will teach knowledge right but I think instruction tuning in its like smaller like non-midtraining form it is very difficult to have the scale and diversity of data needed to reliably teach you know various facts. I think modern midtraining is starting to become a different game. But it's still sort of an emerging object I think. Cool. Okay. So now we get to the part two, right? So part one, this is the quick intro to instruction tuning and SFT. Now we get to reinforcement learning from human feedback, right? The RL part of this lecture. And conceptually, right, and I'm going to take it slow here because I think this is an important conceptual transition, right? We're going to move from the world of I think generative modeling at the very top here which is you know there's a very simple goal in this world which is there's a p star is a reference distribution from which completions are drawn that reference distribution probably looks like some mixture of you know internet data as well as annotator written data but there exists some abstract p star that we're trying to imitate right that's all there is to it so this is pure generative modeling now we're going to move to a second perspective now which is RLHF and in this world I no longer care about matching any distribution right so probabilistic perspectives kind of don't really entirely go out the window but you want to be careful about adopting those because really what we're looking for is we're really just looking for some policy P of Y given X such that we maximize our rewards right there's some reward function R of Y and X that you
43:06 Know take in both my completion and my prompt and then gives me a reward and all I'm looking for is any policy that gives me good rewards, right? And so now LM are not necessarily a model for some underlying distribution. They are policies that give us good rewards. And so why would we go and do RLHF? There are kind of two reasons that we might do this. One of them is on the top one in SFT in order to do this process of imitation we have to get samples from P star and that can be quite expensive, right? And the second one all we need to do is get measurements of rewards R. And SFT data can just be really expensive. You know this is kind of like a caricature of you know various costs that you might have in different stages right you have some compute costs when you train your base model and then you've got you're doing like supervised learning like SFT and then you're going to go collect a bunch of you know pairwise feedback and do RL and do evaluation and so on. So when you do this, right you know, the SFT is just really expensive. You're getting like really expert people to write very long form responses. And you kind of saw how annoying and difficult that was. And Frontier Labs are going to be spending millions on this post-training data. Well, maybe there's a nicer way of collecting data that makes models better. So that's one argument for why we're going to do all the things that we're going to do in the second part of this lecture. There's also a second reason that is equally or maybe even more important that I think you know people do this kind of RL training and one of the
44:46 Things that's very interesting is you know people don't always agree with themselves about what is good. So if you ask somebody to write a summary you know they can write one summary and then you ask them to compare their own summaries to LM written summaries there's a good amount of people that will actually prefer LM written summaries and this was a really surprising result from one of my students papers I guess two years back now where we were benchmarking you know summarization systems and there was one person who is you know of course anonymized annotator one who you know wrote a bunch of summaries and they actually preferred you know the AI summaries actually significantly more than their own and they're like a freelance expert writer or something and we went and like interviewed them and they were like yeah when you asked me to write stuff I just felt like I had to write more with like flowery language but then I like read the AI ones and they just read better you know and I'm sure you've had similar experiences you know where you look at the output and it is actually you know different from your own assessment of how to generate And so there's a, you know, not just cheaper to verify than generate, but actually maybe higher quality to verify than to generate. And so there's this generator validator gap. Okay. So we're going to cover different aspects of this RLHF process. We're going to, you know, talk about how we collect data and like what are things you should worry about if you're in charge of RLHF data collection. And we're going to talk about how we do RLHF. I'm going to talk about two representative algorithms PPO and DPO. I'll defer some of the
Collecting pairwise feedback: guidelines and exercise 2
46:24 More like detailed explanations of PPO till next lecture just for kind of space reasons. And then finally we'll end with some you know things to worry about almost like pitfalls of RLHF at the very end here. Okay. So how do we get pairwise feedback? Right. So pairwise feedback is oh sorry I'll go back here and just like take it a little slower. So when we do this second part of this InstructGPT process, right, like how does this work? Well, we have the model sort of generate its own outputs, right? These are rollouts in RL terms. And then we're going to compare, you know, these different outputs. And, and although four outputs are shown here you know, in kind of standard settings, you often just have a pair of outputs. And let's see, we have A and B. All I want to know is A, is A better than B or not, right? And then given these pairwise feedbacks I'm going to train a reward model that can essentially internally give every single output a scalar value. And then just use that to do reinforcement learning. Like this reward model is now my rewards and I want my model to maximize those rewards, right? Fair fairly hopefully simple pipeline. So how do we collect pairwise feedback data? Well, you know, the obvious simple thing to do is just say, okay, like, you know, I'm just going to make some web app. You know, I've got two different AI responses. And you're going to have a little, you know, four-way check box that like, you know, checks which response is better. I took this from one of the studies that we did, right? All a lot of pairwise feedback responses look, you know, similar to this. But I thought, you know, one thing that would be helpful and useful into like actually getting a
48:02 Sense of like what does this look like for real? I have gone and sort of dug up examples of annotation guidelines from different papers and different places sort of you know talking about this process right so if we look at the InstructGPT guideline this is one of the very few I would say sort of released materials from one of these companies describing their annotation guidelines you know they say okay your job is to evaluate these outputs to ensure that they're helpful truthful and harmless right So those are their three pillars. You've got helpfulness which is like you know writing in clear language you know answer the question they mean to ask like being sensitive to internationality like if someone says football you know they shouldn't assume American football. If it's too confusing ask for clarifications. I want you to be truthful and not hallucinate outputs. And by harmless you know you should sort of be not toxic and be very nice and not say NSFW things. All fairly reasonable things but you can kind of see how sort of there's a the interplay between things like the model spec which openai publishes publicly then there's a very detailed annotation guideline which you know this is not that detailed it's probably much bigger in practice where you would write down these kinds of bullet points you would hand this to annotators and then they would sort of go and make annotations this was you know InstructGPT is not even like kind of production grade this is early days but you see how this process kind of works. The kind of other interesting example you can kind of go look this one up later if that the
49:42 Actual text is too small because I'm not going to read through all of it. I'll just sort of touch on it a bit. There's actually a leaked version of the actual the true annotation guideline apparently for Google Bard that I think was part of some like news story. And you can kind of see very similar things happening here, right? We've got on the top left box like helpfulness like you should address the intent of the user's prompt, you should adhere to any requirements, don't have misleading information, like very similar to the InstructGPT setup. You know, we've got actually here a style box which is, you know, what kinds of style are good or bad. And then we've got, you know, different rating scales for the different responses. And if I remember right, I think the Google Bard folks I think gets like a minute per question to be doing this task which is which is quite difficult. Okay and then for InstructGPT you know they go through like Scale and Upwork and they collect about data from about 40 people like this is really tiny sort of by today's standards but hopefully you kind of get a sense of you know what types of groups are being involved here. Okay, so this is the second part of our interactive exercise. Okay, cool. So that's five minutes. I guess to take a straw poll I think about 27 of you managed to complete this, which is great. Thank you for your participation. You know, how many of you managed to fact check all the facts? So one of the questions yeah the fifth
51:22 One the long word expression completely factually wrong right yes that's right so how many of you managed to fact check you know any or all of the facts you managed to fact check things just one just one okay how many of you managed to check the math in the in the five minutes okay so there a couple people okay good excellent okay that makes me that makes me happy so I you know the point of this exercise partially, I mean, I think maybe you could have guessed what was going to happen here. The shorter ones most of them are essentially taking the longer ones and actually just removing the hallucinations to the best of my ability. And so for the most part, basically, for those of you that pick the longer one are sort of picking the slightly longer but hallucinated ones, like I don't quite remember which of the paragames ones are hallucinated, but many of them are. And so you know we see strong disagreement and actually the longer one gets more votes despite having strong hallucinations. This one I think both is correct. So B is probably the better choice. The two math ones I think you can also back out by the unnaturalness of this construction. The one that's like sort of more conclusive actually is you know going to the wrong conclusion. This is not mathematically fully correct here. And so you know you probably should find it difficult to try to do these judgments this quickly. There are very strong challenges in collecting this kind of pairwise feedback at very large scales just because even though you only have to verify if I'm showing you a math problem or I showing you know
53:01 Something like a very you know strong factually laden you know text you know you're going to have to basically break this down into claims and check each one to know whether one is you know wrong and one is right so this is just like a very labor-intensive and difficult task And you know I gave you five minutes because you know essentially I think the news story that the Google Bard article was associated with was basically that you know the annotators were given one minute per example and you know there was you know big complaints that like this is not enough for us to judge
What crowdsourcing does to your model
53:33 The safety and factual accuracy of the model and you hopefully you know somewhat agree that it is a very difficult task right. So okay the lessons here going back to the you know the flow of the lecture right it's very hard to get high-quality verifiable annotators like I think you are in many ways like pretty high on the bar of like people that are actually motivated to do this because you're just doing it because I asked you to not because I'm paying you to or to get grades. It's very difficult to get people to check correctness especially under time constraints. And the last one I don't know if you know those of you who are doing this, right, but if you put in an online survey like this, you know, someone's just going to like take the whole thing, dump it into GPT-4, and then just like kind of copy the answer straight back onto your pairwise responses. You know, we've had several studies in the past where an annotator had like 95 plus% agreement with GPT-4, and you kind of have to wonder what's going on there. And so it's, you know, despite the fact that pairwise feedback is easier to collect than supervised imitation data there are still significant issues. And I'd be remiss to not point out, you know, the many things that have been written about sort of the kinds of problems that this creates if you try to outsource this to third countries. There's sort of pricing concerns and sort of lots of ethical issues that you all should be aware of. You know, if you're sort of going to be in the future, be a part of these kinds of sort of data collection pipelines, right? You want to make sure to get high quality data and to make also make sure that, you know, people are being paid kind of living wage. The other thing
55:10 To be aware of you know this is in the sort of bias and safety angle because for alignment I think this is a really important thing to also touch on is that you know in some sense RLHF and alignment come at the end of the pipeline and because they come at the end of the pipeline they have very strong influence on model behaviors. One of the papers that actually Percy and my postdoc Shibani and Esin a postdoc on mine worked on was this paper on trying to figure out like how do subjective opinions of LMS align with different you know groups of people and one of the really interesting patterns that we found was actually for InstructGPT like these are old models but you know the now still useful models. These models somehow became more aligned with sort of Southeast Asian religions than before. And then we looked in the appendix of InstructGPT and actually like what's the nationality of the people doing the annotation? It's Filipino and Bangladeshi and 17% American. I was like oh it's kind of surprising but it does circumstantial evidence lines up you know with this kind of thing. And so you do have to be very careful because this is the thing that in some sense goes out in ships. You know others have also noted that you know depending on the annotator what they pay attention to is very different. I really like this paper from Hosking, Blunsom and Bartolo where they basically study you know two different kinds of annotators. One is like the authors and they're like very motivated to like you know judge things correctly with quotes. And then crowd workers and what they really find is you know crowd workers don't pay much attention to factuality. That's kind of this row here. They pay more attention to formatting, right? So depending on the annotator, you're kind
AI feedback takes over; length as the standing confounder
56:47 Of getting different kinds of feedback even though you're asking them the same things, right? And so kind of increasingly people have turned to as with the instruction tuning phase AI feedback where LM generated feedback. And so there's been many works including some of our own that have shown things like oh if you try to get pairwise feedback from GPT-4 it has very strong agreement from like GPT-4's you know estimated win rates of models or responses and sort of the human estimated ones on the y-axis. Same here. The agreement between human to human which is the blue box here is roughly the same as the agreement between GPT-4 and human and it but is much cheaper right so you know there's lots of reasons I think why sort of AI feedback has become popular and it has been used very extensively in RLHF so if you look at UltraFeedback which I think is one of the very popular open source data sets for sort of off policy RLHF you see this if Zephyr 7B I think was a Hugging Face effort I want to say last year to build a big strong open model and the reason why I bring this up I think Zephyr is probably not the most well-known model but one of the things I kind of remember about sort of the Hugging Face model process was you know initially they were really interested in human data collection like they were convinced that like human crowdworker like if you paid them enough and got the right vendors you know would outperform sort of AI generated feedback but kind of late in the process they kind of realized, you know, GPT-4 generated feedback for this kind of thing just worked much better. And so a more
58:26 Modern example of this is Tülu 3, which is a sort of post-training paper/ project out of AI2. And they've done kind of roughly this like they take different prompts. They have lots of different models to generate responses and they have, you know, LM sort of rate these to get chosen versus non-chosen. And this really all just kind of goes back to the classic paper would be the anthropic paper on constitutional AI which I think sort of you know really planted a flag on the ground in terms of AI feedback being used for this kind of alignment process. Finally the last thing I want to talk about for data is length effects. I think you know when we did the annotation the one of the things that we saw was I think many of you saw the longer response and you're like this is more detailed. I like details. This is good. It's not just you. Models and people all have this bias. And so people have found that you know models that people have thought were better were in fact maybe only just also longer. And then AI feedback seems to make models just generally longer. And so this is always a confounder that you want to be careful of you know length as a confounder for general preference. Okay, there was a question there. Could you expand on off policy and on policy? Okay. Yeah, I'll talk about off versus on policy later. You know off policy I think I mentioned it here is going to refer to you collect these like kind of pairwise feedback things separately like they're not collected from the outputs of your
60:03 Model. Maybe your model is involved like for example in Tülu you know you've got all these kind of off policy data on the left that's models that are not your own but you also have on policy data from yourself right so the off policy data kind of tells you about the landscape of places you're not at and the on policy data tells you how to refine yourself yes so the way the sort of human aspect is we're sampling prompts and then asking humans to grade them. Have people ever instead of sampling prompts that you don't have the answer to and then asking for the screen you have like existing in the world from any number of places. There are existing prompts that we know the answers to and we have that data saved. We don't need to do anything to receive that data. Have people done sort of alignment tuning using that style of data instead of asking humans or LMS? Yes. Okay, that's a good question. And it's like why don't we do RLHF on or like can we do RLHF on domains where we know the answer is like one way of putting the question, right? And part of the answer is the next lecture on Thursday is kind of that like where do we really know the answer? Math. We know the answer very well for math and we can do exactly RL against math and it works very well. People also do things like you know here is a long form response from an expert that has already been written. Now can you judge it given this? That helps but one thing that I think is important to keep in mind is
61:40 There's for a lot of these open-ended tasks there's many correct answers and it's very difficult to judge which are correct. Like if I have a new fact in the LM response like is that a correct fact or not? Doesn't really solve those problems. Yes. So like thinking like for instance the anthropic paper you referenced there's this like relatively specific and small constitutional thing but we can even like sample values to make this work for open-ended things like cuz we have any number of sources of like here's a problem here are values that we as a society think are relevant whether that's surveys or politics or you know articles just any number of like sources people do that I know that like anthropic should like let people vote but that seems small yeah like I guess there's lots of different things being mixed in on that comment like you could do like deliberative democracy style stuff the line models that is certainly a thing there's also I guess the thing that feels close to what you're talking about is almost giving the annotator sources almost that are relevant and I'm sure those things help there have been works on like showing people expert written responses when they do the pairwise judgment and so on. But I think it's not a silver bullet like all of this like kind of UI interventions definitely help but it's not necessarily going to in one stroke solve the problem. Okay. Yes. Question is there to like paper where you have like you're using to annotate like your other responses there's like a bias from like
63:22 Yes there's a very strong or like very detectable self preference for most models for their own outputs and so when you use this for evals which many people including I have done you have to be very careful for that self bias. Absolutely. Yes. Okay. Then there's lots of hands, but yes, I think you were first. Balancing the amount of information you get just from having model create its own outputs and then like I guess model to do like its own feedback, right? Is there like a heristic for how many iterations that does that change if you use like the model itself? Yeah, I guess the question of like how much can you extract out of models doing things to themselves is an interesting question, but I guess the kind of in some ways the information theoretic bound so to speak is very high because technically the model like ingests the entire pre-training corpus and that could be stored somewhere in the model and depending on how you prompt it you might get an amazing model out and so it's possible that based on how you're using the model as part of your self-refinement loop you can extract more capabilities out of the model and we don't know what the upper bound of that is because you know basically the inputs are just vast. But I do think there's like practically lots of papers that have studied this question of like how much can self-improvement help like what's the scaling properties and so on and so forth. I think ultimately it's a very empirical question. Okay. Yes. I think the as kind of like problem because compared to SFT so basically SFT is the best answer we have a label for that but
The objective: Bradley–Terry rewards and a KL anchor
65:02 For assume that we don't have label but we have like reward model to give it like preference has like ranking things. Yeah. So maybe I should just go through the next slides because I think that will actually explain it. So not only does it help us get through the rest of this hopefully quickly but also I think it'll explain your questions. So, okay. So, now let's talk about methods. Our goal is to do this thing, right? I want to find a policy to maximize my rewards. So, I need to tell you what the rewards are and I need to tell you what the maximization process is, right? So, let's do that. Now I think as you can probably tell by the frequency with which I'm referring to the InstructGPT paper, the whole sub area of instruction tuning and post-training is very closely tied to InstructGPT. So from the InstructGPT paper you have equation 2 which is this you know objective that describes what we're optimizing. So we've got this r theta of x of y. This is our reward and I'll define that in a moment. And then we've got you know this the second terms the log ratio of my RL policy divided by my SFT output. So what is this object? This is the KL divergence between my RL policy and my original SFT model. So it's saying when I do RL don't move too far from where I started. And then this second line this gamma term this is basically saying keep doing pre-training while you do RL so you don't catastrophically forget if you know keep doing this. Lots of people don't do the second step but this KL thing is a really you know it's a standard thing. It remains even today. Okay. So now what is the reward?
66:39 The reward is this thing kind of at the very top. And so maybe that's a small equation. So I'll talk through what that is, right? So there is a sort of hypothesized model of the world that exists. So what is the hypothesized model of the world? It's that every single output in the world, right? Like that a language model could output. Every single sequence has a scalar value R associated with it. And we don't observe what that R is. And when a person rates it like when they do a pairwise rating of A versus B what they do is they compare the two rewards of those two sequences and based on the difference they'll take a coin flip. Right? So this is a logistic model of the difference of the two rewards. So every sequence is a reward. When I do pairwise comparisons I take the difference and I flip a coin. Right? This is the Bradley–Terry model of human preferences. And this is what's happening. And so when we want to optimize the reward, what we're trying to do is we want to output sort of the sequence that has the highest R. And R is not something we observe. We only observe noisy pairwise comparisons through R theta, right? So that's what we do. So that's the objective, right? So that's what we're trying to optimize. And so now let's talk about the how. And to be clear, I'm only going to talk about PPO, which is kind of the OG algorithm. This is what appears in InstructGPT and a lot of the open AI stuff. I'm going to talk about it only very briefly and then I'll talk about it in much more detail on Thursday because it more naturally belongs there. We're going to do a lot more sort of real RL on that lecture. So at a conceptual level remember what we want to do is we want to optimize the reward
68:18 Of some policy right that's the left term here. And what's a good way of optimizing something? Well let's take some gradients right let's take gradient descent. So that's the very top left equation here. Now you know we can sort of do a little bit of math and if we take the gradient of this object we can write down that this is equivalent to you know the expectation of the reward multiplied by the gradient of P theta. And so this is very natural what you're doing is you take your normal gradients that you normally take in doing like pre-training or whatever. This is saying P theta of Z I want to maximize that probability and I maximize or I multiply that with R. So if my rewards are positive I want to upweight those probabilities. If my rewards are negative I want to downweight those probabilities. Right? This is this is the policy gradient theorem. This is REINFORCE. You know if you've taken RL class or something like it you've definitely seen this before. Now what is PPO? PPO I think is normally a very intuit sorry intimidating object but I think it is actually quite simple. So there's two steps that happen. First, instead of taking a reward, we look at what's called an advantage. An advantage you know, just to sort of gloss over a lot of the details involved an advantage is it's basically a variance reduced version of the reward. You know, if you go through the math, you can notice that I can subtract any constant or in fact, I can subtract any sort of state dependent variable from R and this gradient will still be correct. That means that I can rewrite this reward potentially as sort of after subtracting any baseline values that I want. And let's say we call that the advantage. Now, not only that, maybe
69:55 I want to take multiple gradient steps after sampling from P theta once. What's called essentially you know sampling from one roll out and going almost off policy. To enable that, I have to essentially have importance waiting corrections because the more steps I take, sort of the more stale my original samples become. And so this is what's called TRPO. You basically make corrections for all the gradient steps you take and then you constrain yourself to stay close. Now PPO takes the final step and says instead of explicitly constraining myself to stay close to my old policy using this KL constraint, maybe what I can do is I can just clip the probability ratios and this will naturally incentivize the model to stay close to the original policy. So this is kind of PPO in one slide. I'm not going to go into this with too much more detail because actually this won't be the primary algorithm that I want to sort of just go through the rest of the lecture with. So at least in kind of the open research space and the academic space a lot of the question that I think people were concerned with was can
PPO in one slide, and the DPO derivation
71:00 We get rid of PPO? We will see this theme on Thursday as well. PPO is very complicated. And we have debated but decided against having you implement PPO because it will be suffering. And so lots of people thought can we get rid of PPO and they tried other reasonable things. And I will explain those reasonable things like you know maybe we can SFT on the pairs but for each of the pairs we can prepend like a good token for the chosen good outputs and bad token to the chosen not the bad outputs and then I can just condition on good when I generate that does not work very well. I can train the model only on the preferred output that also does not work super well. I can use a reward model, sample the best one out of those and then train on those. Works okay, but maybe not that great. And so people tried all these variants, but what really stuck was basically DPO. And I think the reason why it caught on as much as it did was because it, you know, removed a lot of the complexity of PPO and worked relatively well. And so you get rid of the reward model that exists in PPO. This is used to calculate the advantage. We get rid of any of the on policy stuff like the importance ratio thing that I was talking about. You just get rid of all of those. Instead, you know, we go back to the basics. We take gradient steps on the log loss of good things. And then we take negative gradient steps on log losses of bad stuff, right? We go back to very simple basic things. And the last part of what I want to talk about today is just deriving the DPO formula, right? So, what is our goal? Our goal is to optimize this quantity at the very top. This is just a
72:39 Rewriting of that InstructGPT equation. I have a reward at the very front and then I have a KL divergence like this keeps me this is pi theta close to my reference. Right? So this is this is very natural. So the first thing I'm going to do is I'm going to assume that my policy pi theta is not actually a neural network. I'm going to assume it's an arbitrary function of any kind. And if I do that then I can sort of write down essentially what the optimal policy looks like. It just look it has this form. It's the exponential of the reward and it's multiplying pyref the reference distribution over here. We can solve for the implied reward by solving for rxy. And the clever part about DPO is now to say okay it basically means that every policy instead of thinking about policies I can think about rewards because the two are one and the same under this nonparametric assumption and so what you do is remember I have these two pieces the left side this is the Bradley–Terry equation from the Stiennon paper the one before InstructGPT and on the right side this is the DPO sort of equivalence I wrote down and then now what I can do is I can plug in these rewards R into this objective and I can minimize this loss. I can say what I want to do is now I want to find a policy such that the implied reward for that policy has the highest probability of generating my pairwise comparisons. Right? So now I've taken an RL problem and I have turned it into a maximum likelihood problem. A problem that's very much similar conceptually to something like pre-training. Right? All we're doing is maximizing the
74:18 Probabilities except what we're doing here is we're maximizing the probabilities of the pairwise comparisons. Right? So those are the key steps. Start by making the nonparametric assumptions. Parameterize the reward via the policy and then optimize it using the supervised losses. I think we're a few minutes over so we'll stop here. I think this is a good place because we got through the derivation of DPO and we'll get through the rest of RLHF at the start of next lecture. Thanks everyone for asking lots of good questions.