Overview and tokenization
Staff, ethos, and the case for building from scratch
00:05 Welcome everyone. This is CS 336, Language Models from Scratch, and this is our core staff. So I'm Percy, one of your instructors. I'm really excited about this class because it really allows you to see the whole language modeling building pipeline end to end, including data, systems and modeling. Tatsu, I'll be co-teaching with him. So I'll let everyone introduce themselves. Hi everyone. I'm Tatsu. I'm one of the co-instructors. I'll be giving lectures in a week or two,
00:37 probably a few weeks. I'm really excited about this class. Percy and I spent a while being a little disgruntled, thinking, what's the really deep technical stuff that we can teach our students today? And I think one of the things is that you've got to build it from scratch to understand it. So I'm hoping that that's sort of the ethos that I take away from that class.
01:13 Hey everyone, I'm Rohith. I actually failed this class when I took it, but now I'm your CA. So when they say anything is possible... Hey everyone, I'm Neil. I'm a third year PhD student in the CS department. I work with Tatsu. Mostly my research is on synthetic data, language models, reasoning, all that stuff. Hey guys, I'm Marcel. I'm a second year PhD. These days I work on health. And he was a topper of many leaderboards from last year. So he's the number to beat. Okay. All right. Well, thanks everyone.
01:45 So let's continue. As Tatsu mentioned, this is the second time we're teaching the class. We've grown the class by around 50%. I have three TAs instead of two. And one big thing is we're making all the lectures available on YouTube so that the world can learn how to build language models from scratch. Okay. So why did we decide to make this course and endure all the pain? So let's ask GPT-4. If you ask it why teach a course on building language models from scratch, the reply is: teaching a course provides foundational understanding of techniques, fosters innovation — the typical kind of generic blather.
02:19 Okay, so here's the real reason. We're in a bit of a crisis, I would say. Researchers are becoming more and more disconnected from the underlying technology. Eight years ago, researchers would implement and train their own models in AI. Even six years ago, you would at least take the models like BERT and download them and fine-tune them. And now many people can just get away with prompting a proprietary model. So this is not necessarily bad, right? Because as you introduce layers of abstraction, we can all do more. And a lot of research has been unlocked by the simplicity of being able to prompt the language model, and I do my fair share of prompting. So there's nothing wrong with that.
02:54 But it's also worth remembering that these abstractions are leaky. So in contrast to programming languages or operating systems, you don't really understand what the abstraction is. It's a string in and string out, I guess. And I would say that there's still a lot of fundamental research to be done that requires tearing up the stack and co-designing different aspects of the data and the systems and the model. And I think really that full understanding of this technology is necessary for fundamental research.
03:27 So that's why this class exists. We want to enable the fundamental research to continue, and our philosophy is: to understand it, you have to build it. So there's one small problem here, and this is because of the industrialization of language models.
Industrialization: frontier scale is out of reach, and small models mislead
04:02 So GPT-4 is rumored to be 1.8 trillion parameters, cost 100 million dollars to train. You have xAI building clusters with 200,000 H100s, if you can imagine that. There's an investment of over 500 billion dollars, supposedly, over four years. So these are pretty large numbers, right? And furthermore, there's no public details on how these models are being built.
04:35 Here from GPT-4, this is even two years ago, they very honestly say that due to the competitive landscape and safety implications, we're going to disclose no details. Okay, so this is the state of the world right now. And so in some sense, frontier models are out of reach for us. So if you came into this class thinking you're each going to train your own GPT-4 — sorry — so we're going to build small language models.
05:08 But the problem is that these might not be representative. And here's two examples to illustrate why. So here's a simple one. If you look at the fraction of FLOPs spent in the attention layers of a Transformer versus the MLP, this changes quite a bit. This is a tweet from Stephen Roller from quite a few years ago, but this is still true. If you look at small models, it looks like the number of FLOPs in the attention versus the MLP layers are roughly comparable.
05:41 But if you go up to 175 billion, then the MLPs really dominate, right? So why does this matter? Well, if you spend a lot of time at small scale and you're optimizing the attention, you might be optimizing the wrong thing, because at larger scale it just gets washed out. This is kind of a simple example because you can literally make this plot without any compute. It's just napkin math.
06:14 Here's something that's a little bit harder to grapple with, and that's just emergent behavior. So this is a paper from Jason Wei from 2022, and this plot shows that as you increase the amount of training FLOPs and you look at accuracy on a bunch of tasks, you'll see that for a while it looks like nothing is happening, and all of a sudden you get these emergent phenomena, like in-context learning.
Mechanics, mindset, intuitions — and the bitter lesson restated
06:47 So if you were hanging around at this scale, you would be concluding that these language models really don't work, when in fact you had to scale up to get that behavior. So don't despair — we can still learn something in this class, but we have to be very precise about what we're learning. So there's three types of knowledge. There's the mechanics of how things work. This we can teach you. We can teach you what a Transformer is. You'll implement a Transformer. We can teach you how model parallelism leverages GPUs efficiently.
07:19 These are just the raw ingredients, the mechanics. So that's fine. We can also teach you mindset. So this is something a bit more subtle and seems a little bit fuzzy, but this is actually in some ways more important, I would say, because the mindset that we're going to take is that we want to squeeze as much out of the hardware as possible and take scaling seriously, right? Because in some sense the mechanics — we'll see later that all of these ingredients have been around for a while, but it was really, I think, the scaling mindset that OpenAI pioneered that led to this next generation of AI models.
07:53 So mindset, I think hopefully we can bang into you, to think in a certain way. And then thirdly is intuitions, and this is about which data and modeling decisions lead to good models. This unfortunately we can only partially teach you, and this is because what architectures and what datasets work at small scales might not be the same ones that work at large scales. But hopefully you got two and a half out of three. So that's pretty good bang for your buck.
08:26 Speaking of intuitions, there's this sort of sad reality of things, that you can tell a lot of stories about why certain things in the Transformer are the way they are, but sometimes you just do the experiments and the experiments speak.
08:59 So for example, there's this Noam Shazeer paper that introduced SwiGLU, which is something that we'll see a bit more of in this class, which is a type of nonlinearity. And in the conclusion, the results are quite good and this got adopted. But in the conclusion there's this honest statement that we offer no explanation, except for this is divine benevolence. So there you go. This is the extent of our understanding.
09:33 Okay, so now let's talk about this bitter lesson that I'm sure people have heard about. I think there's a sort of a misconception that the bitter lesson means that scale is all that matters, algorithms don't matter, all you do is pump more capital into building the model and you're good to go. I think this couldn't be further from the truth. I think the right interpretation is that algorithms at scale is what matters. Because at the end of the day, the accuracy of your model is really a product of your efficiency and the number of resources you put in.
10:10 And actually efficiency, if you think about it, is way more important at larger scale. Because if you're spending hundreds of millions of dollars you cannot afford to be wasteful, in the same way that if you're running a job on your local cluster you might run it again, you fail, you debug it. And if you look at the utilization, I'm sure OpenAI is way more efficient than any of us right now. So efficiency really is important.
10:43 And furthermore, this point is maybe not as well appreciated in the scaling rhetoric, so to speak, which is that if you look at efficiency, which is a combination of hardware and algorithms — but if you just look at the algorithmic efficiency, there's this nice OpenAI paper from 2020 that showed over the period of 2012 to 2019 there's a 44x algorithmic efficiency improvement in the time that it took to train ImageNet to a certain level of accuracy, right?
11:15 So this is huge. And I don't know if you can see the abstract here, but this is faster than Moore's law, right? So algorithms do matter. If you didn't have this efficiency you would be paying 44 times more cost. This is for image models, but there are some results for language as well.
11:49 Okay, so with all that, I think the right framing or mindset to have is: what is the best model one can build given a certain compute and data budget? And this question makes sense no matter what scale you're at, because it's accuracy per resources. And of course if you can raise the capital and get more resources you'll get better models, but as researchers our goal is to improve the efficiency of the algorithms. So, maximize efficiency. We're going to hear a lot of that.
A compressed history, and the three levels of openness
12:24 Okay, so now let me talk a little bit about the current landscape, and a little bit of obligatory history. So language models have been around for a while now, going back to Shannon, who looked at language models as a way to estimate the entropy of English. In AI they really were prominent in NLP, where they were a component of larger systems like machine translation and speech recognition. And one thing that's maybe not as appreciated these days is that if you look back in 2007, Google was training fairly large n-gram models — so five-gram models over two trillion tokens, which is a lot more tokens than GPT-3.
12:56 And it was only recently, I guess in the last two years, that we've gotten to that in token count. But they were n-gram models, so they didn't really exhibit any of the interesting phenomena that we know of language models today. Okay. So in the 2010s, you can think about this as a lot of the deep learning revolution happening and a lot of the ingredients sort of falling into place. So there was the first neural language model from Yoshua Bengio's group back in 2003.
13:28 There were sequence-to-sequence models. This, I think, was a big deal for how do you basically model sequences, from Ilya and the Google folks. There's the Adam optimizer, which is still used by the majority of people, dating over a decade ago. There's the attention mechanism, which was developed in the context of machine translation, which then led up to the famous "Attention Is All You Need," aka the Transformer paper, in 2017.
14:01 People were looking at how to scale mixture of experts. There's a lot of work around the late 2010s on how to essentially do model parallelism, and they were actually figuring out how you could train 100 billion parameter models. They didn't train them for very long because these were more systems work, but all the ingredients were kind of in place by the time 2020 came around.
14:37 So one other trend which was starting in NLP was the idea of these foundation models that could be trained on a lot of text and adapted to a wide range of downstream tasks. So ELMo, BERT, T5 — these were models that were, for their time, very exciting. We kind of maybe forget how excited people were about things like BERT, but it was a big deal.
15:10 And then — I mean, this is an abbreviated history — but I think one critical piece of the puzzle is OpenAI taking these ingredients and applying very nice engineering and really kind of pushing on the scaling laws, embracing it as the mindset piece, and that led to GPT-2 and GPT-3. Google obviously was in the game and trying to compete as well.
15:46 But that sort of paved the way, I think, to another line of work, which is — these were all closed models. So models that weren't released and you can only access via API. But there were also open models, starting with early work by EleutherAI right after GPT-3 came out, Meta's early attempt which didn't work maybe as quite as well, BLOOM, and then Meta, Alibaba, DeepSeek, AI2 and a few others which I have listed have been creating these open models where the weights are released.
16:22 One other tidbit about openness I think is important is that there's many levels of openness. There's closed models like GPT-4. There's open-weight models, where the weights are available and there's actually a very nice paper with lots of architectural details but no details about the dataset. And then there's open-source models where all the weights and data are available, and the paper is where they're honestly trying to explain as much as they can.
16:56 But of course you can't really capture everything in a paper, and there's no substitute for learning how to build it except for doing it yourself. Okay. So that leads to the present day, where there's a whole host of frontier models from OpenAI, Anthropic, xAI, Google, Meta, DeepSeek, Alibaba, Tencent and probably a few others, that dominate the current landscape.
17:30 So we're at an interesting time where — just to reflect — a lot of the ingredients, like I said, were developed, which is good, because I think we're going to revisit some of those ingredients and trace how these techniques work. And then we're going to try to move as close as we can to best practices on frontier models, but using information from essentially the open community and reading between the lines from what we know about the closed models.
Executable lectures, and the course logistics
18:07 Okay, so just as an interlude: what are you looking at here? So this is an executable lecture. It's a program where I'm stepping through and it delivers the content of the lecture. So one thing that I think is interesting here is that you can embed code. You can just step through code, and — I think this is a smaller screen than I'm used to — but you can look at the environment variables as you're stepping through code. So that's useful later when we start actually trying to drill down and give code examples.
18:41 You can see the hierarchical structure of the lecture, like we're in this module and you can see where it was called from main, and you can jump to definitions like supervised fine-tuning, which we'll talk about later. And if you think this looks like a Python program — well, it is a Python program. But I've post-processed it for your viewing pleasure. Okay, so let's move on to the course logistics now.
19:18 Actually, maybe I'll pause for questions. Any questions about what we're learning in this class? Yeah. Would you expect a graduate from this class to be able to lead a team to build a frontier model? So the question is, would I expect a graduate from this class to be able to lead a team and build a frontier model? Of course, with like a billion dollars of capital.
19:52 Yeah, of course. I would say that it's a good step, but there's definitely many pieces that are missing. And I think we thought about — we should really teach a series of classes that eventually leads up to as close as we can get. But I think this is maybe the first step of the puzzle, and there are a lot of things — happy to talk offline about that. But I like the ambition. Yeah, that's what you should be doing, taking the class so you can go lead teams and build frontier models.
20:26 Okay, let's talk a little bit about the course. So here's the website. Everything's online. This is a five-unit class. But I think that maybe doesn't express the level here as well as this quote that I pulled out from a course evaluation: "The entire assignment was approximately the same amount of work as all five assignments from CS 224N plus the final project." And that's the first homework assignment. So, not to scare you off, but just giving some data here.
20:59 So why should you endure that? Why should you do it? I think this class is really for people who have this obsessive need to understand how things work all the way down to the atoms, so to speak. And I think when you get through this class, you will have really leveled up in terms of your research engineering, and the level of comfort that you'll have in building ML systems at scale will just be something.
21:33 There's also a bunch of reasons that you shouldn't take the class. For example, if you want to get any research done this quarter, maybe this class isn't for you. If you're interested in learning just about the hottest new techniques, there are many other classes that can probably deliver on that better than, for example, you spending a lot of time debugging BPE. And this is really a class about the primitives and learning things bottom-up, as opposed to the latest.
22:07 And also, if you're interested in building language models for your own application domain, this is probably not the first class you would take. I think practically speaking, as much as I kind of made fun of prompting, prompting is great, fine-tuning is great. If you can do that and it works, then I think that is something you should absolutely start with. So I don't want people taking this class and thinking, great, any problem, the first step is to train a language model from scratch. That is not the right way of thinking about it.
22:41 Okay. And I know that some of you were enrolled but we did have a cap so we weren't able to enroll everyone. And also for the people online, you can follow along at home — all the lecture materials and assignments are online, so you can look at them. The lectures are also recorded and will be put on YouTube, although there will be some number of weeks lag there.
23:15 And also we'll offer this class next year. So if you were not able to take it this year, don't fret, there will be a next time. Okay. So the class has five assignments. And for each of the assignments we don't provide scaffolding code, in the sense that we literally give you a blank file and you're supposed to build things up, in the spirit of learning and building from scratch. But we're not that mean. We do provide unit tests and some adapter interfaces that allow you to check correctness of different pieces.
23:48 And also the assignment write-up, if you walk through it, does a fairly gentle job of doing that. But you're kind of on your own for making good software design decisions and figuring out what you name your functions and how to organize your code, which is a useful skill, I think. So one strategy for all assignments is that there is a piece of the assignment which is just implement the thing and make sure it's correct.
24:23 That mostly you can do locally on your laptop. You shouldn't need compute for that. And then we have a cluster that you can run on for benchmarking, both accuracy and speed. Right? So I want everyone to embrace this idea that you want to use as small a dataset or as few resources as possible to prototype before running large jobs. You shouldn't be debugging with one-billion-parameter models on the cluster if you can help it.
24:58 There are some assignments which will have a leaderboard, which usually is of the form: do things to make perplexity go down given a particular training budget. Last year it was pretty exciting for people to try different things that you either learn from the class or you read online. And then finally, I guess this year — this was less of a problem last year because I guess Copilot wasn't as good, but Cursor is pretty good.
25:30 So I think our general strategy is that AI tools can take away from learning, because there are cases where it can just solve the thing you want it to do. But I think you can obviously use them judiciously. So use at your own risk — you're kind of responsible for your own learning experience here. Okay. So we do have a cluster. Thank you Together AI for providing a bunch of H100s for us.
26:05 There's a guide — please read it carefully to learn how to use the cluster. And start your assignments early, because the cluster will fill up towards the end of a deadline as everyone's trying to get their large runs in. Okay. Any questions about that? You mentioned it was five units. Can you sign up for less? Right, so the question is, can you sign up for less than five units? I think administratively, if you have to sign up for less, that is possible, but it's the same class and the same workload. Any other questions?
Unit 1 · Basics: tokenizer, Transformer, training loop
26:39 Okay. So in this part I'm going to go through all the different components of the course and just give a broad overview, a preview of what you're going to experience. So remember, it's all about efficiency given hardware and data. How do you train the best model given your resources? So for example, if I give you a Common Crawl dump, a web dump, and 32 H100s for two weeks, what should you do?
27:12 There are a lot of different design decisions. There's questions about the tokenizer, the architecture, systems optimizations you can do, data things you can do, and we've organized the class into these five units or pillars. So I'm going to go through each of them in turn and talk about what we'll cover, what the assignment will involve, and then I'll wrap up.
27:47 Okay. So the goal of the basics unit is just get a basic version of a full pipeline working. So here you implement a tokenizer, model architecture and training. So just to say a bit more about what these components are. A tokenizer is something that converts between strings and sequences of integers. Intuitively you can think about the integers corresponding to breaking up the string into segments and mapping each segment to an integer. And the idea is that your sequence of integers is what goes into the actual model, which has to be a fixed dimension.
28:19 Okay. So in this course we'll talk about the byte-pair encoding, BPE, tokenizer, which is relatively simple and still is used. There are, I guess, a promising set of methods on tokenizer-free approaches. So these are methods that just start with the raw bytes and don't do tokenization, and develop a particular architecture that just takes the raw bytes.
28:53 This work is promising but so far I haven't seen it be scaled to the frontier yet. So we'll go with BPE for now. Okay. So once you've tokenized your strings into a sequence of integers, now we define a model architecture over these sequences. So the starting point here is the original Transformer. That's the backbone of basically all frontier models.
29:29 And here's the architectural diagram. We won't go into details here, but there's an attention piece and then there's an MLP layer with some normalization. So a lot has actually happened since 2017, right? I think there's a sense to which, oh, the Transformer was invented and then everyone's just using the Transformer. And to first approximation that's true — we're still using the same recipe — but there have been a bunch of smaller improvements that do make a substantial difference when you add them all up.
30:03 So for example, there is the nonlinear activation function, SwiGLU, which we saw a little bit before. Positional embeddings: there are new positional embeddings, these rotary positional embeddings, RoPE, which we'll talk about. Normalization: instead of using LayerNorm we're going to look at something called RMSNorm, which is similar but simpler. There's a question of where you place the normalization, which has been changed from the original Transformer.
30:37 The MLP — the canonical version is a dense MLP, and you can replace that with mixture of experts. Attention is something that has actually been getting a lot of attention, I guess. There's full attention and then there's sliding window attention and linear attention. All of these are trying to prevent the quadratic blow-up. There's also lower-dimensional versions like GQA and MLA, which we'll get to in a future lecture.
31:09 And then the most radical thing is other alternatives to the Transformer, like state-space models, like Hyena, where they're not doing attention but some other sort of operation. And sometimes you get best of both worlds by making a hybrid model that mixes these in with Transformers. Okay, so once you define your architecture you need to train.
31:41 So there are design decisions including the optimizer. AdamW, which is basically a variant of Adam fixed up, is still very prominent. So we'll mostly work with that, but it is worth mentioning that there are more recent optimizers like Muon and SOAP that have shown promise. Learning rate schedule, batch size, whether you do regularization or not, hyperparameters — there's a lot of details here.
32:15 And I think this class is one where the details do matter, because you can easily have an order of magnitude difference between a well-tuned architecture and something that's just a vanilla Transformer. So in Assignment 1, basically you'll implement the BPE tokenizer. I'll warn you that this seems to have been a lot of surprising work for people. So, you're warned. And you also implement the Transformer, cross-entropy loss, AdamW optimizer and training loop.
32:49 So again, the whole stack. And we're not making you implement PyTorch from scratch, so you can use PyTorch, but you can't use the Transformer implementation from PyTorch. There's a small list of functions that you can use and you can only use those. Okay, so we're going to have TinyStories and OpenWebText datasets that you'll train on, and then there will be a leaderboard to minimize the OpenWebText perplexity. We'll give you 90 minutes on an H100 and see what you can do.
Unit 2 · Systems: kernels, parallelism, inference
33:23 So this is last year's, so we have the top. This is the number to beat for this year. Okay. All right. So that's the basics. Now after basics, in some sense you're done, right? Like, you have the ability to train a Transformer. What else do you need? So the systems part really goes into how you can optimize this further. So how do you get the most out of hardware?
33:57 And for this we need to take a closer look at the hardware and how we can leverage it. So there are kernels, parallelism and inference — the three components of this unit. So okay, first to talk about kernels: let's talk a little bit about what a GPU looks like. So a GPU, which we'll get much more into, is basically a huge array of these little units that do floating-point operations.
34:33 And maybe the one thing to note is that this is the GPU chip and here is the memory that's actually off-chip. And then there's some other memory like L2 caches and L1 caches on chip. And so the basic idea is that compute has to happen here, your data might be somewhere else, and how do you organize your compute so that you can be most efficient?
35:10 So one quick analogy is: imagine that your memory, where you can store your data and the model parameters, is like a warehouse, and your compute is like the factory. And what ends up being a big bottleneck is just data movement costs, right? So the thing that we have to do is, how do you organize the compute — even a matrix multiplication — to maximize the utilization of the GPUs by minimizing the data movement?
35:43 And there are a bunch of techniques like fusion and tiling that allow you to do that. So we'll get into all the details of that, and to implement and leverage a kernel we're going to look at Triton. There are other things you can do with various levels of sophistication, but we're going to use Triton, which was developed by OpenAI and is a popular way to build kernels.
36:17 Okay, so we're going to write some kernels. That's for one GPU. So now, in general, these big runs take ten thousand if not tens of thousands of GPUs. But even at 8 it kind of starts becoming interesting, because you have a lot of GPUs, they're connected to some CPU nodes, and they are also directly connected via NVSwitch and NVLink.
36:51 And it's the same idea, right? Now the only thing is that data movement between GPUs is even slower. And so we need to figure out how to put model parameters and activations and gradients on the GPUs and do the computation, and minimize the amount of movement. And then we're going to explore different types of techniques like data parallelism and tensor parallelism and so on.
37:26 So that's all I'll say about that. And finally, inference is something that we didn't actually do last year in the class — although we had a guest lecture — but this is important because inference is how you actually use a model, right? It's basically the task of generating tokens given a prompt, given a trained model. And it also turns out to be really useful for a bunch of other things besides just chatting with your favorite model: you need it for reinforcement learning, test-time compute, which has been very popular lately, and even evaluating models — you need inference.
38:02 So we're going to spend some time talking about inference. Actually, if you think about globally the cost that's dedicated to inference, it's eclipsing the cost that is used to train models. Because training, despite it being very intensive, is ultimately a one-time cost, and inference cost scales with every use — and the more people use your model, the more you'll need inference to be efficient.
38:34 Okay. So in inference there are two phases. There's prefill and decode. Prefill is: you take the prompt and you run it through the model and get some activations. And then decode is: you go autoregressively, one by one, and generate tokens. So in prefill all the tokens are given, so you can process everything at once. So this is exactly what you see at training time, and generally this is a good setting to be in because it's naturally parallel and you're mostly compute-bound.
39:06 What makes inference special and difficult is that in this autoregressive decoding you need to generate one token at a time, and it's hard to actually saturate all your GPUs, and it becomes memory-bound because you're constantly moving data around. And we'll talk about a few ways to speed inference up. You can use a cheaper model.
39:39 You can use this really cool technique called speculative decoding, where you use a cheaper model to sort of scout ahead and generate multiple tokens, and then if these tokens happen to be good by some definition of good, you can have the full model just score and accept them all in parallel. And then there are a bunch of systems optimizations that you can do as well.
40:13 Okay, so after the systems — okay, Assignment 2. So you're going to implement a kernel, you're going to implement some parallelism. Data parallel is very natural, so we'll do that. Some of the model parallelism like FSDP turns out to be a bit complicated to do from scratch, so we'll do sort of a baby version of that. But I encourage you to learn about the full version — we'll go over the full version in class, but implementing from scratch might be a bit too much.
Unit 3 · Scaling laws: D* ≈ 20N*
40:45 And then I think an important thing is getting in the habit of always benchmarking and profiling. I think that's actually probably the most important thing — that you can implement things, but unless you have feedback on how well your implementation is going and where the bottlenecks are, you're just going to be flying blind. Okay, so unit three is scaling laws. And here the goal is you want to do experiments at small scale and figure things out, and then predict the hyperparameters and loss at large scale.
41:17 So here's a fundamental question. If I give you a FLOPs budget, what model size should you use? If you use a larger model, that means you can train on less data, and if you use a smaller model, you can train on more data. So what's the right balance here? And this has been studied quite extensively and figured out by a series of papers from OpenAI and DeepMind.
41:51 So if you hear the term "Chinchilla optimal," this is what this is referring to. And the basic idea is that for every compute budget, number of FLOPs, you can vary the number of parameters of your model, and then you measure how good that model is. So for every level of compute you can get the optimal parameter count, and then what you do is you fit a curve to extrapolate.
42:24 And see, if you had let's say 1e22 FLOPs, what would be the parameter size? And it turns out these minima, when you plot them, are actually remarkably linear, which leads to this very simple but useful rule of thumb, which is that if you have a particular model of size N, if you multiply by 20 that's the number of tokens you should train on, essentially. So that means a 1.4 billion parameter model should be trained on 28 billion tokens.
42:57 Okay, but this doesn't take into account inference cost. This is literally how can you train the best model regardless of how big that model is. So there are some limitations here, but it's nonetheless been extremely useful for model development. So in this assignment — this is kind of fun because we define a quote-unquote training API which you can query with a particular set of hyperparameters. You specify the architecture and batch size and so on, and we return you a loss that your decisions will get you.
43:31 Okay. So your job is: you have a FLOPs budget and you're going to try to figure out how to train a bunch of models and then gather the data. You're going to fit a scaling law to the gathered data and then you're going to submit your prediction on what you would choose to be the hyperparameters, what model size and so on, at a larger scale.
44:04 Okay. So this is a case where we want to put you in this position where there are some stakes. I mean, this is not like burning real compute, but once you run out of your FLOPs budget, that's it. So you have to be very careful in terms of how you prioritize what experiments to run, which is something that the frontier labs have to do all the time. And there will be a leaderboard for this, which is minimize loss given your FLOPs budget.
Unit 4 · Data: it does not fall from the sky
44:39 Question about the 2024 links. If we're working ahead, should we expect assignments to change over time, or are these going to be the final assignments? So the question is that these links are from 2024. The rough structure will be the same for 2025. There will be some modifications, but if you look at these, you should have a pretty good idea of what to expect. Okay, so let's go into data now.
45:13 Okay, so up until now you've done scaling laws, you have systems, you have your Transformer implementation, everything — you're really kind of good to go. But data, I would say, is a really key ingredient that differentiates in some sense. And the question to ask here is: what do I want this model to do? Because what the model does is mostly determined by the data. If I train on multilingual data, it will have multilingual capabilities. If I train on code, it'll have code capabilities. It's very natural.
45:45 And usually datasets are a conglomeration of a lot of different pieces. This is from The Pile, which is four years ago, but the same idea holds. You have data from the web — this is Common Crawl — you have Stack Exchange, Wikipedia, GitHub and different sources which are curated. And so in the data section we're going to start talking about evaluation, which is: given a model, how do you evaluate whether it's any good?
46:19 So we're going to talk about perplexity, standardized testing like MMLU. If you have models that generate utterances for instruction following, how do you evaluate that? There are also decisions about whether you ensemble or do chain of thought at test time, and how does that affect your evaluation. And then you can talk about evaluation of entire systems, not just a language model, because language models often get plugged into some agentic system these days.
46:53 Okay, so now after establishing evaluation, let's look at data curation. So this is an important point that people don't realize. I often hear people say, oh, we're training the model on the internet. This just doesn't make sense, right? Data doesn't just fall from the sky and there's the internet that you can pipe into your model. Data has to always be actively acquired somehow. So even if you — just as an example — I always tell people, look at the data.
47:27 And so let's look at some data. So this is some Common Crawl data. I'm going to take 10 documents, and hopefully this works. I think the rendering is off, but you can see this is a random sample of Common Crawl. And you can see that this is maybe not exactly the data — oh, here's some actually real text here. Okay, that's cool.
48:00 But if you look at most of Common Crawl — aside from, this is a different language — but you can also see this is very spammy sites, and you'll quickly realize that a lot of the web is just trash. And, okay, maybe that's not surprising, but it's more trash than you would actually expect, I promise. So what I'm saying is that there's a lot of work that needs to happen in data. So you can crawl the internet, you can take books, archives, papers, GitHub.
48:34 And there's actually a lot of processing that needs to happen. There are also legal questions about what data you can train on, which we'll touch on. Nowadays, a lot of frontier models have to actually buy data, because the data on the internet that's publicly accessible turns out to be a bit limited for the really frontier performance.
49:06 And also I think it's important to remember that this data that's scraped, it's not actually text, right? First of all, it's HTML, or it's PDFs, or in the case of code it's just directories. So there has to be an explicit process that takes this data and turns it into text. Okay, so we're going to talk about the transformation from HTML to text. And this is going to be a lossy process. So the trick is, how can you preserve the content and some of the structure without basically just having HTML?
49:41 Filtering, as you could surmise, is going to be very important, both for getting high quality data but also removing harmful content. Generally people train classifiers to do this. Deduplication is also an important step which we'll talk about. Okay. So Assignment 4 is all about data. We're going to give you the raw Common Crawl dump so you can see just how bad it is. And you're going to train classifiers, dedup, and then there's going to be a leaderboard where you're going to try to minimize perplexity given your token budget.
Unit 5 · Alignment: SFT, then learning from feedback
50:14 So now you have the data, you've done this, built all your fancy kernels, you've trained — now you can really train models. But at this point what you'll get is a model that can complete the next token, right? And this is called a base model, and I think about it as a model that has a lot of raw potential but it needs to be aligned or modified some way. And alignment is a process of making it useful.
50:47 So alignment captures a lot of different things, but three things I think it captures: you want to get the language model to follow instructions, right? Completing the next token is not necessarily following the instruction — it'll just complete the instruction, or whatever it thinks will follow the instruction. You get to specify the style of the generation, whether you want it to be long or short, whether you want bullets, whether you want it to be witty or have sass or not.
51:21 And when you play with ChatGPT versus Grok, you'll see that there's different alignment that has happened. And then also safety. One important thing is for these models to be able to refuse answers that can be harmful. So that's where alignment also kicks in. So there are generally two phases of alignment. There's supervised fine-tuning, and here the goal is — I mean, it's very simple — you basically gather a set of user–assistant pairs, so prompt–response pairs, and then you do supervised learning.
51:55 Okay. And the idea here is that the base model already has the raw potential. So just fine-tuning it on a few examples is sufficient. Of course, the more examples you have, the better the results, but there are papers like this one that show even a thousand examples suffices to give you instruction-following capabilities from a good base model.
52:27 Okay, so this part is actually very simple, and it's not that different from pre-training, because you're just given text and you just maximize the probability of the text. So the second part is a bit more interesting from an algorithmic perspective. So the idea here is that even with the SFT phase you will have a decent model. And now how do you improve it? What you can get there is more SFT data, but that can be very expensive because you have to have someone sit down and annotate data.
52:58 So there, the goal of learning from feedback is that you can leverage lighter forms of annotation and have the algorithms do a bit more work. Okay. So one type of data you can learn from is preference data. So this is where you generate multiple responses from a model to a given prompt, like A or B, and the user rates whether A or B is better.
53:33 And so the data might look like: what's the best way to train a language model, use a large dataset or use a small dataset — and of course the answer should be A. So that is a unit of expressing preferences. Another type of supervision you could have is using verifiers. So for some domains, you're lucky enough to have a formal verifier, like for math or code. Or you can use learned verifiers, where you train an actual language model to rate the response.
54:08 And of course this relates to evaluation. Again, algorithms. This is where we're in the realm of reinforcement learning. So one of the earliest algorithms that was developed that was applied to instruction tuning models was PPO, proximal policy optimization. It turns out that if you just have preference data, there's a much simpler algorithm called DPO that works really well.
54:41 But in general, if you want to learn from verifier data — it's not preference data — so you have to embrace RL fully. And there's this method which we'll do in this class, called Group Relative Preference Optimization, which simplifies PPO, makes it more efficient by removing the value function, developed by DeepSeek, which seems to work pretty well. Okay, so Assignment 5 implements supervised fine-tuning, DPO and GRPO, and of course evaluate.
55:20 Question about Assignment 1. Do people have similar things to say about assignments two or... Yeah, the question is, Assignment 1 seems a bit daunting — what about the other ones? I would say that Assignment 1 and 2 are definitely the most heavy and hardest. Assignment 3 is a bit more of a breather, and assignments 4 and 5, at least last year, were, I would say, a notch below Assignment 1 or 2. Although I don't know, it depends — we haven't fully worked out the details for this year.
Efficiency as the through-line, and what changes when the regime does
55:55 Yeah, it does get better. Okay, so just a recap of the different pieces here. Remember, efficiency is this driving principle, and there's a bunch of different design decisions. And I think if you view everything through a lens of efficiency, a lot of things kind of make sense.
56:28 And importantly, it's worth pointing out, we are currently in this compute-constrained regime — at least this class, and most people who are somewhat GPU-poor. So we have a lot of data but we don't have that much compute, and so these design decisions will reflect squeezing the most out of the hardware. So for example, data processing: we're filtering fairly aggressively because we don't want to waste precious compute on bad or irrelevant data. Tokenization: it's nice to have a model over bytes, that's very elegant, but it's very compute-inefficient with today's model architectures.
57:00 So we have to do tokenization as an efficiency gain. Model architecture: there are a lot of design decisions there that are essentially motivated by efficiency. Training: the fact that most of what we're going to do is just a single epoch — this is clearly, we're in a hurry. We just need to see more data as opposed to spend a lot of time on any given data point. Scaling laws is completely about efficiency: we use less compute to figure out the hyperparameters.
57:32 And alignment is maybe a little bit different, but the connection to efficiency is that if you can put resources into alignment then you actually require smaller base models. Okay. So there are sort of two paths. If your use case is fairly narrow, you can probably use a smaller model, you align it or fine-tune it and you can do well. But if your use cases are very broad, then there might not be a substitute for training a big model. So that's today.
58:04 So increasingly now, at least for frontier labs, they're becoming data-constrained, which is interesting because I think the design decisions will presumably completely change. Well, I mean, compute will always be important, but I think the design decisions will change. For example, taking one epoch of your data I think doesn't really make sense if you have more compute — why wouldn't you take more epochs, at least, or do something smarter?
58:36 Or maybe there will be different architectures, for example, because the Transformer was really motivated by compute efficiency. So that's something to ponder. Still, it's about efficiency, but the design decisions reflect what regime you're in. Okay, so now I'm going to dive into the first unit. Before that — any questions?
59:15 Do you have a Slack? The question is, do we have a Slack? We will have a Slack. We'll send out details after this class. Yeah. Will students auditing the course also have access to the same materials? The question is, will students auditing the class have access to all the online materials and assignments — and we'll give you access to Canvas so you can watch the lecture videos.
59:48 Yeah. What's the grading of the assignments? What's the grading of the assignments? Good question. So there will be a set of unit tests that you will have to pass. So part of the grading is just, did you implement this correctly. There will be also parts of the grade which will be, did you implement a model that achieved a certain level of loss or is efficient enough. In the assignment, every problem part has a number of points associated with it, and so that gives you a fairly granular level of what grading looks like.
Tokenization: the interface, and the compression ratio
60:23 Okay, let's jump into tokenization. So Andrej Karpathy has this really nice video on tokenization, and in general he makes a lot of these videos that actually inspired a lot of this class, on how you can build things from scratch. So you should go check out some of his videos. So tokenization, as we talked about it, is the process of taking raw text, which is generally represented as Unicode strings, and turning it into a set of integers essentially, where each integer represents a token.
60:56 Okay. So we need a procedure that encodes strings to tokens and decodes them back into strings. And the vocabulary size is just the number of values that a token can take on — the range of the integers. Okay, so just to give you an example of how tokenizers work, let's play around with this really nice website which allows you to look at different tokenizers.
61:30 And just type in something like "hello" or whatever. And one thing it does is it shows you the list of integers. This is the output of the tokenizer. It also nicely maps out the decomposition of the original string into a bunch of segments. And a few things to note. First of all, the space is part of a token. So unlike classical NLP, where the space just kind of disappears, everything is accounted for.
62:03 These are meant to be reversible operations, tokenization. And by convention, for whatever reason, the space is usually preceding the token. Also notice that "hello" is a completely different token than " hello", which might make you a little bit squeamish, and it can cause problems, but that's just how it is.
62:38 Question: is the space being leading instead of trailing intentional, or is it just an artifact of the BPE process? So the question is, is the space before intentional or not? So in the BPE process, I will talk about — you actually pre-tokenize, and then you tokenize each part, and I think the pre-tokenizer does put the space in the front.
63:11 So it is built into the algorithm. You could put it at the end, but I think it probably makes more sense to put it in the beginning. But actually, I guess it could go either way. That's my sense. Okay, so then if you look at numbers, you see that the numbers are chopped down into different pieces. It's a little bit interesting that it's left to right. So it's definitely not grouping by thousands or anything semantic.
63:43 But anyway, I encourage you to play with it and get a sense of what these existing tokenizers look like. So this is a tokenizer for GPT-4o, for example. So there are some observations that we made. So if you look at the GPT-2 tokenizer, which we'll use as a reference — okay, let me see if I can — okay, hopefully this is — let me know if this is getting too small in the back.
64:17 You could take a string, and if you apply the GPT-2 tokenizer, you get your indices. So it maps strings to indices, and then you can decode to get back the string, and this is just a sanity check to make sure that it round-trips. Another thing that's interesting to look at is this compression ratio, which is if you look at the number of bytes divided by the number of tokens.
64:50 So how many bytes are represented by a token — and the answer here is 1.6. Okay, so every token represents 1.6 bytes of data. Okay, so that's just the GPT-2 tokenizer that OpenAI trained.
Characters, bytes, words: three attempts that fail
65:23 To motivate BPE, I want to go through a sequence of attempts. Suppose you wanted to do tokenization. What would be the simplest thing? The simplest thing is probably character-based tokenization. A Unicode string is a sequence of Unicode characters, and each character can be converted into an integer called a code point. Okay, so "a" maps to 97. The world emoji maps to 127,757, and you can see that it converts back.
66:01 Okay. So you can define a tokenizer which simply maps each character into a code point. Okay. So what's one problem with this? Compression ratio is one. The compression ratio is one. Well, actually the compression ratio is not quite one, because a character is not a byte. But it's maybe not as good as you want. One problem with that: if you look at some code points, they're actually really large, right?
66:36 So you're basically allocating one slot in your vocabulary for every character uniformly, and some characters appear way more frequently than others. So this is not a very effective use of your budget. Okay. So the vocabulary size is huge. The bigger problem is that some characters are rare and this is inefficient use of the vocab. Okay, so the compression ratio is 1.5 in this case, because it's the number of bytes per token, and a character can be multiple bytes.
67:11 Okay, so that was a very naive approach. On the other hand, you can do byte-based tokenization. Okay, so Unicode strings can be represented as a sequence of bytes, because every string can just be converted into bytes. So "a" is already just one byte, but some characters take up as many as four bytes, and this is using the UTF-8 encoding of Unicode. There are other encodings but this is the most common one.
67:45 So let's just convert everything into bytes and see what happens. So if you do it into bytes, now all the indices are between 0 and 256, because there are only 256 possible values for a byte by definition. So your vocabulary is very small, and each byte is — I guess not all bytes are equally used, but you don't have that many sparsity problems.
68:19 But what's the problem with byte-based encoding? Long sequences. Yeah, long sequences. So in some ways I really wish byte encoding would work. It's the most elegant thing. But you have long sequences, your compression ratio is one — one byte per token. And this is just terrible. A compression ratio of one is terrible because your sequences will be really long. Attention is quadratic, naively, in the sequence length.
68:51 So you're just going to have a bad time in terms of efficiency. Okay, so that wasn't really good. So now the thing that you might think about is, well, maybe we have to be adaptive here, right? Like, we can't allocate a character or a byte per token, but maybe some tokens can represent lots of bytes and some tokens can represent few bytes. So one way to do this is word-based tokenization, and this is something that was actually very classic in NLP.
69:26 So here's a string, and you can just split it into a sequence of segments, okay, and you can call each of these tokens. So you just use a regular expression. Here's a different regular expression that GPT-2 uses to pre-tokenize, and it just splits your string into a sequence of strings.
69:57 And then what you do with each segment is that you assign each of these to an integer, and then you're done. Okay. So what's the problem with this? Yeah. So the problem is that your vocabulary size is sort of unbounded. Well, maybe not quite unbounded, but you don't know how big it is, right? Because on a given new input, you might get a segment that you've never seen before. And that's actually kind of a big problem.
70:31 This is actually — word-based is a really big pain, because some real words are rare, and it's really annoying because new words have to receive this UNK token. And if you're not careful about how you compute the perplexity, then you're just going to mess up. So word-based, I think it captures the right intuition of adaptivity, but it's not exactly what we want here.
BPE: train the vocabulary from corpus statistics, merge by merge
71:05 So here we're finally going to talk about BPE encoding, byte-pair encoding. So this was actually a very old algorithm, developed by Philip Gage in 1994 for data compression, and it was first introduced into NLP for neural machine translation. So before, papers that did machine translation — or basically all NLP — used word-based tokenization, and again word-based was a pain.
71:38 So this paper pioneered this idea: well, we can use this nice algorithm from 1994 and we can just make the tokenization round-trip and we don't have to deal with UNKs or any of that stuff. And then finally this entered the language modeling era through GPT-2, which was trained using the BPE tokenizer.
72:10 Okay, so the basic idea is instead of defining some preconceived notion of how to split up, we're going to train the tokenizer on raw text. That's the basic insight, if you will. And so organically, common sequences that span multiple characters we're going to try to represent as one token, and rare sequences are going to be represented by multiple tokens. There's a slight detail, which is that for efficiency the GPT-2 paper uses a word-based tokenizer as a preprocessing step to break it up into segments and then runs BPE on each of the segments, which is what you're going to do in this class as well.
72:44 The BPE algorithm is actually very simple. So we first convert the string into a sequence of bytes, which we already did when we talked about byte-based tokenization. And now we're going to successively merge the most common pair of adjacent tokens over and over again. So the intuition is that if there's a pair of tokens that shows up a lot, then we're going to compress it into one token. We're going to dedicate space for that. Okay, so let's walk through what this algorithm looks like. So we're going to use "the cat in the hat" as an example, and we're going to convert this into a sequence of integers.
73:19 These are the bytes. And then we're going to keep track of what we've merged. So remember, merges is a map from two integers, which can represent bytes or other pre-existing tokens, and we're going to create a new token. And the vocab is just going to be a handy way to represent the index to bytes. Okay. So we're going to run the BPE algorithm. I mean, it's very simple. So I'm actually just going to run through the code.
73:52 You're going to do this a number of times — the number is three in this case. We're going to first count up the number of occurrences of pairs of bytes. So, hopefully this doesn't become too small. So we're going to just step through this sequence and we're going to see that — okay, so what's 116, 104? We're going to increment that count. 104, 101, increment that count. We go through the sequence and we're going to count up the bytes.
74:26 Okay, so now after we have these counts, we're going to find the pair that occurs the most number of times. So I guess there are multiple ones, but we're just going to break ties and say 116 and 104. Okay, so that occurred twice. So now we're going to merge that pair. So we're going to create a new slot in our vocab, which is going to be 256. So so far it's 0 through 255, but now we're expanding the vocab to 256.
75:00 And we're going to say every time we see 116 and 104, we're going to replace it with 256. Okay? And then we're going to just apply that merge to our training set. So after we do that, the 116, 104 became 256, and this 256 — remember — occurred twice.
75:32 Okay, so now we're just going to loop through this algorithm one more time. The second time, it decided to merge 256 and 101. And now I'm going to replace that in indices. And notice that the indices are going to shrink, right? Because our compression ratio is getting better as we make room for more vocabulary items and we have a greater vocabulary to represent everything.
76:06 Okay, so let me do this one more time. And then the next merge is 257 and 32. And this is shrinking one more time. Okay. And then now we're done. Okay. So let's try out this tokenizer. So we have the string, "the quick brown fox". We're going to encode into a sequence of indices. And then we're going to use our BPE tokenizer to decode. Let's actually step through what that looks like.
76:39 This — well, actually maybe decoding isn't actually interesting. Sorry, I should have gone through the encode. Let's go back to encode. So encode: you take a string, you convert to indices, and you just replay the merges — and importantly, in the order that they occur. So I'm going to replay these merges and then I'm going to get my indices. Okay. And then verify that this works.
77:15 Okay, so that was pretty simple. And because it's simple, it was also very inefficient. For example, encode loops over the merges. You should only loop over the merges that matter. And there are some other bells and whistles, like there are special tokens, pre-tokenization. And so in your assignment, you're going to essentially take this as a starting point — or, I mean, I guess you should implement your own from scratch. But your goal is to make the implementation fast. And you can parallelize it if you want. You can go have fun.
77:49 Okay, so summary of tokenization. A tokenizer maps between strings and sequences of integers. We looked at character-based, byte-based, word-based — they're highly suboptimal for various reasons. BPE is a very old algorithm from 1994 that still proves to be an effective heuristic. And the important thing is that it looks at your corpus statistics to make sensible decisions about how to best adaptively allocate vocabulary to represent sequences of characters.
78:24 And I hope that one day I won't have to give this lecture, because we'll just have architectures that map from bytes — but until then, we'll have to deal with tokenization. Okay. So that's it for today. Next time we're going to dive into the details of PyTorch and give you the building blocks, and pay attention to resource accounting. All of you have presumably implemented PyTorch programs, but we're going to really look at where all the FLOPs are going. Okay, see you next time.