Scaling laws 2
Why a second scaling lecture
00:05 Okay. So let's get started. Today's the second and last of the scaling laws lectures. And today it's going to be a little bit more of a case study and details oriented lecture. I'm going to cover two separate kinds of things. The first one is I'm going to go through a couple of papers where people have done careful scaling law studies as part of their model building. And I'm going to use that as a way to convey to you how modern large language model builders use scaling laws as part of their design process. Right? So motivation from last time and today is what's the best practice for scaling a large model. We want to have large language models with nice hyperparameters and good architecture choices. And you know, I've already told you about Chinchilla and using scaling laws to validate some of this, but I think you know, you should have rightfully skeptical questions about scaling laws, right? Like it's curve fitting on a log log plot. Is it really as good as you know, I said it was last lecture. So, does Chinchilla's approach to scaling laws actually work? you're finding this out in your assignments. You know, if you fit an isoFLOP, is that really telling you about the right token trade-offs? Can you use this stuff to really set optimal learning rates and should we be picking sort of particular architectures or parameterizations to scale nicely. So the last paper or the newest paper we talked about with like lots of detailed scaling studies in last lecture was the DeepMind Chinchilla paper right and after that ChatGPT happened and kind of the competitive
After Chinchilla, the frontier went quiet
01:45 Landscape of large language model building really changed and people just stopped publishing anything about you know data and scaling and all these things right it was sort of very secretive I've talked to you know people at some of the frontier labs before and ask them, oh, you know, like what are you guys doing for scaling? And they're like, no, we will not tell you anything about what we do for scaling. And so, we have to sort of rely on other sources for how scaling happens in practice. And there have been, several, you know, competently executed large-scale models that have done scaling. And so, you know, last year, in this lecture, I covered Cerebras-GPT, DeepSeek LLM and MiniCPM. And as a as a nice side note, you know, last year I had to really strongly justify why I was covering these like Chinese models, so to speak. But this year, thankfully, hopefully you're all already excited to hear about DeepSeek rather than me trying to convince you that this is the right thing to listen to. In the years since then, I've looked at a lot of the models that have come out. Actually the haul in terms of new scaling law insights and papers is actually much sparser. I'll briefly mention some results from Llama 3 which came out at the later end of last year. Hunyuan-Large, which is the model from China. And then MiniMax-01 which is a linear time sort of hybrid attention model from or long context model that came out this year. And all three of those have some scaling studies, but really nothing quite as extensive as DeepSeek or MiniCPM, which have really been the gold standard, I think, for modern scaling law studies. So, that's one part of what I want to talk about today. I want to make
03:25 Sure you guys have an understanding of, you know, what scaling looks like in a real, you know, semi-production model. And the other thing I want to talk about which is an important deep dive I think is the muP method that I mentioned last time. So, muP is this approach. Just as a recap of last lecture is you know when we train these models as we make them bigger we need to change certain hyperparameters right on the left hand side of this plot here you see that as you make models wider in this case like a MLP you make them wider the optimum learning rate sort of shifts downward so you need smaller learning rates for these bigger models and that's a really big problem potentially because then you need to hyperparameter tune your learning rates at the very large scale right and that's going to be very computationally expensive. It's going to be a huge problem. If on the other hand, if we could sort of parameterize our model differently so that you know the learning rate that's optimal just stayed the same forever across all the scales, you know, that's great. That's really simplified our search process, right? We would like all of our hyperparameters and really choices in general to remain stable across scales, right? That's the ideal. And muP is a is a very interesting class of approaches. And you know, it teaches us some pretty interesting sort of ways of thinking about the problem. So, I'm going to actually go through some of the details in the in the math. In the years since last time I taught this, there were a couple of very nice tutorials on muP that came out. So, I'm going to follow those because they have math that's pretty easy to follow. And then I'll talk about some work that has come out doing sort of third-party validation and evaluation
05:04 Of muP style methods. So, okay. The focus of the first part of this lecture which is the case study is going to be on three models. I talked about three additional more modern models but actually the details in those are much more sparse and I think the lessons you learn are primarily from these three papers here. So that's going to be my focus for the first part of this lecture. So, I'm going to talk about Cerebras-GPT, MiniCPM, and DeepSeek. And each one of these will have actually a pretty different mix of scaling strategies. And it'll also have different things to teach us about
Cerebras-GPT: muP for predictable scaling
05:36 How to get scaling, right. So, we'll get started. Cerebras-GPT is the first of the models and sort of scaling things that I want to talk about. It's a large family of models. It's trained 0.1 to 13 billion parameter models. Trained with the Chinchilla recipe. So roughly the same number of like token to parameter counts as is optimal. And you know they have a interesting core finding. The Cerebras folks actually are pretty interested in a lot of these like scaling and parameterization studies and they have a really interesting core finding which is they scale up this muP thing that I mentioned before and they find that it makes scaling a lot more stable and a lot more pleasant to deal with. And just to kind of show you the punch line, right? You've got test loss on the Pile. And you've got sort of the scaling curves here of like the Cerebras-GPT in blue. This is with standard parameterization. You've got muP in orange. This is the model that they you know also train with using the maximum update parameterization. And they show that it scales more nicely if not better than things like Pythia or GPT-J. So that's nice. And the thing that I want to emphasize here is that this is kind of one of the few if not first public validations of muP. We know that all of or most of the labs that are doing LM scaling pay close attention to how they parameterize their networks. Their initializations as a function of the scale of the model as well as things like per layer learning rates are things that people
07:12 Pay close attention to make scaling much more stable. And so things like muP are pretty important in this space. Llama 4 for example the paper for that isn't out and I don't know if they it will be out but they talk about a technique they call MetaP which is a variant of this as well. So what they show is that when they train models using sort of standard parameterization they find that you know they have sort of big oscillations kind of around the predicted scaling point. So that's the this dash line. You know, they have kind of oscillations due to the fact that, for example, they have to adjust the learning rate as a function of scale. And so it's hard for them to kind of really get the predicted performance sort of exactly right, which is this dash line, using their scaling recipe. On the other hand, what they find is you know if you have sort of the Cerebras-GPT — sorry — muP scaling then you get this orange line which is much closer to the scaling law fit for this muP version. And so their claim here at least is that using this alternative parameterization allows them to get much more predictable scaling much more nice hyperparameter tuning. We're going to see this in more detail. I'll return to this slide again once I've sort of gone through the mathematical derivation of muP. But in case you're you're ever interested in implementing this thing. Some of the Cerebras-GPT folks and in general, the kind of artifacts that the Cerebras research folks are putting out is very helpful for muP because they have this big table in the appendix that really just tells you exactly the difference between the standard
08:48 Initialization and parameterization or SP and the maximum update version or muP. And you'll see that, you know, I'll just give you kind of the one-liner version. Basically, every non-embedding parameter is initialized with one over the width. And then the learning rates per layer are scaled down by one over kind of the width, right? So, so the interesting difference from standard parameterization, even if you're already doing sort of one over width scaling on the initialization, is actually there's per layer learning rates that are different. And I'm going to get to that later. I'm going to do a full derivation of this result. But you can kind of see here there this kind of like nice quick reference and also if you want to implement this thing this gives you very easy ways of implementing muP. Another interesting thing that we also see in some of the other scaling strategies is that you combine these strategies like muP which makes hyperparameter selection stable with very aggressive scaling. So what they do here is they scale down their experiments all the way down to 40 million parameters. They do extensive hyperparameter search on this proxy model. And then they scale things back up using sort of muP to try to keep hyperparameters as stable as possible. And so this is sort of what they see in their small scale hyperparameter search. Each one of these dots is a model run and there's sort of a hyperparameter associated with each one of these and then they pick the minimum across these runs giving them essentially their sort of hyperparameter grid. This is a very like clean approach to hyperparameter selection. It's unclear whether you know this kind of this level of aggressive scaling down
10:24 Is really what you want to do if you want to train these like really large models. But this is kind of one strategy that we see also in MiniCPM and DeepSeek like training much smaller surrogate models and then trying to figure out how to stably scale them back up. And that's going to be kind of a theme that we see throughout. And yeah, if folks have questions please stop me actually. Maybe I'll stop here for a moment in case anyone has questions for the Cerebras-GPT piece. Although maybe it'll be clearer once I talk about the muP derivation later in this lecture. Okay. There is another paper
MiniCPM: small models, enormous token budgets
10:59 I want to talk about MiniCPM or another you know artifact I guess. For whatever reason, I think MiniCPM hasn't been talked about quite as much. Especially in sort of like western academic circles. But at least for me, this was one of the first sort of releases or papers I saw coming out of like a Chinese research group where they had done some like really cool in-depth, you know, scaling and other kinds of research. It really felt like, you know, stuff coming out of the frontier, right? And to give you a overview of what they do, their goals here is they want to train, you know, relatively small language models, but use a lot of compute to train really good small language models. That's their ostensible goal. And in doing so, they do a lot of careful scaling computations. They also once again use muP to stabilize and simplify scaling. When they sort of end up scaling these models, not in size, but in terms of the amount of data. And to try to convince you that you know this is a paper worth following right you know at the time that they were trained this was a remarkably good 1.2 to 2.4B models. It beats out, you know, most of the 2B models that were out there. And it matched many of the modern 7B models. At least modern as of 2024, standards. I mean, now of course you've got even better 7B models. The arms race is fierce. But this should give you a sense that at least, you know, given the amount of compute and technology available, you know, back in mid 2024, this was actually really at the frontier and they did something right to get models of this quality. And so much like Cerebras they do
12:39 Essentially they have to have you know some strategy to get scaling right so stepping back right you're going to do a really big model run what do you have to do you have to pick hyperparameters you have to make sure those hyperparameters scale nicely and then you scale up your model right so we can do the same thing as the Cerebras-GPT folks we can try to pick hyperparameters at a small scale hope that they sort of stay stable and then scale everything up and the way to do that would to use something like muP and this has exactly the same kinds of strategy at play here. You see so for the embedding you don't really do anything you just scale it by a constant whenever you have some sort of like residual connection like an MLP you scale it by the square root of the number of layers you initialize it by sort of fan in one over the base width and then the learning rates are also scaled by the width of the model. We see basically the same strategy or the same kinds of scaling factors appear as the Cerebras-GPT case, right? And they also end up with very similar parameters as Cerebras-GPT, the same kinds of scale_emb settings similar learning rates off by a factor of two or so. But generally you end up in similar places as these kinds of hyperparameters. And then what you do is once you have this you're sort of relying on your optimum learning rates to remain stable. So you're just going to keep those roughly fixed. And we know that the aspect ratio is a pretty important thing. So we just fix that after figuring out what the right one is. And then you scale up the overall model size going all the way from you know 9M or 30M all the way to half or one billion parameter models, right? And
Fitting the optimal batch size
14:22 So what they have is like a roughly 5x or maybe a little bit more compute savings going from the smallest models that they've got to the largest sort of pilot run models that they have. And now you can use this and then you can sort of figure out whether you have sort of optimal batch sizes as a function of scale. So you know you want to figure out crit which is the critical batch size. If you remember correctly this is the critical batch size is roughly the diminishing returns point right. So as models get bigger, their losses get lower. As their loss gets lower, you can make use of bigger and bigger batch sizes. So the critical batch size is roughly telling you for the given model size and scale that I'm operating at, what is what is a appropriate global batch size for me to be training these models with. And so much like the Kaplan paper, they follow a similar kind of recipe. You know, the plots look different from the Kaplan paper, but the underlying strategy is kind of the same. What they're trying to figure out is what is the critical batch size for training or the optimal batch size in this case of for training different models and they're trying to find relatively predictable scaling relationships between the batch size and for example the data size or the loss size and vertical columns here sort of represent a single training curve and then sort of the quadratics are sort of being fitted to try to identify the minimum right so the red line here is the is the minimum across all these points as we go upwards and this is trying to tell us the optimum batch size for a particular choice of model size and data set size. And then at this point you know you can follow the same logic as the Kaplan paper for identifying the batch
16:01 Sizes. Basically you reproduce the same kind of plot if you remember the Kaplan paper and the critical batch size discussion from two lectures ago. If not you can kind of pull up the lecture slides. You'll remember that basically the thing that's highly predictable is the loss that you're trying to train to the terminal loss and the batch size of like the critical batch size point. And so we see that once again much like in Kaplan you see a log log linear relationship here between the target loss or the terminal loss and the batch size that you want right and so from this you know you can kind of figure out what batch size you're going to get right because if you have a particular target scale you can use scaling laws to figure out what is the loss that I expect to get once you know the loss that you expect to get you can use that to back out what batch size you can kind of operate at right so there's a fairly clean trend polynomially increase the batch size as loss decreases now batch sizes do sort of shift around as a function of target loss and thus compute. So we have to fit a scaling law for that guy. But we already did muP and so in theory right if the approach works at all what we
Does the optimal learning rate actually stay put?
17:08 Should now get is we should get that the optimum learning rate here is stable. So on this plot we're seeing essentially different model sizes from sort of small models in the light colors to their biggest models in the dark colors. And you see them sort of varying different learning rates. The big models they're only running for a little bit for compute reasons. But what you see is a is a fairly clear trend and you know once again very consistent with some of the earlier results that we've seen in Kaplan et al. Where you have a relatively wide basin and then sort of sharp increases as your model becomes very unstable. Right? But the important thing here is that the minimum remains fixed across relatively large orders of magnitude. Right? From your small model to the big model, the minimum or at least tied with the minimum is at the exact same point at roughly 10^ the -2 learning rate. And so this is a nice sort of piece of evidence or some validation that properly scaling your model initialization and properly scaling your per layer learning rates allow you to avoid tuning learning rates over and
The cosine trap and the WSD schedule
18:13 Over or even fitting scaling laws on learning rates in order to try to predict what the optimal learning rate is. Okay. And then the final thing is, you know, you might want to figure out essentially model size to data trade-offs. If you're training small models, you're going to be probably overtraining your models or at least you want to justify to yourself why you're training on so many tokens. And so you might want to re replicate something like the Chinchilla analysis. So the MiniCPM people had a really cool or nice innovation. others have done similar things but I think they were the first to really popularize this in the LM setting especially in the context of Chinchilla style scaling is kind of the following so let's say I want to fit a Chinchilla scaling law when I do that what do I need to do well I need to vary the number of tokens and I need to vary model sizes right and so when I do that I'm going to fix a model size and I'm going to train a model for longer and longer it would be nice if I could sort of early stop and take the checkpoints of this model and have that be sort of the difference or changes to the data set size, right? Because earlier checkpoints see less data. It would be nice if I could use a single run to collect all this sort of data scaling things. Unfortunately, what I'm showing here, right, is that kind of the cosine learning rates for different data targets are different, right? So, if you have a very small amount of data, you have a cosine that goes up very quickly or sorry, that goes up the warm-up is always the same, but a very fast cool down, right? You train for a little bit and then you come down very quickly. If you have a lot of data, then you're going to very slowly come down to the end. And so your learning rates between a small data training run and a
19:51 Big data training run will be different. Right? This is a very key point. Right? Lots of people get tripped up by this. You cannot use a single run of a cosine learning rate model and try to get early checkpoints and reason about data scaling behavior based on that. Right? This is this bites people all the time. And so in order to avoid this, what you would normally need to do is you need to train a model from start to every single end point, right? So you have to train it to every single target. And so this kind of takes you to n squared runs, right? Like some of the runs are small, but you have to basically run lots of runs, each one with a target termination point rather than using a single run and collecting checkpoints. It's it feels kind of senseless that we have to do this. So, the MiniCPM folks, popularized this idea of a of a WSD or warm-up stable decay learning rate. And so this plot on the left really shows you, you know, what's going on here. Normally, what we would train with is something that looks like this cosine learning rate shown in yellow here, right? It goes up, there's a warm-up period, usually very short, to get to your full learning rate, and then there's a cosine that goes all the way down to kind of your termination point. And maybe you stay at your minimum learning rate. This is all of course optional. You might terminate here as well. You might go all the way to zero, right? And so cosine learning rate looks like this. And the issue here of course is that if I have a different target, the cosine is going to be totally different. So everything past the warm-up can't be reused. Now if you look at the this new WSD, which is basically a trapezoid learning rate, what it has is three phases. It's got a warm-up phase that's the same as a cosine. It's got a stable phase that's flat. And then it's got a decay phase
21:30 That rapidly cools down the model down to its minimum learning rate. And of course, you can have variations of this. You can go up, down, and then stay stable at your minimum. You know, you can do any of these variations. But I think in general, the simplest form to think about is warm up, stable, decay, terminate, right? Why is this nice? This is nice because you can reuse the stable part, right? So the thing that you do is if you want to do Chinchilla in almost one run, what you do is you sort of warm up. You have a stable run all the way to the end and then you cool down and then if you want to figure out oh how would my model have been if I used less data you rewind the checkpoints and then you do another cool down right and now you've got a exact warm-up stable decay learning rate shape without having done the training from the beginning right so this is a very nice thing the fact that the stable part essentially is flat allows you to do Chinchilla style scaling or data scaling in a single training run or for mostly the cost of a single training run and a lot of people now do Okay. So they work very well. MiniCPM I think popularized this and you know I think a lot of people have since then adopted it and we see a lot of WSD style schedules in many places. They you see curves that look kind of like this. If you have a cosine learning rate schedule, you'll see essentially relatively predictable smooth decay towards your terminal loss like this yellow line here. If you train with WSD, you'll see much funkier learning curves that look like, you know, the curves that I have here above them, the darker lines, right? So, you've got your warm-up phase, which
23:05 Doesn't really show up in this training curve. It's so short. Then you got your stable phase where it sort of goes down normally, and then as soon as you hit your decay phase, like the cool down part, your loss really rapidly drops off until you've hit your sort of either zero or minimum learning rate point, at which point you've gotten your terminal loss, right? So these losses may look very disturbing to you but they are actually pretty normal when you're training with these kinds of sort of rapid cooldown learning curves. And maybe the point to make here is at every single token count you see that the warm-up stable decay curve you know the minimum point beats or matches the cosine learning rate. That's not always the case. There can sometimes be cases where cosine works better WSD works better. But in general, I think a lot of the I think things that people say here is that the two learning rates are roughly comparable, but WSD has the additional nice advantage that you don't have to worry about your termination point. You can repeatedly cool down to get checkpoints of different data counts. Okay. Cool. Okay. And then of course
Cheaper Chinchilla, and the 192-to-1 fit
24:09 There's other things that have appeared for trying to estimate Chinchilla. Some folks a collaboration of like UW — formerly UW and Apple folks had this paper on estimating sort of the Chinchilla penalty that is when you keep adding more and more data you know how much worse is your loss than if you had you know scaled according to Chinchilla. So you have your sort of teal line here which is sort of m equals 20 you know 20 tokens to parameters and you can sort of think about okay what happens if I train with 320 tokens to parameters well then you have a separate parallel scaling line and then you have another line which is the circles which is what if I train with 640 or the darker one is there and so the thing that they show is actually instead of doing this WSD style thing another thing you could do is you could try to figure out okay how much does my model degree grade as a function of sort of higher tokens to parameter ratios. Well, that turns out also to have a fairly predictable shape. And you can sort of extrapolate that based on sort of degradation at small training runs. I don't think I've seen large scale training runs using this idea, but it's kind of an additional cool thing to know that essentially you could do Chinchilla in almost one training run by sort of extrapolating the excess token penalty at a small scale as well. So okay going back to MiniCPM now we have the tools that we need. We have the WSD learning rate which allows us to essentially do one training run and that one training run allows us to you know have both variation sorry not that one training run allows us to vary data as we go along and then we have multiple
25:47 Training runs for different model sizes that gives us all that we need to do Chinchilla analysis. And they use method one and method three if you remember what those are. Method one is you overlay all of the learning curves and you take the lower envelope and the lower envelope of all the training curves is supposed to be roughly a power law. And then method three is you basically jointly fit this equation two you have here. You hypothesize this two variable scaling law and you kind of fit it to all the data that you have in kind of curve fitting style fashion. And then that you know allows you to solve for the optimum token to data ratio through that fit. So they do both of that. They do see for Chinchilla method one fairly clear although not perfectly linear trends that allow them to essentially go from compute to token ratios. And their primary approach that they use to justify a lot of their design decisions is the method three. It's the curve fitting. So the kind of contours that you see here is the curve that they fit. The dots that they have here is the small scale runs that they did to fit the Chinchilla parameters. And just to you know sort of justify what they do they find very high token to parameter ratios like so high that I feel like this is an outlier that doesn't really agree very closely with most of the other literature. They argue that Llama style architecture should all have a higher ratio because of you know improved data quality and improved model efficiency but their token to parameter ratio estimates are really high 192
27:24 Tokens per parameter which I don't think I've seen anyone else derive I think other people have done replications of Chinchilla I don't think anyone's ever really done or argued for 192 tokens to parameter regardless you know we have seen that recent models like Llama 3 have significantly higher data to model ratios. We also don't really see diminishing returns. Like these models aren't way worse than the equivalent Chinchilla scaled like Llama 2 models. This kind of suggests that with careful optimization and careful tuning, we should be able to go far beyond like the 20 times model size rule of thumb, right? So if there's one thing you take away from this set of sort of, you know, these last two slides, maybe not necessarily that you should trust whatever scaling law fits that MiniCPM vid. But rather that sort of the Chinchilla analysis isn't really you know a strong constraint right like 20 times model size is just a starting point you should be you know feeling free to significantly increase that token to parameter ratio finally the curve fits that they get are generally pretty good looking so this is the scaling lock curves for essentially data and model size scaling and perplexities on code in English they do have some really weird outliers that I don't really understand why they get these but their sort of fitted scaling laws are generally pretty good as they increase the amount of data on their relatively small model. So this is one example of a you know large scale training run scaling recipe. So I'll stop here. Things like WSD are probably new to you. So you know if you have any questions
29:04 Please feel free to ask or any of the other bits including like the Chinchilla replication and muP and so on. Oh okay sure the main adaptation of muP is in terms of initializing the weight of parameters right so the question was the main change in muP was initialization so there's two things that will happen when you derive and you implement muP one of them will be the initialization will change and the other thing will be that the learning rate changes or the learning rate changes per layer right and that is probably a more exotic object than many of you are used to the initialization actually is not that different. If you're already using like a standard like Kaiming-style initialization, that's already one over the fan in one sorry one over the square root of the fan in which is going to be already the right thing. Whereas the learning rate normally like unless you're doing something really exotic, you're using like a global constant learning rate, right, everywhere. So that's going to be a big difference from what you're you're normally training with. So you can think of that as like the practical difference for a lot of the muP implementations. Yes, was kept constant. We saw that the curve was very close to the cosine decay. Yeah. So you're talking about this curve and you're saying like oh when we're in the stable phase of WSD like when we're up here the curve remains pretty close and like why is that? Well, it's kind of close, but also not really, right? Like if you look at this last curve over here, you
30:42 Know, there's a big gap before we enter the decay phase between cosine and WSD. And I think this is like one of the pretty interesting, you know, mysteries about deep learning optimizers. Clearly you need a stable phase to kind of get you far from your initialization, but also the cool down phase is what gets you most of your gains and your losses, right? If you don't cool down, this is a gigantic loss in your losses. So the cool down is actually really critical. And you know, a lot of the gains from cosine versus here, like this relative gap, this is all from cool down, right? And so a
DeepSeek LLM: fit the hyperparameters instead
31:15 Lot of the optimizer learning rate design is about this balance between like how do I keep learning rates high to travel far from my initialization but still have, you know, good decay on my learning rate to be able to anneal my loss down to a very low value. So the other paper I want to talk about is DeepSeek. This is the original DeepSeek LLM paper from 2024. And in many ways if you read the original DeepSeek LLM paper, you'll kind of you can know that these are like very serious science people, you know, because they do a lot of very careful scaling ablations and they're really trying to get it right when they scale up. And that's kind of an attitude that's shared amongst sort of the players that get scaling right. So you know they have seven and 67b parameter models. You know at the time very high performance relative to you know Llama, which is really the primary competitor at the time. And you know there at the time I guess you know Llama 2 and Mistral were kind of the big players. DeepSeek comes in and they're able to match the performance. Not quite the flashy impact of DeepSeek V3 coming in and sort of matching OpenAI's GPT-4o, but you know, for a first time sort of attempt, this is a pretty remarkable result. And so, let's kind of dig in and try to understand, all right, like what did DeepSeek do that allowed them to go from essentially zero to you know, at least for open source state-of-the-art at the time, right? So, and I think DeepSeek more than you know most other players maybe the
32:53 Only comparable one being MiniCPM is very open about a lot of the experiments they did and their the approach they use to choose a lot of these hyperparameters. So immediately we see one difference between DeepSeek v1 and MiniCPM and also Cerebras-GPT which is that they don't use any muP. And they're going to directly try to estimate both the optimal batch size and the optimum optimal learning rate. So it's like a really direct method you might call it and requires kind of a strong belief in scaling laws. So what they do is they you know take two relatively small models and they have they run a grid over different batch sizes, they run a grid over different learning rates and they get losses across this grid. They do the same thing at a larger scale and you know you can get kind of the optimum batch size and learning rate right and so they're saying all right well this is a pretty wide basin so we don't maybe have to be you know too scared about messing this up and so then what they do is they know that all right this the choice of learning rate and batch size are both relatively forgiving but we do want to get the order of magnitude of these things correct right so how do we get the order of magnitude of these things correct Well, you know what we're going to do is we're going to train a bunch of models with different amounts of you know non-embedding flops and we're going to change essentially across a grid the parameters that I had before both the batch size and the learning rate and by varying these you know we're going to have the optimum
34:29 Batch size and the optimum learning rate sorry across these different scales right so you can imagine basically making these grids across many different flop scales and basically marking down a are for each one. Perhaps unsurprisingly because it's the scaling law lectures, these things seem to kind of follow a scaling law line. At least for the batch size, things seem more clear and you can kind of fit a line to here and you can, you know, extrapolate out to the big models that you're going to train what your optimum batch sizes should kind of look like, right? They do the same thing with learning rate. And they sort of fit this line and they say, "Oh, these are the two learning rates we're going to use." it might be because the points are being plotted on top of each other, but I find this line to be particularly not particularly like somewhat suspicious looking. I mean, I could have probably fit a horizontal line and that would have also looked okay. This one, I don't know, even as a as a scaling law enthusiast, I'm not quite sure I would, you know, bet my life on this one to pick the learning rate, but they did. And you know, that's how they get the learning rate. now they also kind of follow best practices at the time. They do a Chinchilla style analysis and they use once again a WSD style learning rate. Where you know they're trying to essentially minimize the amount of repeated work that they do. They do kind of something a little bit weird or a little bit more non-standard where what they're doing is you know they do warm up, they do stable and then they do two sets of decay steps decaying down to zero. So it's like two, you know, decay phases consisting of kind of like 10% plus 10%. And they sort of analyze
36:07 Different choices of that decay phase. And they kind of doesn't seem to matter very much, but generally speaking it, you know, it's about 20% of the total compute budget is going to be spent on that cool down phase. And so they also show once again that it matches cosine learning rates. But once again the advantage here is that we can do Chinchilla style analysis for very cheap in contrast to the learning rate sorry learning rate fits you know Chinchilla style analysis just fits really cleanly. I think this is a a broad lesson when you look at lots of people's scaling laws. I think the stuff on hyperparameters always looks a little noisy and tenuous. But the isoFLOP analysis from all the players look always like very nice. And so this is you know replication of the Chinchilla result. You see you know different compute scales. We see different quadratics. We draw a line through the bottom of the quadratics. We get you know exactly the kinds of sort of optimum sorry optimum flops per token and optimum token size as a function of training flops. Right? So this gives us a very straightforward way of analyzing the token size to model size trade-offs. And this allows them sort of do everything from scratch, right? Of course, I think, you know, as a side commentary, I think it's really nice that they're kind of redoing a lot of this. Like they could have certainly cargo culted Chinchilla and just picked 20 tokens per parameter but they said no like let's actually go and do the scaling law analysis and like let's actually make sure that you know that the token sizes are relatively appropriate for
37:45 Us. Okay. And then they have you know a fitted scaling law at the very end. This is in some ways not surprising because this is after they fix their scaling strategy. They do predictable scaling. They try to predict what happens on the 7B and the 67b models. You know, it's unsurprising in many ways, but very nice that they're able to extrapolate out from about 10 to the 20 to 10 to the 24 and actually nail the prediction on the basis of the scaling law. Right? So, it's a very nice thing to be able to see that we can actually get predictive measures of model capabilities before we actually train them. So, that's kind of the DeepSeek part. Anyone have questions for kind of the DeepSeek strategy and kind of what they did and any of the other pieces? I think most of this I think WSD was probably the newest thing that I've sort of mentioned today. The other thing that DeepSeek does is directly fitting a scaling law to the optimum learning rate and batch sizes rather than using something like yes. Do they have a global learning rate? Yeah. So they're tuning that global learning Cool. Okay. Yeah. Once we know the problem. Yeah. So, so the question was like do people redo this kind of analysis for new frontier models? To be honest, I'm not actually sure and I'm beginning to think that a lot of people maybe don't like exactly replicate some of this. Because we see that in the newer paper is just increasingly less scaling details. Like even from DeepSeek
39:24 For example like DeepSeek v2 and then v3 we see a lot of emphasis on the new parts of each paper like so for DeepSeek v2 we see a lot of emphasis on like MLA and like the architectural improvements and then DeepSeek v3 we see a lot of the systems components being emphasized like the low bit training but we don't see for example in either of those any additional new scaling law studies and so I think my guess is that there's not much new there like maybe they're replicating it just to make sure it works but nothing new to report. And I think that will kind of be captured in the next
Llama 3, Hunyuan-Large, MiniMax-01, and the recipe recap
39:58 Couple slides where I'm going to talk about scaling laws and papers and models from the last year or so. So I did a little brief survey. But actually there's nothing that is at the level of detail of either MiniCPM or DeepSeek like those are really still I think the most detailed open studies into scaling that we have in 2025. Cool. Okay. So, you know, Llama 3 was probably one of the bigger model releases in the past year since I last taught this class. And they do have some pretty interesting scaling bits. for one, you know, just the question right now of like, do people actually replicate these analyses once they've run them once? Well, kind of yes. Llama 3, you know, redo the isoFLOP style scaling Chinchilla scaling laws and they find you know roughly that the optimum ratio if I if I got the calculation right is about 39 to1 right and I do think this is interesting because you know Chinchilla got the 20 to1 parameter ratio I think many of us have trained models at the Chinchilla ratio in our research and so on you know it's quite clear that the 20 isn't really that stable like other people that have in fitting it have been getting generally slightly higher ratios than before. And that might point to things like improved algorithmic efficiency in sort of architectures that learn better from data. It might mean something else like improved data quality. All of those are kind of moving parts. So, it's hard to know like what's leading to these slightly different ratios, but the results seem fairly clear. The fits are relatively good and they do get a 40
41:33 To1 ratio. The other thing which is close to the data scaling stuff that I mentioned the early parts of my first scaling lecture. One of the interesting things that the Llama 3 folks do is they try to essentially correlate compute into NLS like log loss and then correlate those NLS back into downstream accuracies. Right? And so, the thinking that they're trying to do here is they would like to not really scale against log likelihoods. That's not really a thing they truly care about, right? They care about improving, I don't know, benchmark numbers on MMLU or LAMBADA or whatever other benchmarks that they've decided to hill climb on, right? And so if that's the case, then what they're going to need is to have a conversion factor going from these NLS per character, these, you know, perplexities or equivalent to perplexities, and then map them into accuracies. And so they've done some studies in Llama 3 essentially trying to relate these two fitting sigmoids showing that you know if you fit essentially these small models and you fit some Llama 2 models and you fit a sigmoid on the whole thing you can accurately predict the performance of Llama 3 405b on the basis of those fits. It's interesting. I think they say that they use these kinds of ideas for data selection. But I think it's a there's not that much details there and it's unclear whether this is like a really core object when Llama 3 was being trained or whether this was kind of a sidecaling thing that was like just of interest to the authors. Another recent work that has come out sort of yet another
43:14 Chinese LLM that's nicely executed is Hunyuan-Large. Hopefully I didn't really butcher the pronunciation there. They are training and so because they're training they want to kind of redo the Chinchilla style analysis they fit once again they do isoFLOP analysis they fit quadratics they figure out the minimums and then they're able to get a different token to parameter ratio so they get a 96 to1 data to active parameter ratio these ratios are obviously going to be quite different because they're training there's lots of differences about the architectures we don't really expect the same thing as Chinchilla right and so we do actually see in various papers Essentially replications of Chinchilla happen again and again because a lot of these people are very interested in understanding like how far can I push the token to parameter size ratio. We would like to stay on the higher end of that right like have more data than parameters because then people will actually use our models or our models will be cheap to serve right so for all those reasons people have been replicating Chinchilla I think this is one of the best replicated results in scaling in many ways the actual 20 to1 parameter ratio isn't the thing that you know consistently replicates but the fact that you can do isoFLOP and fit the minimum and get these like very predictable tradeoffs in flops to minim optimal act optimum parameters is quite clean and consistent in the replications. Okay. The last one which is honestly a little bit more of an exotic scaling law over the last year is MiniMax-01 which came out pretty recently. So MiniMax-01
44:55 Is a kind of linear time or long context language model released by another sort of Chinese startup. And their interest is well what we're going to do is we're going to take softmax attention which is quadratic and you know they have this thing called lightning attention which is a kind of linear attention or linear yeah linear attention layer which is linear time. And then you know they have a hybrid version of this model and they want to figure out like all right, how much cost am I paying in terms of the performance of the model going from softmax to linear to hybrid attention. And so they do things like they basically replicate method one for Chinchilla where they're looking at the lower envelope of the loss curves as they train. They look at essentially the implied optimal model size and the implied optimal token count as they go. And roughly the conclusion that they draw from this is that you know the lightning and the hybrid models you know roughly perform the same as the softmax attention and thus they can they're you know okay to train long context models on the basis of these architectures. We've seen these kinds of plots occur very often in research papers. Like if you look at the Mamba paper or the Mamba-2 paper or any or the Delta paper or any of these other kinds of linear time complexity RNN papers, you'll see plots that look a lot like this where they say, "Oh, the full attention scaling and my linear attention scaling are basically the same as a function of compute." But this is, you know, I would say like a kind of a rare case of this same plot being produced almost at scale from a major sort of artifact release. Okay, so putting all that
46:37 Together, right, I know that was like a bunch of mini case studies that I went through fairly quickly. But I want to sort of step back and recap it a little bit, right? We've seen several common ingredients being used in these scaling recipes. We've seen Cerebras DeepSeek MiniCPM and then the few new papers since so Cerebras-GPT and MiniCPM both use muP as a way to make hyperparameters more stable across scale. And you know they essentially MiniCPM especially has a nice WSD schedule which is a thing they popularize to be able to do Chinchilla style scaling. Cerebras doesn't bother to replicate Chinchilla. DeepSeek seek does a little bit different thing. They assume that most hyperparameters just don't change with scale but they do a full scaling analysis on batch size and learning rate and then they use the scaling laws as a way to figure out optimum scaling. You know I've already noted that some of the scaling looks a little bit more suspicious than others but really this is a way to at least get the order of magnitude hopefully right. They use isoFLOP analysis. You know, they replicate Chinchilla once again to figure out the model sizing and to make sure they're kind of in the right order of magnes. You know, Llama 3 and Hunyuan-Large do isoFLOP analysis only. Llama 3 does a little bit more, but that's basically it. And then MiniMax-01 does the more interesting thing of basically justifying architecture choices through the lens of a scaling law. But we see generally speaking that there's a few different things that get replicated like Chinchilla and learning rate and batch size are really the things that people are really deeply concerned about when they're scaling models up and they sort of do things like fixed aspect ratio and just scale the total model size up and that's
What muP actually is: two Theta(1) conditions
48:16 Generally the way that people handle a lot of the moving pieces of scaling up. Okay. Any questions about the case studies pieces? Actually, I'm going to stay here and just make sure I've covered any questions that people might have. Okay, cool. Okay, so the second and kind of last part of this lecture is going to be understanding up. Hopefully through the case studies you've seen that essentially getting the learning rate right is one of the core concerns that people have. And also the batch size. But in general, I think we want to have scale and variant hyperparameters. And it is the case that you know our choice of initialization and our choice of per layer learning rates are essentially arbitrary, right? Like there's no reason why we have to initialize them one way and not the other. And so if we could manipulate those sort of three variables to get scale and variance in our learning rates, that would just be really wonderful, right? Like that would make our lives way easier and it would make, you know, small scale experiments much more possible. So you know I'll also talk you know first through the math of this like how it's derived what's the justification what are the core conceptual objects behind trying to make models scale predictably and then I want to talk about a pretty nice preprint by an independent researcher on basically just a bunch of ablations on up like what makes it break what is it robust to does it work on a real you know transformer language model these kinds of questions are explored
49:53 Pretty well in this preprint that I'll talk about at the very end here. So, okay, what is muP anyway? I feel like maybe I've jumped the gun for the last two lectures because I've mentioned what this is without really giving you the core conceptual object that it's based off of. On the other hand, I think I'm I'm justified in doing this because I think most of the literature doesn't explain muP that clearly either. They're just like, yeah, just scale the initialization by one over the width and scale the per layer learning rate by one over the width. That's muP. But I think the ideas behind muP are pretty interesting and worth discussing because I think they speak to some core objects that recur in deep learning in general. So, I'm going to be basing my slides off this preprint or paper. If you're interested in kind of reading about muP, I would point you to this one. I think this and another blog post called a practitioner guides to muP I think are the two kind of readable descriptions of what the sort of this paradigm is. Okay. So I'm going to base myself off this. The math is for whatever reason not exactly the same across these different presentations. So I'll I'll you know clarify that I'm basing the math off this one. So muP is based off of this the following relatively simple ideas, right? So there's two things that we think should happen when we're training a neural network, right? So you know when we scale a neural network, we're going to make the in this case, let's just say only the width, the width of the network bigger, right? I'm going to fix the layer size or the sorry the depth and I'm going to make the width bigger as we go. Now if I do that, as I make the width bigger, I want the activations at
51:31 Initialization to remain you know, big theta of one, right? I want it to remain roughly constant, you know, bounded above and below by universal constant, you know, roughly constant as I make the width bigger. It shouldn't blow up. It shouldn't vanish, right? Seems like a pretty natural thing to want, right? You don't want your activations to get too big. This is per coordinate. Now, the second assertion I want is that, you know, I'm going to initialize my model and I'm going to take a single gradient step. And when I take that single gradient step, I want to make sure that the change in activation should also be big theta of one, right? So both of these seem like very natural conditions, right? Because if you violate these, it's going to mean that, you know, as I make the models bigger, either the initial activations will blow up or vanish or after one gradient step, my activations will either blow up or vanish, right? Those are both bad conditions, right? And as a note, right, I'm talking about individual activations like coordinates. And so if you're thinking about norms, right, of an entire vector of activations, right, that should look like, you know, big theta of square root of NL, right? Because each one of these are going to be roughly independent. So the norm is
Deriving the initialisation and the learning rate
52:37 Going to look like the square root of the width the number of elements in my in my width coordinate. So I can derive you know muP from those two conditions. So the first condition which is that I want my activation to remain stable imposes sort of constraints on the initialization right so I'm going to walk you through a very simple example right so I'm going to consider a deep linear network so this is h of l so this is the activations at layer little l and that's going to be a function of the weight matrix at layer l and the activations from the previous layer no nonlinearities no fancy stuff right like you know it's all square just forget all this complexities if you want complexities you can go read the preprint they'll explain in slightly handwavy terms why those things don't matter now the initialization I'm going to pick a gausian initialization right so it's going to be zero centered it's going to be a rectangular size of the of the sizes that depend on the sizes of my activations and then I'm going to have one hyperparameter which is the noise scale of this matrix at this layer. Sorry, there should be a little L on this sigma. So now what can we say? Well, I want to understand the size of H of L, you know, at initialization. So, how can we do that? Well, one thing we can do is we can consider sort of the limiting behavior of this system, right? I'm going to take the basically little n of L and little N of L minus one to infinity. And if I do that, this W is going to concentrate. It's a random Gaussian matrix. If you remember your
54:18 Your random matrix theory, actually that's not a prerequisite for the course, but you know if you know some basic random matrix theory, you know that the operator norm of a g gausian matrix is going to roughly concentrate to this object, right? It's going to be sigma which is the noise scale times, you know, the square root of both of the coordinates added. And importantly, you know, you can write down roughly that this equivalence is true, right? So the activations at layer L the norm of that is going to be approximately equal to the operator norm of WL times the activation norm of H of L minus one. Right? And this is roughly assuming that W of L is independent of H of L minus one which is true at initialization. So I think you can basically make that a right arrow if you'd like. Now I'm going to say I'm going to pick a particular choice of sigma which is going to be square<unk> nl<unk> nl minus one times this object. You can simply think of it as this right hand side thing. This is the exact form. This is kind of the more asymptotic form that you can think of this as, but really it's just one over the square root of the fan in of your layer times the minimum of one and sort of the aspect ratio of your model in case your fan in is much larger than your fan out then this sort of kicks in. Okay. So let's say that I pick this sigma what happens right roughly one over the square root of my fan in. So now what happens? I can I can plug this back in to this formula, the matrix concentration limit and also this approximation here and I can sort of inductively prove that every layer is going to have the right sort of activation size. So let's just go through all the layers and assume that up until layer L minus one I have this
55:58 Property right so that's the inductive assumption at layer L minus one I have that my activation norm is square root of N L minus one. Okay, so that's just an assumption. Now if this is true then I just plug all of these in right so I plug in square root of nl minus one into this component into wl operator norm I plug in the limit and then for sigma I plug in this expression over here you see that this inverse cancels this and then you're going to get exactly that h of l the l2 norm of h of l is equal to square root of n of l. So this is the thing that we wanted because before remember we said we want to make sure that the activations remains big theta of one which means that the norm should be square root of n of l. So that's exactly what we get plus some lower order terms right. So you know this is a fairly clear step-by-step argument that shows you what the right thing to do is for initializations. I want to pick one over the square root of the fan in plus a small correction factor in order to make sure that my activations do not blow up at initialization. Right? I'll pause here for a moment in case someone has questions. I feel like this is actually maybe the first like real math that we've done in the class. So maybe it's a bit of a context switch for people. I did not warn you that I was going to talk about a bit of math. Okay. Is this all relatively clear for people? One over square root of fan-in. Yes. Okay. I'm gonna assume that everyone's on board with one over square root of fan-in. Okay. So now we're gonna derive the second part of muP. Right. So the first part of muP
57:36 Was about initializations. The second part of muP is going to be about learning rates. Right? And so how are we going to think about learning rates? Well to think about learning rates I'm going to look at the second condition. The second condition A2 which says when I take one gradient step past initialization what needs to happen is that my activation sorry my update size needs to remain constant. It can't blow up. It can't vanish. Okay. So what does that mean? So if I have an update of delta WL on the weights at layer L what where does that come from? Well that comes from let's say I'm doing SGD. That's going to come from this expression. It's going to be a learning rate times L which is my loss, the gradient of L which is my loss and then the activations transposed. In the case that my batch size is one, this is a rank one object, right? This is a rank one update to delta of L, right? And because it's rank one, you know, there's a nice easy expression. The change of WL times the activation of the previous layer is equal to the norm of the change in WL the operator norm of this thing times the L2 norm of H of L minus one right and now you know combine this with the fact that the change in activation at layer L is kind of this expression you can convince yourself that this is true you can write this out by sort of figuring out what the actual final activation is at layer L after the update and that and canceling out wlh of ll which is a shared term across left and right then you'll get this expression you'll get that you know what is the update in h of l this is the object that we want to keep roughly square root of n of l right the norm of
59:15 This object so let's look through each of these terms and look at what the magnitude of this is the first term here wl delta h of l minus one this you know we can assume is going to be controlled from the inductive assumption because this is exactly the sort of the delta h of l that we have plus the condition a1 argument right condition a1 basically says sorry condition a1 is going to say that delta of h of l minus one is going to be square root nl and then wl is going to maintain that norm the more complicated parts is going to be these two arguments the second and third terms that we have here delta wl h of ll minus one and delta wl delta hl minus one sorry that's that's quite the mouthful they all have the same order of magnitude actually. And the only thing that we need to really figure out is this expression here. What is the product of the previous layers norm times the operator norm of delta WL? Right? Because we don't really know how big the update is going to be in the weight matrix W if we knew that. All very straightforward stuff. Okay. And so the remaining argument is actually relatively straightforward. Even though this is actually like a you know complicated jumble of things the intuition is actually very clear. The intuition for this is says okay what do I really need to figure out? The one thing I really need to figure out is this expression here. How much does the weight at layer L change right? If I can figure that out then I can sort of derive all the relevant quantities and solve for the learning rate. That's at a high level that's our strategy here. And so how can we possibly figure out a after one gradient step how much delta WL moves right that's really the key
60:58 Question well there's an additional sort of sneaky assumption that then shows up here and the assumption is something like this if our learning is well behaved then after a single gradient step then the change in the loss delta of L this quantity has to also be big theta of one right and why is that well because we don't want the size of our losses, the update, the decrease in our losses to kind of blow up or go to zero as the width goes to infinity, right? We want essentially our improvement in losses to remain roughly the same order of magnitude no matter how big our models get. That's a stronger assumption than what we've had before. But assuming that's true, then essentially we can say, okay, the change in the loss is kind of like multiplying the gradient with the change in the weights. This left side is O of one. We know how big this delta of L should look like. So now or sorry we know how big this delta wl looks like. Now we can solve for the gradient size. And once we have that we can plug that in here. We know delta wl we know the gradient of l. We know the size of h of l from condition a1. And now we can solve for the learning rate. And that's exactly what you get at the bottom here. And the you know if you work through the arithmetic the final result that you get here is that the learning rate for SGD is equal to the fan out over the fan in right so lots of steps involved and lots of like substitution and slightly sketchy bigo notation being substituted into the equations here. But once we do that we're going to end up with a very simple formula. Note that this is true for SGD. And those of you that have kind of been paying attention and like kind of staring at
62:34 This equation are probably, you know, internally complaining. You're like, you have you have misled us because, you know, in a transformer, well, what's NL over NL minus one for like a MLP that actually is like a four, right? Because you've got a factor of four between DFF and D model, right? And so this thing doesn't really change. It's just a constant in most models, right? Unless your aspect ratios are like dramatically changing through your network. The reason why muP is different from standard parameterization is because this derivation is for SGD where the parameterizations look very similar between muP and SP. If you do the exact same derivation for Adam, you're going to find that actually you're going to get slightly different things, which is that it's going to be one over the fan in rather than the fan out over the fan in. Okay, so here's the recap. I have sort of dragged you through hopefully willingly the derivation of the basic what people call the spectral conditions that define muP. But now I will give you the kind of one slide highle takeaway of that result. Right? So you know when we want to do something like muP if we following the guidelines from before directly what we will end up with is the following blue box. At initialization, you know, you set yourself to one over the square root of fan in times a correction factor. That's, you know, one if your fan is smaller than your fan out, but you know, square root of the ratio otherwise. And then this is a simple initialization for the scale of your gausian. For your learning rate, if you're doing SGD, then you set it to fan out over fan in. But if you're doing
64:12 Adam, that's going to be slightly different. It's going to be one over the fan in. Now in case you already sort of know the standard Kaiming initialization and so on off top of your head you know you can mentally compare what this looks like to the standard parameterization. So in a standard parameterization if you're doing it right you should probably be already setting your gausian to initialize to one over the square root of fan-in so that's good that's already perfectly set but your learning rates are probably being set globally to a constant this is fine for SGD not so fine for Adam where the really big difference between sp and muP comes in. So okay that brings us right back to kind of the Cerebras-GPT paper. Now we have all the context we need to understand all the operations they do. If you look once again at the column over here of muP the embedding layer is special. It doesn't really do essentially like any scaling because embeddings are one hot so their norms don't scale linearly with the number of vocab elements. But ignoring that basically you see that all the layers get scaled down by one over the width. You know that's the initialization rule and then the learning rate rules are scaled by one over the width as well right so this is once again the learning rate rule for Adam right so if you're using Adam that's exactly the right thing to do and that's also exactly what they do in Cerebras-GPT so hopefully that's clear and hopefully this gives you a sense of you know both the interestingness of manipulating per layer learning rates to get more predictable scaling and also maybe an appreciation of this idea of trying to control activations and
65:53 Updates as a function of model width right like you know I'll pause for a moment there and just mention right that's like a very successful idea from physics right lots of physicists think about ideas like renormalization as I take limits of certain things I want things to remain stable I want them to not blow up or go to zero this is an exact application of that idea that's that's kind of an interesting use of that okay any questions about I don't know muP derivation or Cerebras-GPT or any of the other things. Yes, there's no assumption about any architecture, right? So any transformer or any model? Yeah. So that is part of the subtlety. The question was you know what's the architecture assumptions? Well I mean technically there's a even stronger assumption here. Oh why did I go? There's an even stronger assumption here which is that I'm assuming things are a deep linear network, right? I'm just multiplying matrices repeatedly. You know this is the kind of silliest network that you can have. Basically there are arguments for why adding nonlinearities are fine. There are arguments for how you would take the same arguments and apply them to the attention layer. There are arguments for why much more complex things are needed for a gated linear unit. So each one of those architecture pieces needs a careful analysis in order to have a you know corresponding object. Yes. How are these like ends determined? Like it looks like they're indexed by layer. Right. Right. Right. Right. Right. So, so N subl is just the output of a matrix multiply and N l minus one is the input. So for example, if you have a MLP, you're going to have a matrix multiply that takes you from D model dimension to four times D model like the DFF dimension, right? So that
67:32 Would give you a NL over NL minus one of four for example. So all the different matrix shapes are giving you the NL and NL minus one. Oh, the fan in and the fan out of a matrix. Yeah. Exactly. Yeah. Yeah. So, so the input and output dimensions are determining all these objects. Okay. Excellent. It's also just in and out of the Yeah, fan in and the fan out are like the input and output dimensions. Yeah, I was using those terms exchangeably, but I should have been a little bit more clear. Oh, yes. Since DeepSeek only has a global learning rate, does it mean that they don't have order one updates? So, so the question was like since DeepSeek uses a global learning rate, does that mean they don't have an order one update? So, you know, all this argument is asymptotic, right? It's basically saying as I scale my width out to infinity, things will kind of be big or small. And I mean if you if you look at the muP plot for example, you do you do kind of see this, right? You see that the learning rates have to shift as the model gets larger in order to compensate for the fact that the updates are getting bigger and bigger, right? What's empirically, you know, been seen is if you do nail the learning rate, you don't need muP, right? Like it's not like muP is necessary for you to train a good model. It's really just an attempt to try to keep this shift as small as possible so you can use the same learning rate, you know, throughout scaling. And if you go back to DeepSeek, you know, if you remember
69:07 The, scaling law that I was, you know, being a slight hater for, you'll see that, you know, they too have learning rates that go down as a function of scale in order to try to compensate for the fact that the bigger models are going to have bigger updates, right? And so you know to respond more directly to the question yes in the case of DeepSeek right as we scale the model up you know our activation updates will get bigger so we have to shrink the global learning rate or we should shrink the global learning rate to compensate for that cool.
Does muP survive contact with real architectures?
69:40 Okay nice questions. Okay so that was kind of the conceptual somewhat mathematical components of muP. Now I want to talk about the empirical aspects of muP. And so I'm going to talk through a preprint or I think this one's being published at COLM a large scale exploration of mu-transfer. And I like this because it's got a bunch of ablations and I think I'm a sucker for ablations. So I'll present any paper that has large scale ablations in the course. And so they do essentially with muP as we've described it. Just look at the right hand side which is the more relevant piece. You know, they're scaling down the variances. They're scaling down the learning rates by the width, the global width of the models M. And they're primarily keeping the depth fixed, which is a little bit of an unusual scaling regime because usually you'd see scale depth and width together, but they really want to do a controlled experiment where they're only looking at width variations, and they want to see if muP precisely nails scaling in this regime. There's also a little bit of a of a kind of weird subtlety that all of the muP papers seem to do which is that if you remember your CS224N lecture you know you remember that there's like a scaling on the attention activations you scale you know you do your inner product and then you scale it down by one over the square root of D you know and I told you kind of this was a magic constant that was like the right thing to do you know muP and other papers use one-over-d scaling instead of one over square root d for various arguments related to activation and update size stability. And so that's another thing that I think was
71:19 Worth pointing out because you might sort of not initially think of that as being something that's related to muP. Okay. Architecture is mostly similar to the standard you know transformer stuff. And as I already mentioned before, they only consider width scaling, right? So they take a standard transformer trained auto reggressively on, you know, pre-training text and they want to basically make the model wider and wider and wider on the MLPS and sort of the model residual stream dimensions. They're going to make that bigger and bigger and bigger. And what they want is for the optimum learning rate to remain the same as they scale the width up. And if it remains stable, then that's the big victory for muP, right? So the game is hopefully clear to everybody. You just want to scale with I want my learning rate that's optimal to stay the same. So question number one is you know does it work? Well the answer is yes. So we have different width 128 512 2048 we have different learning rates across the columns. And you know the sort of idealized strategy here is we run a sweep of learning rates at the small scale. I pick the smallest scale and I scale that up and hopefully that base learning rate remains optimal and yeah it seems like you know learning rates transfer very reliably across model sizes if we're doing this like you know somewhat precise with scaling and so then I think you start asking questions of all right very similar to the to the previous question that was just asked of like okay when does muP break right so you can ask that question in theory but you can also ask that question in practice, right? So, I'm just going to try all sorts of
72:59 Modern variations to architectures that people do and then I'm going to ask, you know, does this hyperparameter transfer thing continue to hold under these variations or not? And, you know, the paper is quite nice because they just go through a lot of different stuff. They'll vary the activations, they'll vary the batch sizes, the initializations, the RMS norm gains. They'll even use like really exotic optimizers like sort of sign gradient style stuff. And then they'll also vary the regularizers. So which one of these prevents learning rate transfer? So the first one you know which I think is probably relevant if you were kind of looking at that deep linear network and saying oh no one just multiplies matrices together there's like nonlinearities in between right so does muP work when we change nonlinearities around well SwiGLU squared ReLU and the baseline sort of muP approach with ReLU all have the same minimal learning rate. So no changes at all. You know we just see that for example SwiGLU and squared ReLU just do better than baseline muP. Unsurprising sort of agrees with a lot of what we've learned in the course right we might vary the batch sizes because we know that batch sizes are kind of going to be sensitive to scale like we've seen MiniCPM and we've seen DeepSeek basically fit scaling laws to batch sizes to try to get what the optimum batch size was. You know once again we see that you know as we scale up batch sizes by four up or down you know learning optimum learning rates remain stable. What about initializations right you know there are some initializations that you know people vary like for example some people set
74:37 The query matrix to zero so that all the different items get uniform attention maybe that's more stable. Some people sort of the unembedding layer at the very top they'll scale differently based on either you use standard parameterization or muP maybe that matters a lot turns out neither of those do you know the center column the optimum learning rate remains optimal in all of these cases you know what is it not robust to well you know it's not going to work for every single case for example if you add sort of learnable gains that turns out to break muP. Right? So you need to remove the biases. But you know if you if you remove them the muP works if you add them back in don't necessarily work. Similarly you can try sort of more exotic optimizers. Lion is an optimizer that takes like the sign of the gradient updates which to me feel a little bit crazy but I think this was searched this was found through like evolutionary search or something like this to find like the fastest optimizer. If you use this kind of a more crazy optimizer, it really breaks down. And I think this is what you expect, right? muP is designed to adapt to a very particular optimizer like AdamW to control the update sizes. So, you know, if you're using a totally different optimizer, I don't know why you'd expect the learning rates to transfer. So, maybe expected that this thing fails. And then sort of finally, what is it, you know, also not robust? Turns out if you really have much stronger weight decay, muP actually starts to fail. And so this is one of the few significant muP failures that are in there. A lot of the other ones are kind of just like, oh, we maybe expected that or that's not standard to
76:17 Do. You know, weight decay is something that you actually do. Okay, so muP seems generally useful like if you take standard parameterization like kind of going back to the baseline, right? You might ask like all right but what if I just do you know standard baseline stuff you know you can't use the same learning rate your the same learning rate results in you know significantly worse losses at 2048 right like your model just blows up gives you basically you know degenerate losses you would have been very sad scaling up at the same learning rate you know and we see also that the learning rate needs to scale down predictably as a function of the width on the other hand you know even if you scale up all the way to a 10b parameter model you know you see that the base loss remains the same so they do one large scale experiment and they see that the learning rate remains sort of ideal at the 2 to the negative6 level which is a kind of cool validation right so they do the whole study at a medium to small scale they do one big sort of hero run and then the learning rate remains optimal so the empirical results on that look somewhat promising the fact that meta used it for Llama 4 is also quite nice but as far as I know it's not a
Scaling in the wild: the closing checklist
77:22 Consensus that people use muP. Okay, so putting it all together you know how do you scale in the wild? I have never trained a you know 70B model at you know super-Chinchilla sizes and so we're going to have to rely a lot on case studies and we saw several examples of scaling in the wild. We saw people setting things like model hyperparameters especially learning rate and batch sizes using scaling laws. We saw you know people using things like muP or assume stability to try to avoid search over these spaces. And then also the use of things like alternative learning schedules like WSD can decrease the amount of compute that you need in order to fit a lot of these scaling laws. So, that's all I got.