Architectures, hyperparameters
Transcript: cleaned auto-captions with timestamps
Lecture 1 gave you tokenization, lecture 2 gave you the resource accounting. This lecture answers the question those two leave dangling: when you sit down to actually type a model definition, what do you type, and why is it that and not something else? The course's stated method is hands-on experience; this lecture is its explicit second-best — you cannot train 19 large models to find out which knobs matter, so you read 19 people's papers and treat their agreements as evidence and their disagreements as the map of what is actually open.
Outline, with timestamps
- 00:05 — Framing: the details other courses spare you, and why they matter.
- 01:46 — Two transformers: the 2017 original vs. the variant you implement in Assignment 1.
- 03:29 — The method: ~19 dense releases in a year, read as an evolutionary record.
- 05:40 — Pre-norm vs. post-norm: the one thing everybody agrees on.
- 09:11 — The "double norm": normalizing after the block too (Grok, Gemma 2, OLMo 2).
- 11:23 — LayerNorm → RMSNorm, and why 0.17% of the FLOPs is 25% of the runtime.
- 15:53 — Dropping bias terms everywhere.
- 19:17 — Activations and gating: ReLU → GeLU → SwiGLU, and the 2/3 sizing rule.
- 28:57 — Serial vs. parallel blocks.
- 32:15 — Position embeddings, and the construction that makes RoPE work.
- 40:33 — Hyperparameters that don't vary: FFN ratio, head budget, aspect ratio, vocabulary.
- 56:12 — Regularization: why weight decay survived and dropout didn't.
- 65:00 — Stability tricks: z-loss, QK-norm, logit soft-capping.
- 75:12 — Attention variants: MQA/GQA, KV-cache arithmetic intensity, sliding windows.
The method: nineteen models as a natural experiment
Most architecture advice you get is either a single paper's ablation on a 200M-parameter model, or somebody's confident vibe. Neither generalizes reliably to the scale you care about. Hashimoto's move is to sidestep both: build a spreadsheet with one row per released dense model from 2017 to 2025 and one column per design decision — norm placement, norm type, activation, position embedding, serial/parallel, vocabulary size, aspect ratio — and then look at where the columns are monochrome and where they are speckled (03:29).
This is an evolutionary argument, with the strengths and weaknesses of one. The strength: nobody spends eight figures on a pretraining run without some ablation behind each choice, so a monochrome column encodes a large amount of aggregate private evidence. The weakness: labs copy each other, so a column can go monochrome through imitation as easily as through discovery — and the lecture is honest that some of these choices have never been seriously tested at scale (49:30). Read the table holding both: default to consensus because it is cheap and safe, but know which items are backed by a published ablation and which by nothing but copy-paste.
"The theme of the class is the best way to learn is hands-on experience. But the theme of this lecture, because we can't train all these transformers, is to learn from the experience of others."— Tatsunori Hashimoto, 01:10
The practical payoff arrives immediately. The transformer you are asked to build in Assignment 1 is not the 2017 one: the norm moves in front of each block, position information comes from RoPE, the FFN is a SwiGLU, and every linear layer loses its bias (01:46). Four changes, four columns of the table that have gone monochrome. The rest of the lecture is the justification for each.
Normalization: the settled question, and the lesson hiding inside it
The 2017 transformer applied LayerNorm after each sublayer's residual add, putting a normalization operator directly on the residual path. Pre-norm moves it inside the branch, so the stream from embedding to output logits is a clean sum of identities and branch outputs, untouched (06:15). Asked why this is better, Hashimoto declines to claim a proof and gives the honest intuition: an unobstructed identity path is what makes gradients propagate through very deep stacks, and a normalizer in the middle interferes with exactly that (10:16). The original selling point in Nguyen & Salazar and Xiong et al. was that pre-norm lets you delete learning-rate warmup; the modern one is different — pre-norm is what keeps a large run from spiking and what lets you use a learning rate worth using (07:24). The lone exception in the survey is OPT-350M, post-norm apparently by accident.
The genuinely new development since the 2024 edition of this lecture is the "double norm": if the objection to post-norm is that it sits on the residual stream, then a norm placed after the branch but still outside the stream is not objectionable at all, and you can have both. Grok and Gemma 2 normalize before and after each sublayer; OLMo 2 does only the non-residual post-norm (09:43). Reports are that it is a bit more stable at large scale.
The type of norm is the other consensus. RMSNorm drops the mean subtraction and the learned shift, keeping only y = x / sqrt(mean(x²) + ε) · γ. Every argument for it is a systems argument, and this is where the lecture's most portable idea lives. If you have absorbed that matrix multiplies are essentially all of the FLOPs, you would predict norms are irrelevant to runtime, and the profiling in Ivanov et al. agrees on the first half: tensor contractions are about 99.8% of the FLOPs, normalization plus softmax about 0.17%. But that 0.17% of the arithmetic is roughly 25% of the wall-clock, because these are memory-bound elementwise passes over activations (14:12). Fewer operations and one fewer parameter tensor to stream is therefore a real, if modest, win — Narang et al. measured 3.5 → 3.68 steps/sec and a slightly better final loss (15:19).
"Even though tensor contraction … that's like 99.8% of the flops. If you have things like the softmax operation or LayerNorms … they're 0.17% of the flops — actually, they're 25% of the runtime."— Tatsunori Hashimoto, 14:12
The same reasoning, pushed one step further, deletes the bias terms from every linear layer and from the norm: FFN(x) = σ(xW₁)W₂. They cost parameters to move and buy nothing measurable in quality, and — a claim Hashimoto flags as empirically solid but poorly understood — removing them tends to stabilize the largest runs (16:28). Cohere's Command A and Command R+ are the notable holdouts still using LayerNorm.
Activations and gating
The zoo (ReLU, GeLU, Swish, ELU, GLU, GeGLU, ReGLU, SeLU, SwiGLU, LiGLU) is the part of deep learning Hashimoto openly says he did not want to have to care about, and then does (19:17). The arc is short: the original transformer and T5 used ReLU; the GPT line used GeLU (GELU(x) = x·Φ(x), a ReLU with a soft, slightly non-monotone knee); essentially everything after 2023 uses a gated unit.
Gating is a one-line change with a real idea in it. Instead of FF(x) = max(0, xW₁)W₂, you add a second up-projection and multiply elementwise: FF_ReGLU(x) = (max(0, xW₁) ⊗ xV)W₂. Swap the nonlinearity and you get GeGLU (Google's line: T5 v1.1, mT5, LaMDA, Phi-3, Gemma 2/3) or SwiGLU with swish(x) = x·sigmoid(x) (LLaMA 1/2/3, PaLM, Mistral, OLMo, most models post-2023). The lecture files gating alongside residual connections and normalization as one of the few structural motifs that keep re-earning their place across architectures (21:30).
The bookkeeping detail everyone trips over: the extra V matrix means a gated FFN has three weight matrices where the ungated one had two, so to stay parameter-matched you shrink the hidden width by 2/3. Starting from the standard d_ff = 4·d_model, that lands you at d_ff = (8/3)·d_model ≈ 2.67·d_model, which is exactly why LLaMA 1, Qwen 14B, DeepSeek 67B and Yi 34B all sit near 2.68 (24:55).
Evidence is unusually good here by 2020s standards. Shazeer's GLU-variants paper sweeps them all on GLUE-style tasks and reports standard deviations, and the gated variants win consistently; Narang et al.'s much larger survey of transformer modifications finds the same, with the gated rows the ones that survive (26:39). Hashimoto is careful to separate "better" from "necessary": GPT-3 has no gating, Nemotron-4 340B uses a squared ReLU, Falcon 2 11B uses plain ReLU, and all three are strong models (28:21).
The last architectural variant is parallel blocks: instead of x → attn → mlp in series, compute x + attn(x) + mlp(x) with both branches from the same input. GPT-J introduced it, PaLM ran it at 540B, and Cohere Command A, Command R+ and Falcon 2 11B still use it. Done right you share the norm and fuse the branches' matmuls into wider ones — a systems win, at the plausible cost of expressiveness, since you are adding two computations rather than composing them (30:03). It has not spread: most of the last year's models are serial, and there is no serious public ablation either way.
Position: the criterion that picks RoPE
Position embeddings are the column where the table is most interesting historically and most boring now. Sinusoids (original transformer), learned absolute vectors (GPT-1/2/3, OPT), learned relative biases added into the attention logits (T5, Gopher, Chinchilla), a brief ALiBi phase — and then from about 2023 onward, RoPE, everywhere, in all 19 surveyed models (32:15).
What makes the derivation worth following is that it starts from a criterion, not a trick. You want an embedding f(x, i) such that ⟨f(x,i), f(y,j)⟩ = g(x, y, i−j): attention is an inner product, so demand the inner product see only the relative offset (33:24). Every predecessor fails specifically. Additive sinusoids expand into cross terms like ⟨PE_i, v_y⟩ that leak absolute position; learned absolute embeddings obviously fail; T5-style relative biases are relative but are added to logits rather than realized as an inner product, so they violate the form.
The construction falls out of one fact: inner products are invariant under a shared rotation. Rotate every token's vector by an angle proportional to its position, and two tokens two apart keep the same angle between them wherever they sit — so the inner product sees only the gap (35:00). In d dimensions "which rotation" is underdetermined, and RoPE takes the simplest usable answer: pair up coordinates and rotate each 2-D block by m·θ_k, with a θ_k schedule spanning fast and slow frequencies exactly as sinusoids do, so some pairs encode nearby structure and others long-range (36:41). The θs are fixed, not learned — which is why the rotations pose no optimization difficulty; a fixed rotation is just a fixed matmul (40:33).
The implementation detail that catches people: RoPE is not applied once at the bottom of the network. It is applied to the queries and keys inside every attention layer, right after the QKV projections and before the dot product (37:49). That is what enforces the relative-only property; add it at the embedding layer and later layers mix absolute information back in. Two things then explain the sweep: RoPE works well even at small scale and short context, and — more decisively for production — a rotation schedule is a continuous knob, so an entire family of context-extension methods exists for rescaling it after training (38:56).
The hyperparameters nobody varies
Four numbers, four consensuses, and one repeat offender.
| FFN ratio | d_ff = 4·d_model ungated, 8/3 gated. Kaplan et al. sweep it and find a broad basin from about 1 to 10 where loss is near-optimal, so 4 is a safe interior point rather than a magic value (45:32). |
| Head budget | n_heads · d_head = d_model — adding heads splits the same width rather than growing it. Nearly universal; T5 is the outlier at 16× (48:21). |
| Aspect ratio | d_model / n_layers between roughly 100 and 200 — GPT-3/OPT/Mistral/Qwen at 128, LLaMA and Chinchilla at 102, PaLM 540B at 156, BLOOM at 205, GPT-2 at 33 and T5-11B at 43 as the shallow-and-wide old guard. |
| Vocabulary | 30–50k for monolingual (GPT-2/3 at 50257, LLaMA at 32000), 100–250k for multilingual and production (GPT-4 at 100276, PaLM at 256000, Command A at 255000, mT5 at 250000) (54:31). |
The head-budget consensus is the weakest of the four, and the lecture says so. Bhojanapalli et al. argued that slicing d_model across many heads makes each head's attention matrix low-rank enough to bottleneck expressiveness — a real theoretical concern that practice has simply not confirmed (49:30). Treat it as "held constant by convention," not as validated.
Aspect ratio is the one where systems constraints, not loss, do the deciding. Pipeline parallelism cuts the model by layers and tolerates slower interconnect; tensor parallelism cuts each matrix and demands very fast interconnect — so your network topology has an opinion about deep vs. wide (50:37). On the modelling side, Kaplan's sweep puts the loss optimum near 100 and, usefully, shows it barely moves across 50M, 274M and 1.5B parameters — fix it once and scale (52:54). Tay et al. add the caveat: at equal parameters depth barely changes pretraining loss, but at equal FLOPs deeper models looked better on downstream fine-tuned accuracy. Loss and downstream are not the same objective (53:59).
T5 is the recurring rule-breaker and the lecture enjoys it: the 11B model sets d_model = 1024 against d_ff = 65536, a 64× multiplier, chosen to get very fat, hardware-friendly matmuls (44:27). It worked — proof that none of these numbers is a law of nature. But T5 v1.1 quietly walked it back to a standard 2.5 multiplier on GeGLU and came out better, about as clear a retraction as the literature ever prints (46:42). Gemma 2 at 8× and Gemma 3 / SmolLM at 4× show the space is still being probed.
Regularization that isn't regularizing
Pretraining is roughly the least overfitting-prone setting in machine learning: a single pass, over more tokens than you have parameters, on data you will never see twice. By that logic you should not regularize at all. The table says otherwise — older models used dropout 0.1 (original transformer, GPT-2/3, T5, OPT), newer ones dropped it, but almost everyone still applies weight decay (57:50).
Andriushchenko et al. supply the resolution, and it is genuinely counterintuitive. Varying weight decay does not move the train-to-validation gap; there is no overfitting for it to control. What it does is interact with the learning-rate schedule. Under a cosine decay, a high-weight-decay run trains visibly worse while the learning rate is high, then optimizes very rapidly during the cool-down and finishes ahead (59:28). Weight decay in an LLM is an optimizer intervention wearing a regularizer's name.
"You don't weight decay because you want to regularize the model, which is kind of what it was designed for. You're weight decaying in order to get actually better training losses."— Tatsunori Hashimoto, 60:01
Dropout's disappearance has no equally clean explanation — Hashimoto says he has not seen an analysis showing it helps training loss, and there is no overfitting story to justify it. Qwen 14B is the notable model still reporting it. A caveat he flags twice: on dropout in particular, papers frequently just don't say, and open models' silence has usually meant zero, but closed models may differ.
Stability and inference: where the last year's work actually went
Asked what is new since the previous edition of the lecture, the answer is not the core architecture — it is stability. OLMo 2's paper opens with the motivating picture: an acceptable-looking loss curve sitting above a gradient-norm curve that is a forest of spikes, the kind of run that eventually explodes and cannot be resumed (66:06). Instabilities can come from anywhere, but the interventions cluster on one culprit — softmax, which exponentiates and then divides, and appears twice: over the vocabulary at the output, and over positions inside attention (67:46).
For the output softmax the fix is the z-loss, traceable to Devlin et al.'s self-normalizing MT models in 2014. Add α·log²Z(x) to the objective, pushing the partition function toward 1; if it succeeds, the log and the exponential cancel and the numerically dangerous operation effectively disappears (68:22). PaLM brought it into language modelling with α = 1e-4; Baichuan 2, DCLM and OLMo 2 followed (70:02).
For the attention softmax the fix is QK-norm: normalize queries and keys before the dot product. Note the different theory of the problem — z-loss controls the normalizer, QK-norm bounds the softmax's inputs so the exponentials can never get large in the first place (71:11). It came out of vision and multimodal work (Dehghani et al.'s 22B ViT, then Chameleon and Idefics) and back-propagated into text models: Gemma 2, DCLM, OLMo 2. Hashimoto's one joke of the lecture is the summary — LayerNorm keeps working everywhere you put it (72:21). Logit soft-capping, cap·tanh(logits/cap), is the third option (Gemma 2, OLMo 2) and the weakest: in the NVIDIA stability study's table, soft-capping made perplexity worse relative to the baseline while QK-norm improved it, because QK-norm lets you push the learning rate harder (74:04). And yes, QK-norm stays on at inference — it has learned scale parameters, so removing it is a different model (74:37).
The final segment is the course's clearest example so far of inference economics reshaping training-time architecture. At training time, attention has arithmetic intensity on the order of (1/k + 1/(bn))⁻¹ — batch and sequence are both large, so the GPU stays fed. At decode time you produce one token at a time, the big matmuls become skinny matrix–vector products, and with a KV cache the intensity degrades to about (n/d + 1/b)⁻¹ (81:19). That n/d term is the whole problem: the only ways to shrink it are a shorter context or a bigger model, and you want neither. MQA attacks it structurally — many query heads, a single shared key/value head, so the cache you stream per step shrinks by a factor of the head count (82:24). GQA is the same idea with a dial: share K/V across groups of query heads, trading expressiveness against cache bandwidth. Shazeer measured a small perplexity cost for full MQA; Ainslie et al. found GQA close to free (83:34).
The long-context story ends the lecture. Sparse and strided patterns go back to Child et al. and GPT-3; sliding-window attention (Mistral) restricts each layer to a local neighbourhood and lets depth grow the effective receptive field (84:40). The 2025 synthesis, in Cohere's Command A and in LLaMA 4 and Gemma, is an interleave: three sliding-window layers with RoPE for local structure, then one full-attention layer with no position embedding at all. Cheap, because full attention runs a quarter of the time — and it extrapolates, because the layer that sees the whole sequence has no positional encoding to extrapolate (85:46). That is how a model advertises 10M tokens of context.
What you build with this
No new assignment starts here. This lecture is the retroactive justification for Assignment 1: Basics (handout PDF), whose model-definition section asks for exactly the four departures from the 2017 transformer surveyed above: pre-norm placement, RMSNorm without bias, a SwiGLU feed-forward sized at 8/3·d_model, and RoPE applied to Q and K inside each attention layer. If your RMSNorm or RoPE implementation fails the provided tests, the relevant explanations are at 11:23 and 37:49 respectively — the RoPE test in particular fails for people who apply the rotation at the embedding layer rather than per-attention-layer. The hyperparameter section is what makes the assignment's suggested configuration non-arbitrary, and the leaderboard is where you find out whether deviating from those defaults on a fixed compute budget was a good idea.
Looking forward: the arithmetic-intensity argument at 77:56 is the setup for the GPU lecture (L5) and Assignment 2, and the parallelism constraints on aspect ratio at 50:37 are the setup for the distributed-training lectures. The one architectural direction deliberately deferred is sparsity — the next lecture takes up mixture-of-experts and DeepSeek V3.
Supporting materials, verified
- Attention Is All You Need — Vaswani et al. (2017) · The baseline the whole lecture measures against: post-norm LayerNorm, sinusoidal positions, ReLU FFN, biases everywhere. Every one of those four has since been replaced.
- Transformers without Tears: Improving the Normalization of Self-Attention — Nguyen & Salazar (2019) · The pre-norm case, framed as removing warmup. The slides render the authors as "Salazar and Ngyuen"; the paper is Nguyen & Salazar, IWSLT 2019.
- On Layer Normalization in the Transformer Architecture — Xiong et al. (2020) · The gradient-attenuation analysis behind the pre-norm figures at 07:24.
- Root Mean Square Layer Normalization — Zhang & Sennrich (2019) · The RMSNorm definition itself; field map extra, since the lecture uses the equation without citing the source.
- Data Movement Is All You Need: A Case Study on Optimizing Transformers — Ivanov, Dryden, Ben-Nun, Li & Hoefler (2020; MLSys 2021) · Source of the 99.8%-of-FLOPs / 25%-of-runtime split. The slide labels it "Ivanov et al 2023" and the lecturer recalls the title as "Memory Movement…"; both are slips for this paper.
- Do Transformer Modifications Transfer Across Implementations and Applications? — Narang et al. (2021; EMNLP) · The RMSNorm step-time and loss numbers, and independent corroboration of the gated-FFN result. Cited on the slides as 2020; the arXiv posting is February 2021.
- GLU Variants Improve Transformer — Shazeer (2020) · The original GeGLU/SwiGLU/ReGLU sweep, with standard deviations. This is the paper that made the 2/3 sizing convention standard.
- RoFormer: Enhanced Transformer with Rotary Position Embedding — Su et al. (2021) · RoPE. Read §3 for the relative-inner-product criterion the lecture reconstructs at 33:24.
- Scaling Laws for Neural Language Models — Kaplan et al. (2020) · Cited here not for the scaling laws but for the hyperparameter appendices: the d_ff/d_model basin and the aspect-ratio sweep across three model scales.
- Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers — Tay et al. (2021) · Depth vs. width, and the observation that pretraining loss and downstream accuracy disagree about it.
- Low-Rank Bottleneck in Multi-head Attention Models — Bhojanapalli et al. (2020) · The argument against n_heads·d_head = d_model that practice has ignored.
- Why Do We Need Weight Decay in Modern Deep Learning? — Andriushchenko et al. (2023) · The weight-decay-as-optimizer-intervention result; the source of the constant-vs-cosine learning-rate figures.
- Fast and Robust Neural Network Joint Models for Statistical Machine Translation — Devlin et al. (2014) · Self-normalization, the ancestor of the z-loss, eight years before PaLM.
- PaLM: Scaling Language Modeling with Pathways — Chowdhery et al. (2022) · Where z-loss, parallel blocks, SwiGLU and a 256k vocabulary all appear in one 540B model.
- 2 OLMo 2 Furious — OLMo team (2025) · The lecture's main modern source on stability: the gradient-norm figure, non-residual post-norm, QK-norm and z-loss all in one open, reproducible report.
- Scaling Vision Transformers to 22 Billion Parameters — Dehghani et al. (2023) · Where QK-norm comes from, before it crossed back into text models.
- Methods of Improving LLM Training Stability — Rybakov et al., NVIDIA (2024) · The uncited slide at 74:04: the head-to-head where QK-norm beats the baseline and soft-capping does not.
- Fast Transformer Decoding: One Write-Head is All You Need — Shazeer (2019) · MQA, and the KV-cache bandwidth argument for it.
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Query Checkpoints — Ainslie et al. (2023) · GQA, including how to convert an existing MHA checkpoint rather than retrain.
- Generating Long Sequences with Sparse Transformers — Child et al. (2019) · The strided/local factorized attention patterns GPT-3 used.
- Mistral 7B — Jiang et al. (2023) · Sliding-window attention in a shipped model, plus GQA; the receptive-field-grows-with-depth argument.
- Command A: An Enterprise-Ready Large Language Model — Cohere (2025) · The interleaved 3×(sliding-window + RoPE) : 1×(full attention, no position embedding) pattern, and one of the few recent models still on LayerNorm and parallel blocks.
- Gemma 2: Improving Open Language Models at a Practical Size — Gemma team (2024) · Field map extra: double norm, GeGLU, logit soft-capping and interleaved local/global attention in one report — the single best cross-check on this lecture's stability section.
- DataComp-LM: In Search of the Next Generation of Training Sets for Language Models — Li et al. (2024) · Field map extra: DCLM, cited here for adopting both z-loss and QK-norm, and worth reading ahead of the data lectures.
Exercises
- Derive and verify the 2/3 rule code — (1) Write the parameter count of an ungated FFN with d_ff = 4·d_model and of a gated one with three matrices. (2) Solve for the gated d_ff that matches it; confirm you get 8/3·d_model. (3) Implement both in PyTorch at d_model = 1024 and assert the parameter counts agree to within the rounding your implementation does. (4) Look up d_model and intermediate_size in the HuggingFace configs for LLaMA 1 7B, Mistral 7B and Qwen 14B and compute the ratios. A good answer explains why Mistral's 3.5 is not a violation of the rule but a deliberate choice to spend more parameters in the FFN, and notes that real models round d_ff to a multiple convenient for their kernels.
- Test RoPE's defining property code — (1) Implement RoPE as a function f(x, i) over a d = 64 vector. (2) Draw random x, y and assert that ⟨f(x,i), f(y,j)⟩ is equal for all pairs with the same i−j, to floating-point tolerance. (3) Repeat with sinusoidal additive embeddings and show the equality fails; identify which cross term is responsible. (4) Now apply your RoPE once at the embedding layer of a two-layer toy transformer instead of inside each attention layer, and re-run the test on the layer-2 queries. A good answer reports that the property survives the first attention layer and dies after the residual add and MLP, which is precisely why the lecture insists RoPE goes inside every attention layer.
- Find the 25% code — (1) Build one pre-norm transformer block (RMSNorm, attention, SwiGLU) and run it on a realistic batch. (2) Count analytic FLOPs per operator and confirm the matmuls are ≳99% of them. (3) Profile with torch.profiler and record CUDA time per operator. (4) Compare the FLOP share and the time share for the norms and softmax. (5) Swap RMSNorm for LayerNorm and re-measure. A good answer shows the two rankings disagreeing, attributes the gap to memory traffic rather than arithmetic, and states the elementwise-pass bytes moved per norm — this is the empirical version of the callout above, and the number you will re-derive properly in lecture 5.
- Extend the survey by one row — Pick a dense model released after this lecture and fill in its row: norm placement and type, activation, position embedding, serial/parallel, d_ff/d_model, n_heads·d_head vs. d_model, d_model/n_layers, vocabulary size, KV-head count, attention pattern, and whether it reports z-loss, QK-norm or soft-capping. Work from the config file and the technical report, not from a blog summary. A good answer names every cell where the model departs from this lecture's consensus, quotes what the report says about why, and flags at least one cell the report simply does not disclose — the disclosure gaps are themselves a finding, and are why Hashimoto repeatedly hedges that closed models may not match the open pattern.