Assembling the block: heads, feed-forward, residuals, LayerNorm
Transcript: this part, with timestamps
Part 5 ended with a self-attention head that works in a notebook cell. That is not yet a model. The lecture's next move is deliberately unglamorous engineering: take the head, put it in the network, and then keep adding the smallest piece that unblocks the next improvement — until the thing on screen is, line for line, the decoder half of the 2017 paper. Nothing here is a new idea about attention. Everything here is about making a stack of attention layers actually optimize.
Outline, with timestamps
- 79:11 — Inserting a single head: the Head module, tril as a buffer, wired between the embeddings and the output layer.
- 80:49 — Cropping the context in generate: positional embeddings make block_size a hard ceiling for the first time.
- 81:20 — First result: lower the learning rate, run longer, 2.5 → 2.4.
- 81:59 — Multi-head attention: four heads of 8 dimensions, concatenated back to 32.
- 82:53 — Why four narrow channels beat one wide one; the group-convolution analogy. 2.28.
- 84:25 — Feed-forward layers: the paper's position-wise MLP, and why the tokens needed one. 2.24.
- 86:26 — The block: interleave communication and computation, then repeat it three times — and watch it fail to improve.
- 86:48 — Residual connections, from the 2015 ResNet paper.
- 89:03 — The residual pathway picture: fork off, compute, add back; addition as a gradient distributor.
- 90:35 — Projections into the stream, and the 4× inner dimension of the MLP. 2.08, and the first overfitting.
- 92:51 — LayerNorm and its relationship to BatchNorm: the same code, normalizing rows instead of columns.
- 95:12 — Pre-norm: the one place this implementation knowingly departs from the paper, plus the final ln_f. 2.06.
One head, wired into the network
The head from part 5 becomes an nn.Module with three bias-free linear layers — key, query, value — each mapping n_embd → head_size. Bias is off because these are pure projections; there is nothing for a per-feature offset to do that the layer they feed into cannot also do.
The one piece of PyTorch trivia worth internalizing is register_buffer. The lower-triangular mask is a constant, so it is not a Parameter — the optimizer must not touch it. But it still has to move with .to(device) and show up in state_dict. A buffer is exactly that: tensor state owned by the module, excluded from gradients. Registering it as a plain attribute would leave it on the CPU the moment you train on a GPU.
class Head(nn.Module):
""" one head of self-attention """
def __init__(self, head_size):
super().__init__()
self.key = nn.Linear(n_embd, head_size, bias=False)
self.query = nn.Linear(n_embd, head_size, bias=False)
self.value = nn.Linear(n_embd, head_size, bias=False)
self.register_buffer('tril', torch.tril(torch.ones(block_size, block_size)))
self.dropout = nn.Dropout(dropout)
def forward(self, x):
# input of size (batch, time-step, channels)
# output of size (batch, time-step, head size)
B,T,C = x.shape
k = self.key(x) # (B,T,hs)
q = self.query(x) # (B,T,hs)
# compute attention scores ("affinities")
wei = q @ k.transpose(-2,-1) * k.shape[-1]**-0.5 # (B, T, hs) @ (B, hs, T) -> (B, T, T)
wei = wei.masked_fill(self.tril[:T, :T] == 0, float('-inf')) # (B, T, T)
wei = F.softmax(wei, dim=-1) # (B, T, T)
wei = self.dropout(wei)
# perform the weighted aggregation of the values
v = self.value(x) # (B,T,hs)
out = wei @ v # (B, T, T) @ (B, T, hs) -> (B, T, hs)
return out
Walk the shapes once. Input (B, T, C). Keys and queries come out (B, T, hs); transposing the last two dimensions of k gives (B, hs, T), so the matmul yields one T×T affinity matrix per batch element: (B, T, T). masked_fill writes -inf wherever the triangle is zero — that is row t, column t' with t' > t, the future — and softmax over dim=-1 turns each row into a distribution over the past that sums to 1, with the -inf entries going to exactly 0. Then wei @ v contracts the T axis away and returns (B, T, hs). The slice self.tril[:T, :T] is not decoration: at generation time the context is shorter than block_size, and the mask has to be cropped to match.
Two knock-on changes. First, generate now has to crop its input — idx_cond = idx[:, -block_size:] — because the position embedding table only has block_size rows and will index out of range on token 9 with a context window of 8. The bigram model had no such constraint; it never looked backwards at all. Second, the learning rate has to come down. The bigram script trains happily at 1e-2; with attention in the graph that diverges, so Karpathy lowers it and raises max_iters to compensate (81:20). Validation loss moves 2.5 → 2.4. It is a real improvement and an unimpressive one — a single 32-dimensional communication channel is not much of an architecture.
Four narrow heads instead of one wide one
MultiHeadAttention is four lines: build a ModuleList of num_heads heads, run them all on the same x, and concatenate the outputs along the channel dimension. The arithmetic that makes it invisible from the outside is head_size = n_embd // n_head. At this point in the video n_embd is 32 and n_head is 4, so each head produces (B, T, 8) and the concatenation restores (B, T, 32). In the final gpt.py the same identity reads 384 / 6 = 64.
Notice what this does not cost. Four heads of width 8 have exactly as many projection parameters as one head of width 32, and the aggregated value vector is the same size either way. What you buy is four independent T×T attention matrices instead of one. Each can specialize: one head can be a "find the nearest vowel" circuit, another a "what happened two characters ago" circuit, and they do not have to share a single softmax. Karpathy compares it to a grouped convolution — one big operation split into independent groups over channels (83:23). Take that as an analogy about the grouping, not a claim about convolution.
Result: 2.28. He is careful not to oversell it — the samples still read like noise. The number is the evidence, not the text.
Communication, then computation
Up to here the model attends and then immediately computes logits. Every token gathers information from its past and then, with no further processing, has to commit to a prediction. That is the gap the feed-forward layer fills:
the tokens looked at each other but didn't really have a lot of time to think on what they found from the other tokensKarpathy, 85:26
The fix is the paper's "position-wise feed-forward network," which is a two-layer MLP applied to each position separately: Linear(C, 4C) → ReLU → Linear(4C, C). The word doing the work is position-wise. nn.Linear only ever touches the last dimension, so a (B, T, C) tensor is processed as B×T independent vectors — no information crosses the time axis. That is worth stating plainly and remembering: in the entire Transformer, attention is the only operator that moves information between positions. Everything else — the MLP, the norms, the output layer — is a per-token map.
The first version he trains has no 4× widening at all, just a linear and a ReLU at width 32, and it gets 2.24 (86:26). The 4× multiplier arrives later, at 91:36, read straight off the paper's numbers: d_model 512, inner dimension 2048. If you are matching the video against the repo, the class is spelled FeedFoward — the typo is in the real file, and copying it is how you keep imports working.
Now the pattern is visible, and the Block just names it: multi-head attention (communication) followed by a feed-forward net (computation), packaged so you can stack it. nn.Sequential(*[Block(n_embd, n_head=n_head) for _ in range(n_layer)]), three deep at this stage.
Depth stops working, and residuals fix it
Stacking three blocks does not help. Karpathy is honest that he is not diagnosing this so much as recognizing the shape of it — the network is now deep enough that deep-network optimization problems are plausibly what he is hitting (88:01). He does not measure gradient norms or show a vanishing-gradient plot; he reaches for the standard fix. If you want the rigorous version, that is the ResNet paper's job, not this lecture's.
The fix is the skip connection from He et al. 2015, and the mental model he gives for it is the one worth keeping. Picture a single wide bus of activations — the residual stream — running from the embeddings at the top to the logits at the bottom, always shape (B, T, C). Every sublayer forks off the bus, computes something, and adds its result back. The bus itself is never multiplied, never normalized, never reshaped. In backward mode, addition hands the incoming gradient to both of its inputs unchanged, so the gradient from the loss reaches the token embedding table without passing through a single weight matrix:
you have this gradient superhighway that goes directly from the supervision all the way to the input, unimpededKarpathy, 89:34
In code this is two characters of change per sublayer — x = x + self.sa(...) instead of x = self.sa(x) — plus one new module each. MultiHeadAttention gains self.proj = nn.Linear(head_size * num_heads, n_embd), and FeedFoward's second linear plays the same role. This is the part people misread: at these settings head_size * num_heads equals n_embd exactly, so the projection is 32→32 and changes no shape whatsoever. It is there because writing into a shared stream is its own job — the sublayer computes in its own space, and the projection decides how that result gets expressed in the stream's basis, including scaling it down toward nothing if that is what training wants.
Karpathy adds that residual branches are typically initialized to contribute almost nothing, so the network starts out close to the identity and the blocks "come online over time." That is true of good practice but is not what gpt.py does — its _init_weights gives every nn.Linear the same std=0.02, and the file's own comment flags better init as "not covered in the original GPT video." nanoGPT does implement it, scaling every c_proj weight by 1/sqrt(2 * n_layer) per the GPT-2 paper (model.py L142–L145). Worth knowing so you do not go looking for it in the lecture's file.
With residuals, projections and the 4× MLP together, validation loss drops to 2.08 — and for the first time training loss pulls ahead of validation loss. That gap is the reason dropout shows up in part 7.
LayerNorm, and the pre-norm swap
The second depth trick is the "Norm" in the paper's "Add & Norm." Karpathy derives LayerNorm by editing the BatchNorm he wrote in an earlier video: change the reduction dimension from 0 to 1 and you are done (94:10). BatchNorm normalizes each feature across the batch — the columns of a (32, 100) matrix. LayerNorm normalizes each example across its features — the rows. Same formula, transposed intent.
The consequences all fall out of that one change. Because the statistics never cross examples, there is no running mean or variance to maintain, no momentum, and no difference between training and inference — all the machinery that makes BatchNorm fiddly disappears. What survives is the learnable gain and bias, so the output is unit-Gaussian only at initialization; training is free to move it, and generally does.
In the model it is nn.LayerNorm(n_embd) applied to a (B, T, C) tensor. The argument tells it to reduce over the last dimension only, which at this stage means the mean and variance are computed over 32 numbers, per token. B and T both act as batch dimensions (96:13) — token 3 of sequence 0 is normalized entirely on its own, using nothing from token 4 and nothing from sequence 1. It is a per-token rescaling of the feature vector, nothing more.
The placement is the interesting part, and it is the one deliberate departure from the 2017 paper in the whole build. Figure 1 puts the norm after the sublayer and the addition: x = LayerNorm(x + sublayer(x)). What everyone actually does now, and what gpt.py does, is normalize the input to the sublayer and leave the stream untouched: x = x + sublayer(LayerNorm(x)). This is the pre-norm formulation. It preserves the property the residual stream was introduced for — the identity path from loss to embeddings has nothing sitting on it — which is why pre-norm Transformers train without a learning-rate warmup and post-norm ones typically need one (Xiong et al. 2020 is the analysis; the video predates leaning on it).
Pre-norm has one obligation attached: since nothing normalizes the stream itself, the last block's output arrives at the output layer unnormalized. Hence self.ln_f, a final LayerNorm between the block stack and lm_head, which Karpathy adds as an afterthought at 97:17. It is not optional — every real implementation has it.
Adding the norms takes 2.08 to 2.06. He flags that honestly as slight, and expects the payoff to grow with depth and width. Part 7 is where that bet gets tested.
The ladder of losses
| Change | Val loss | Timestamp |
|---|---|---|
| Where part 5 left off (embeddings + positions, no attention in the script) | ~2.5 | — |
| One self-attention head, head_size = n_embd = 32 | 2.4 | 81:20 |
| Multi-head: 4 heads × 8 dims, concatenated | 2.28 | 83:53 |
| + per-token feed-forward (plain 32 → 32 + ReLU) | 2.24 | 86:26 |
| 3 blocks stacked, no residuals | no improvement | 88:01 |
| + residual connections, projections, 4× MLP inner dim | 2.08 (train < val) | 92:06 |
| + pre-norm LayerNorms and final ln_f | 2.06 | 96:46 |
The numbers, here versus the finished script
| Symbol | In the video during this part | Final gpt.py |
|---|---|---|
| batch_size (B) | 32 | 64 |
| block_size (T) | 8 | 256 |
| n_embd (C) | 32 | 384 |
| n_head | 4 | 6 |
| head_size = n_embd // n_head | 8 | 64 |
| n_layer | 3 | 6 |
| MLP inner dimension | 4 × 32 = 128 | 4 × 384 = 1536 |
| learning_rate | lowered from the bigram's 1e-2 | 3e-4 |
| dropout | not yet | 0.2 (part 7) |
The code at the end of this part
class MultiHeadAttention(nn.Module):
""" multiple heads of self-attention in parallel """
def __init__(self, num_heads, head_size):
super().__init__()
self.heads = nn.ModuleList([Head(head_size) for _ in range(num_heads)])
self.proj = nn.Linear(head_size * num_heads, n_embd)
self.dropout = nn.Dropout(dropout)
def forward(self, x):
out = torch.cat([h(x) for h in self.heads], dim=-1)
out = self.dropout(self.proj(out))
return out
class FeedFoward(nn.Module):
""" a simple linear layer followed by a non-linearity """
def __init__(self, n_embd):
super().__init__()
self.net = nn.Sequential(
nn.Linear(n_embd, 4 * n_embd),
nn.ReLU(),
nn.Linear(4 * n_embd, n_embd),
nn.Dropout(dropout),
)
def forward(self, x):
return self.net(x)
class Block(nn.Module):
""" Transformer block: communication followed by computation """
def __init__(self, n_embd, n_head):
# n_embd: embedding dimension, n_head: the number of heads we'd like
super().__init__()
head_size = n_embd // n_head
self.sa = MultiHeadAttention(n_head, head_size)
self.ffwd = FeedFoward(n_embd)
self.ln1 = nn.LayerNorm(n_embd)
self.ln2 = nn.LayerNorm(n_embd)
def forward(self, x):
x = x + self.sa(self.ln1(x))
x = x + self.ffwd(self.ln2(x))
return x
Line by line against the repo. L92–L104, MultiHeadAttention: nn.ModuleList (not a plain Python list) so the heads' parameters are registered; torch.cat(..., dim=-1) glues num_heads tensors of (B, T, head_size) into (B, T, num_heads*head_size); self.proj is the write into the residual stream. L106–L119, FeedFoward: widen 4×, ReLU, narrow back — the second nn.Linear is this sublayer's projection, folded into the Sequential rather than given its own attribute. L121–L136, Block: head_size = n_embd // n_head is the invariant that lets the concatenation land back on n_embd; two separate LayerNorms because each has its own gain and bias; and the two x = x + ... lines with the norm inside the call are the pre-norm formulation. The nn.Dropout calls are part 7's addition — delete them and you have exactly the model this part ends on.
And the spine that consumes it, L160–L169:
def forward(self, idx, targets=None):
B, T = idx.shape
# idx and targets are both (B,T) tensor of integers
tok_emb = self.token_embedding_table(idx) # (B,T,C)
pos_emb = self.position_embedding_table(torch.arange(T, device=device)) # (T,C)
x = tok_emb + pos_emb # (B,T,C)
x = self.blocks(x) # (B,T,C)
x = self.ln_f(x) # (B,T,C)
logits = self.lm_head(x) # (B,T,vocab_size)
Five lines, and the shape never changes until the last one. pos_emb is (T, C) and broadcasts across the batch when added to (B, T, C). self.blocks is the nn.Sequential stack; every block preserves (B, T, C), which is why they compose at all. ln_f is the pre-norm tax being paid. Only lm_head changes the last dimension, from n_embd to vocab_size = 65.
Where people get stuck
- The scaling constant: C or head_size? On screen at 80:05 the head divides by C**-0.5, and Karpathy pins a correction in the description saying it should be head_size. It is easy to miss because at that exact moment the single head has head_size == n_embd == 32 and the two are the same number. They diverge the second multi-head arrives (8 vs 32), and the repo has it right: wei = q @ k.transpose(-2,-1) * k.shape[-1]**-0.5. The variance argument from part 5 is about the length of the vectors being dotted, which is head_size.
- Why does self.proj exist if it changes no shape? Because it is not a shape fix. head_size * num_heads == n_embd by construction, so the projection is square. Its job is to let the sublayer choose how to write into the residual stream, and in particular to be able to write almost nothing at initialization. Delete it and the model still runs — it just loses the knob that makes deep stacks behave. The same logic explains why FeedFoward's final 4C → C linear is described as a projection rather than just the second half of an MLP.
- The diagram says "Add & Norm," the code says norm-then-add. The paper's Figure 1 is post-norm; gpt.py is pre-norm, and Karpathy calls the swap out explicitly at 95:12. If you are reading the paper alongside the code, expect this one mismatch and no others in the decoder path. It also explains an apparent extra layer: pre-norm requires a final ln_f that the original diagram has no box for.
- LayerNorm on (B, T, C) does not normalize over time. A common wrong picture is that LayerNorm smooths the sequence. nn.LayerNorm(n_embd) reduces over the last dimension only — 32 numbers for one token, at this stage — and treats both B and T as batch axes. Neighbouring tokens have no effect on each other's normalization. If it did reduce over T, it would leak the future into the past and quietly break the causal mask you spent part 4 building.
Go deeper, verified
- Attention Is All You Need — Vaswani et al. (2017) · §3.2.2 is multi-head, §3.3 is the position-wise feed-forward with its 512/2048 numbers, §3.1 and Figure 1 are the Add & Norm placement this part deliberately reverses. The whole of P06 is these three subsections in code.
- Deep Residual Learning for Image Recognition — He, Zhang, Ren, Sun (2015) · the skip connection, and the empirical case that depth without it makes training worse, not just slower. Read §3.1–3.2 for the argument Karpathy compresses into one minute.
- Layer Normalization — Ba, Kiros, Hinton (2016) · §2–3 derive it exactly as the transpose of batch normalization and explain why sequence models are where BatchNorm falls down.
- On Layer Normalization in the Transformer Architecture — Xiong et al. (2020) · field map extra. The gradient analysis behind pre-norm: post-norm has exploding gradients at the output layer at initialization and needs warmup; pre-norm does not. This is the paper the video's "it is now more common" is standing on.
- nanoGPT model.py — Karpathy · the production version of this exact block: L94–L106 is the same Block with GELU instead of ReLU, L142–L145 is the residual-projection init the lecture only describes, and L29–L76 folds all heads into one batched matmul — the answer to Karpathy's EX1, walked through in P07.
- Building makemore Part 3: Activations & Gradients, BatchNorm — Karpathy (2022) · the earlier lecture he edits live to produce LayerNorm. If the dim-0-to-dim-1 change went past you, this is the hour that makes it obvious.
- Root Mean Square Layer Normalization and GLU Variants Improve Transformer — Zhang & Sennrich (2019); Shazeer (2020) · field map extra. What replaced the two boxes in this part. Modern stacks drop the mean-subtraction and the bias from LayerNorm (RMSNorm) and replace ReLU-MLP with a gated SwiGLU at roughly 8/3× width. Both are drop-in swaps into Block, which is the point.
Exercises
- Ablate the residual streamcode — in the Colab, set n_layer = 4 and train two variants for the same number of steps: Block.forward as written, and one with the additions removed (x = self.sa(self.ln1(x)); x = self.ffwd(self.ln2(x))). Record train and val loss at every eval_interval and plot both curves. Then repeat at n_layer = 1. A good answer shows the gap widening with depth — at one layer the two are nearly indistinguishable, at four the no-residual run is clearly worse — and names the mechanism: with no addition node, the gradient reaching the embedding table has passed through every weight matrix in the stack.
- Post-norm versus pre-normcode — change Block.forward to the paper's ordering, x = self.ln1(x + self.sa(x)) and x = self.ln2(x + self.ffwd(x)), and drop ln_f (post-norm does not need it). Train at n_layer = 6 against the unmodified model. Then re-run the post-norm version with a 200-step linear learning-rate warmup. A good answer reports three curves and observes that post-norm's disadvantage is largest in the first few hundred steps and largely closed by warmup — which is the practical reason the field moved, not a claim that pre-norm has a better optimum.
- Where does the 4× go? — sweep the feed-forward multiplier over 1, 2, 4 and 8 at fixed n_embd = 32, n_layer = 3. For each, write down the total parameter count (sum(p.numel() for p in m.parameters())), the fraction of parameters living in the MLPs versus the attention heads, and the final validation loss. A good answer notices that the MLP dominates the parameter count at 4×, that the loss improvement is sublinear in width, and can say what a fair comparison would be instead — matched parameters, e.g. more layers at 2× against fewer at 8×.