LET'S BUILD GPT // FIELD MAP
← field map
PART 06 · THE TRANSFORMER79:11–97:49 · 19 min

Assembling the block: heads, feed-forward, residuals, LayerNorm

Andrej Karpathy · Let's build GPT (2023) · part 06 of 8

Transcript: this part, with timestamps

TL;DR — Attention alone is one operator, not an architecture. This stretch surrounds it with the four pieces that make it trainable: several narrow heads run in parallel and get concatenated, a per-token MLP gives each position time to process what it gathered, residual connections give the gradient a clean path from the loss back to the embedding table, and LayerNorm keeps each token's 32 features at a sane scale. Validation loss walks from 2.5 to 2.06 in five steps. The thing to remember: a Transformer block is communication, then computation, both written back into a residual stream that nothing ever multiplies.

Part 5 ended with a self-attention head that works in a notebook cell. That is not yet a model. The lecture's next move is deliberately unglamorous engineering: take the head, put it in the network, and then keep adding the smallest piece that unblocks the next improvement — until the thing on screen is, line for line, the decoder half of the 2017 paper. Nothing here is a new idea about attention. Everything here is about making a stack of attention layers actually optimize.

Outline, with timestamps

One head, wired into the network

The head from part 5 becomes an nn.Module with three bias-free linear layers — key, query, value — each mapping n_embd → head_size. Bias is off because these are pure projections; there is nothing for a per-feature offset to do that the layer they feed into cannot also do.

The one piece of PyTorch trivia worth internalizing is register_buffer. The lower-triangular mask is a constant, so it is not a Parameter — the optimizer must not touch it. But it still has to move with .to(device) and show up in state_dict. A buffer is exactly that: tensor state owned by the module, excluded from gradients. Registering it as a plain attribute would leave it on the CPU the moment you train on a GPU.

class Head(nn.Module):
    """ one head of self-attention """

    def __init__(self, head_size):
        super().__init__()
        self.key = nn.Linear(n_embd, head_size, bias=False)
        self.query = nn.Linear(n_embd, head_size, bias=False)
        self.value = nn.Linear(n_embd, head_size, bias=False)
        self.register_buffer('tril', torch.tril(torch.ones(block_size, block_size)))

        self.dropout = nn.Dropout(dropout)

    def forward(self, x):
        # input of size (batch, time-step, channels)
        # output of size (batch, time-step, head size)
        B,T,C = x.shape
        k = self.key(x)   # (B,T,hs)
        q = self.query(x) # (B,T,hs)
        # compute attention scores ("affinities")
        wei = q @ k.transpose(-2,-1) * k.shape[-1]**-0.5 # (B, T, hs) @ (B, hs, T) -> (B, T, T)
        wei = wei.masked_fill(self.tril[:T, :T] == 0, float('-inf')) # (B, T, T)
        wei = F.softmax(wei, dim=-1) # (B, T, T)
        wei = self.dropout(wei)
        # perform the weighted aggregation of the values
        v = self.value(x) # (B,T,hs)
        out = wei @ v # (B, T, T) @ (B, T, hs) -> (B, T, hs)
        return out

Walk the shapes once. Input (B, T, C). Keys and queries come out (B, T, hs); transposing the last two dimensions of k gives (B, hs, T), so the matmul yields one T×T affinity matrix per batch element: (B, T, T). masked_fill writes -inf wherever the triangle is zero — that is row t, column t' with t' > t, the future — and softmax over dim=-1 turns each row into a distribution over the past that sums to 1, with the -inf entries going to exactly 0. Then wei @ v contracts the T axis away and returns (B, T, hs). The slice self.tril[:T, :T] is not decoration: at generation time the context is shorter than block_size, and the mask has to be cropped to match.

Two knock-on changes. First, generate now has to crop its input — idx_cond = idx[:, -block_size:] — because the position embedding table only has block_size rows and will index out of range on token 9 with a context window of 8. The bigram model had no such constraint; it never looked backwards at all. Second, the learning rate has to come down. The bigram script trains happily at 1e-2; with attention in the graph that diverges, so Karpathy lowers it and raises max_iters to compensate (81:20). Validation loss moves 2.5 → 2.4. It is a real improvement and an unimpressive one — a single 32-dimensional communication channel is not much of an architecture.

Four narrow heads instead of one wide one

MultiHeadAttention is four lines: build a ModuleList of num_heads heads, run them all on the same x, and concatenate the outputs along the channel dimension. The arithmetic that makes it invisible from the outside is head_size = n_embd // n_head. At this point in the video n_embd is 32 and n_head is 4, so each head produces (B, T, 8) and the concatenation restores (B, T, 32). In the final gpt.py the same identity reads 384 / 6 = 64.

Notice what this does not cost. Four heads of width 8 have exactly as many projection parameters as one head of width 32, and the aggregated value vector is the same size either way. What you buy is four independent T×T attention matrices instead of one. Each can specialize: one head can be a "find the nearest vowel" circuit, another a "what happened two characters ago" circuit, and they do not have to share a single softmax. Karpathy compares it to a grouped convolution — one big operation split into independent groups over channels (83:23). Take that as an analogy about the grouping, not a claim about convolution.

Result: 2.28. He is careful not to oversell it — the samples still read like noise. The number is the evidence, not the text.

Communication, then computation

Up to here the model attends and then immediately computes logits. Every token gathers information from its past and then, with no further processing, has to commit to a prediction. That is the gap the feed-forward layer fills:

the tokens looked at each other but didn't really have a lot of time to think on what they found from the other tokensKarpathy, 85:26

The fix is the paper's "position-wise feed-forward network," which is a two-layer MLP applied to each position separately: Linear(C, 4C) → ReLU → Linear(4C, C). The word doing the work is position-wise. nn.Linear only ever touches the last dimension, so a (B, T, C) tensor is processed as B×T independent vectors — no information crosses the time axis. That is worth stating plainly and remembering: in the entire Transformer, attention is the only operator that moves information between positions. Everything else — the MLP, the norms, the output layer — is a per-token map.

The first version he trains has no 4× widening at all, just a linear and a ReLU at width 32, and it gets 2.24 (86:26). The 4× multiplier arrives later, at 91:36, read straight off the paper's numbers: d_model 512, inner dimension 2048. If you are matching the video against the repo, the class is spelled FeedFoward — the typo is in the real file, and copying it is how you keep imports working.

Now the pattern is visible, and the Block just names it: multi-head attention (communication) followed by a feed-forward net (computation), packaged so you can stack it. nn.Sequential(*[Block(n_embd, n_head=n_head) for _ in range(n_layer)]), three deep at this stage.

Depth stops working, and residuals fix it

Stacking three blocks does not help. Karpathy is honest that he is not diagnosing this so much as recognizing the shape of it — the network is now deep enough that deep-network optimization problems are plausibly what he is hitting (88:01). He does not measure gradient norms or show a vanishing-gradient plot; he reaches for the standard fix. If you want the rigorous version, that is the ResNet paper's job, not this lecture's.

The fix is the skip connection from He et al. 2015, and the mental model he gives for it is the one worth keeping. Picture a single wide bus of activations — the residual stream — running from the embeddings at the top to the logits at the bottom, always shape (B, T, C). Every sublayer forks off the bus, computes something, and adds its result back. The bus itself is never multiplied, never normalized, never reshaped. In backward mode, addition hands the incoming gradient to both of its inputs unchanged, so the gradient from the loss reaches the token embedding table without passing through a single weight matrix:

you have this gradient superhighway that goes directly from the supervision all the way to the input, unimpededKarpathy, 89:34

In code this is two characters of change per sublayer — x = x + self.sa(...) instead of x = self.sa(x) — plus one new module each. MultiHeadAttention gains self.proj = nn.Linear(head_size * num_heads, n_embd), and FeedFoward's second linear plays the same role. This is the part people misread: at these settings head_size * num_heads equals n_embd exactly, so the projection is 32→32 and changes no shape whatsoever. It is there because writing into a shared stream is its own job — the sublayer computes in its own space, and the projection decides how that result gets expressed in the stream's basis, including scaling it down toward nothing if that is what training wants.

Karpathy adds that residual branches are typically initialized to contribute almost nothing, so the network starts out close to the identity and the blocks "come online over time." That is true of good practice but is not what gpt.py does — its _init_weights gives every nn.Linear the same std=0.02, and the file's own comment flags better init as "not covered in the original GPT video." nanoGPT does implement it, scaling every c_proj weight by 1/sqrt(2 * n_layer) per the GPT-2 paper (model.py L142–L145). Worth knowing so you do not go looking for it in the lecture's file.

With residuals, projections and the 4× MLP together, validation loss drops to 2.08 — and for the first time training loss pulls ahead of validation loss. That gap is the reason dropout shows up in part 7.

LayerNorm, and the pre-norm swap

The second depth trick is the "Norm" in the paper's "Add & Norm." Karpathy derives LayerNorm by editing the BatchNorm he wrote in an earlier video: change the reduction dimension from 0 to 1 and you are done (94:10). BatchNorm normalizes each feature across the batch — the columns of a (32, 100) matrix. LayerNorm normalizes each example across its features — the rows. Same formula, transposed intent.

The consequences all fall out of that one change. Because the statistics never cross examples, there is no running mean or variance to maintain, no momentum, and no difference between training and inference — all the machinery that makes BatchNorm fiddly disappears. What survives is the learnable gain and bias, so the output is unit-Gaussian only at initialization; training is free to move it, and generally does.

In the model it is nn.LayerNorm(n_embd) applied to a (B, T, C) tensor. The argument tells it to reduce over the last dimension only, which at this stage means the mean and variance are computed over 32 numbers, per token. B and T both act as batch dimensions (96:13) — token 3 of sequence 0 is normalized entirely on its own, using nothing from token 4 and nothing from sequence 1. It is a per-token rescaling of the feature vector, nothing more.

The placement is the interesting part, and it is the one deliberate departure from the 2017 paper in the whole build. Figure 1 puts the norm after the sublayer and the addition: x = LayerNorm(x + sublayer(x)). What everyone actually does now, and what gpt.py does, is normalize the input to the sublayer and leave the stream untouched: x = x + sublayer(LayerNorm(x)). This is the pre-norm formulation. It preserves the property the residual stream was introduced for — the identity path from loss to embeddings has nothing sitting on it — which is why pre-norm Transformers train without a learning-rate warmup and post-norm ones typically need one (Xiong et al. 2020 is the analysis; the video predates leaning on it).

Pre-norm has one obligation attached: since nothing normalizes the stream itself, the last block's output arrives at the output layer unnormalized. Hence self.ln_f, a final LayerNorm between the block stack and lm_head, which Karpathy adds as an afterthought at 97:17. It is not optional — every real implementation has it.

Adding the norms takes 2.08 to 2.06. He flags that honestly as slight, and expects the payoff to grow with depth and width. Part 7 is where that bet gets tested.

The ladder of losses

ChangeVal lossTimestamp
Where part 5 left off (embeddings + positions, no attention in the script)~2.5—
One self-attention head, head_size = n_embd = 322.481:20
Multi-head: 4 heads × 8 dims, concatenated2.2883:53
+ per-token feed-forward (plain 32 → 32 + ReLU)2.2486:26
3 blocks stacked, no residualsno improvement88:01
+ residual connections, projections, 4× MLP inner dim2.08 (train < val)92:06
+ pre-norm LayerNorms and final ln_f2.0696:46

The numbers, here versus the finished script

SymbolIn the video during this partFinal gpt.py
batch_size (B)3264
block_size (T)8256
n_embd (C)32384
n_head46
head_size = n_embd // n_head864
n_layer36
MLP inner dimension4 × 32 = 1284 × 384 = 1536
learning_ratelowered from the bigram's 1e-23e-4
dropoutnot yet0.2 (part 7)
Carry this away: a Transformer block is two sublayers with the same signature — (B,T,C) → (B,T,C) — one that mixes across time and one that does not, each reading a normalized copy of the residual stream and adding its result back. Every architectural variant you will meet later (RoPE, RMSNorm, SwiGLU, GQA, MoE) swaps out the contents of one of those two boxes. The bus, the fork, and the add stay.

The code at the end of this part

class MultiHeadAttention(nn.Module):
    """ multiple heads of self-attention in parallel """

    def __init__(self, num_heads, head_size):
        super().__init__()
        self.heads = nn.ModuleList([Head(head_size) for _ in range(num_heads)])
        self.proj = nn.Linear(head_size * num_heads, n_embd)
        self.dropout = nn.Dropout(dropout)

    def forward(self, x):
        out = torch.cat([h(x) for h in self.heads], dim=-1)
        out = self.dropout(self.proj(out))
        return out

class FeedFoward(nn.Module):
    """ a simple linear layer followed by a non-linearity """

    def __init__(self, n_embd):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(n_embd, 4 * n_embd),
            nn.ReLU(),
            nn.Linear(4 * n_embd, n_embd),
            nn.Dropout(dropout),
        )

    def forward(self, x):
        return self.net(x)

class Block(nn.Module):
    """ Transformer block: communication followed by computation """

    def __init__(self, n_embd, n_head):
        # n_embd: embedding dimension, n_head: the number of heads we'd like
        super().__init__()
        head_size = n_embd // n_head
        self.sa = MultiHeadAttention(n_head, head_size)
        self.ffwd = FeedFoward(n_embd)
        self.ln1 = nn.LayerNorm(n_embd)
        self.ln2 = nn.LayerNorm(n_embd)

    def forward(self, x):
        x = x + self.sa(self.ln1(x))
        x = x + self.ffwd(self.ln2(x))
        return x

Line by line against the repo. L92–L104, MultiHeadAttention: nn.ModuleList (not a plain Python list) so the heads' parameters are registered; torch.cat(..., dim=-1) glues num_heads tensors of (B, T, head_size) into (B, T, num_heads*head_size); self.proj is the write into the residual stream. L106–L119, FeedFoward: widen 4×, ReLU, narrow back — the second nn.Linear is this sublayer's projection, folded into the Sequential rather than given its own attribute. L121–L136, Block: head_size = n_embd // n_head is the invariant that lets the concatenation land back on n_embd; two separate LayerNorms because each has its own gain and bias; and the two x = x + ... lines with the norm inside the call are the pre-norm formulation. The nn.Dropout calls are part 7's addition — delete them and you have exactly the model this part ends on.

And the spine that consumes it, L160–L169:

    def forward(self, idx, targets=None):
        B, T = idx.shape

        # idx and targets are both (B,T) tensor of integers
        tok_emb = self.token_embedding_table(idx) # (B,T,C)
        pos_emb = self.position_embedding_table(torch.arange(T, device=device)) # (T,C)
        x = tok_emb + pos_emb # (B,T,C)
        x = self.blocks(x) # (B,T,C)
        x = self.ln_f(x) # (B,T,C)
        logits = self.lm_head(x) # (B,T,vocab_size)

Five lines, and the shape never changes until the last one. pos_emb is (T, C) and broadcasts across the batch when added to (B, T, C). self.blocks is the nn.Sequential stack; every block preserves (B, T, C), which is why they compose at all. ln_f is the pre-norm tax being paid. Only lm_head changes the last dimension, from n_embd to vocab_size = 65.

Where people get stuck

Go deeper, verified

Exercises

  1. Ablate the residual streamcode — in the Colab, set n_layer = 4 and train two variants for the same number of steps: Block.forward as written, and one with the additions removed (x = self.sa(self.ln1(x)); x = self.ffwd(self.ln2(x))). Record train and val loss at every eval_interval and plot both curves. Then repeat at n_layer = 1. A good answer shows the gap widening with depth — at one layer the two are nearly indistinguishable, at four the no-residual run is clearly worse — and names the mechanism: with no addition node, the gradient reaching the embedding table has passed through every weight matrix in the stack.
  2. Post-norm versus pre-normcode — change Block.forward to the paper's ordering, x = self.ln1(x + self.sa(x)) and x = self.ln2(x + self.ffwd(x)), and drop ln_f (post-norm does not need it). Train at n_layer = 6 against the unmodified model. Then re-run the post-norm version with a 200-step linear learning-rate warmup. A good answer reports three curves and observes that post-norm's disadvantage is largest in the first few hundred steps and largely closed by warmup — which is the practical reason the field moved, not a claim that pre-norm has a better optimum.
  3. Where does the 4× go? — sweep the feed-forward multiplier over 1, 2, 4 and 8 at fixed n_embd = 32, n_layer = 3. For each, write down the total parameter count (sum(p.numel() for p in m.parameters())), the fraction of parameters living in the MLPs versus the attention heads, and the final validation loss. A good answer notices that the MLP dominates the parameter count at 4×, that the loss improvement is sublinear in width, and can say what a fair comparison would be instead — matched parameters, e.g. more layers at 2× against fewer at 8×.
Next: P07 Scaling up: hyperparameters, dropout, the final gpt.py, nanoGPT · Back to the map.