LET'S BUILD GPT // FIELD MAP
← field map
PART 04 · BUILDING SELF-ATTENTION62:00–71:38 · 10 min

The crux: self-attention

Andrej Karpathy · Let's build GPT (2023) · part 04 of 8

Transcript: this part, with timestamps

TL;DR — Part 3 built a machine that averages every earlier token equally. That average is a fixed, hand-chosen weighting, and it is the wrong one: a token should pull hard on the few earlier tokens that matter to it and ignore the rest. Self-attention makes those weights learned and data-dependent by having each position emit two small vectors — a query ("what am I looking for") and a key ("what do I contain") — and defining the weight from position t to position s as the dot product of query t with key s. The same triangular mask and the same softmax from part 3 are reused unchanged; only the numbers going into them stopped being zeros. One more vector, the value, decides what actually gets shipped when a token is attended to. That is the entire mechanism: softmax(mask(q @ kᵀ)) @ v.

The lecture has spent twenty minutes building an elaborate way to compute a running average — for-loops, then a lower-triangular matrix, then a softmax over a matrix of zeros — and Karpathy has been careful to say each time that the numbers in that matrix are placeholders. This part is where the placeholders get filled in. Everything structural was already in place at the end of part 3; what arrives here is the one idea that turns the scaffolding into a transformer, and it is small enough to fit in fifteen lines. Karpathy flags it as the single most important stretch of the video, and he is right: parts 5 through 7 are notes, plumbing and scale on top of what happens in these ten minutes.

Outline, with timestamps

Where part 3 leaves us, and why that is not enough

At the start of this stretch the toy tensor has been widened from two channels to thirty-two, so we are looking at B,T,C = 4,8,32: four independent sequences, eight positions each, a thirty-two-dimensional vector per position. The aggregation code from part 3 builds an 8×8 matrix of zeros, masks its upper triangle to -inf, softmaxes each row, and matrix-multiplies the result against x. Because every unmasked entry in a row was the same number, softmax hands back the uniform distribution over the allowed positions, and the matmul computes, for position t, the mean of positions 0…t.

That is a legitimate way to move information backwards in time, and it is also a terrible prior. Averaging says every earlier token is equally relevant to me, always, regardless of what any of us are. A character in the middle of a word cares enormously about the two characters before it and essentially not at all about a character forty positions back — and which forty-positions-back token it cares about changes with the content. Karpathy's example is deliberately linguistic and deliberately vague:

If I'm a vowel then maybe I'm looking for consonants in my past, and maybe I want to know what those consonants are, and I want that information to flow to me.Karpathy, 63:22

Note the shape of the wish. It is not "let me learn a fixed 8×8 weight matrix" — that would be a learned but still positional prior, the same for every sentence, and it would break the moment the context got longer. It is "let the weight from me to you be computed from what you and I currently are." The weights have to be a function of the activations, not parameters in their own right. That constraint is what forces the whole query/key construction; everything else follows from it.

Query, key, value: three questions each token answers

The trick is to give every position a small vocabulary for advertising itself and for asking about others, and then to define compatibility as a dot product in that space. Concretely, three learned linear maps run over the channel dimension, applied identically and independently at every position:

Three points about this that are easy to skate past. First, they are bias-free, which for a single head is close to a convention: a constant added to every key shifts every score in a column by the same amount, and softmax over the row can partially absorb that, so the bias buys little and later gets subsumed by LayerNorm anyway. Second, they are applied at every position in parallel with no interaction — after the projections, position 5 still knows nothing about position 3. All the communication in a transformer happens in exactly one place, the matmul that comes next. Third, head_size is a genuinely new hyperparameter, unrelated to n_embd: in the notebook it is 16 while C is 32. That asymmetry is the seed of multi-head attention in part 6 — six heads of size 64 concatenating back up to 384 channels.

The asymmetry between query and key is worth dwelling on, because it is the part people quietly assume away. q and k are produced by different weight matrices, so the affinity from i to j is q_i · k_j, and the affinity from j to i is q_j · k_i. There is no reason for these to be equal. Attention is a directed relation, not a similarity — "I find you interesting" and "you find me interesting" are separate claims, and the model can and does learn them separately. If the two projections were tied, the score matrix would be symmetric and the head would lose half its expressive range.

The matmul that does the talking

With k and q both (B, T, head_size), the affinity matrix is every query dotted against every key. In tensor form that is q @ k.transpose(-2, -1). Two details in that line trip people up. The transpose is by negative index, -2 and -1, because dimension 0 is the batch and must be left alone — k.transpose(0, 1) or a bare k.T would scramble batch with time. And @ on three-dimensional tensors is a batched matmul: PyTorch treats the leading dimension as a stack of independent matrix multiplies, so what actually happens is four separate 8×16 by 16×8 products, one per sequence, never mixing them.

StepExpressionShapeWhat it is
inputx(4, 8, 32)(B, T, C) — token embedding plus position embedding, from part 3
keysk = key(x)(4, 8, 16)one head_size advertisement per position
queriesq = query(x)(4, 8, 16)one request per position
scoresq @ k.transpose(-2,-1)(4, 8, 8)(B,T,T) — row t, column s is how much t wants s
maskwei.masked_fill(tril == 0, -inf)(4, 8, 8)strictly-upper triangle set to -inf: the future is deleted
weightsF.softmax(wei, dim=-1)(4, 8, 8)each row a distribution; row t has t+1 nonzero entries
valuesv = value(x)(4, 8, 16)what each position is willing to broadcast
outputout = wei @ v(4, 8, 16)(B, T, head_size) — not C; part 6 fixes that

The single most important line in that table is the fourth, because of what it does not change: the shape. Part 3's wei was (T, T) — one matrix, broadcast identically across all four sequences, since it was built from constants. Now it is (B, T, T): every sequence in the batch gets its own affinity matrix, because every sequence contains different characters at different positions. Karpathy pauses on exactly this when he prints the tensor at 67:29 — the rows are no longer uniform, and batch element 0 no longer looks like batch element 1. (The auto-captions garble his conclusion there into "so this is not data dependent"; he is saying the opposite.) That extra B is the whole difference between a hand-designed smoothing filter and a learned one.

Mask, then softmax — in that order

The mask and the softmax are inherited unchanged from part 3, which is a nice piece of engineering economy, but their interaction is worth spelling out because it is the piece students most often reimplement wrongly. tril is a lower-triangular matrix of ones; masked_fill(tril == 0, float('-inf')) writes -inf into every position where tril is zero, which is the strictly-upper triangle — the future. Then softmax exponentiates: exp(-inf) = 0, so those entries contribute exactly zero to the row sum and exactly zero to the normalized weights. The masked positions are not merely down-weighted, they are gone, and no gradient flows back through them.

Order matters and so does the axis. Masking before the softmax is what makes the surviving weights renormalize to sum to one; masking after would leave rows summing to less than one and shrink later positions' outputs. And dim=-1 normalizes across the key axis, so each query's weights over the tokens it can see form a distribution. Softmaxing over dim=-2 instead would normalize each key's outgoing attention — a coherent-looking tensor of the same shape that trains to nothing useful, and one of the more painful silent bugs in this code.

Karpathy makes this visible by temporarily deleting the mask and the softmax at 69:01, so you can see the raw dot products, which at initialization sit roughly in the range −2 to +2 and are as happy to be negative as positive. Then he puts the two lines back: clamp the future, exponentiate, normalize. Those raw scores are also where the missing 1/√head_size factor belongs — part 5's sixth note — and its absence here is deliberate. Without the scale, the variance of the scores grows with head_size (about 16 in this toy), softmax starts to saturate toward one-hot, and the head degenerates into copying a single earlier token. The version-4 cell is correct in structure and slightly wrong in calibration, and Karpathy fixes it eight minutes later.

Why we aggregate v and not x

The last move is the least obvious and the easiest to drop when reimplementing from memory. Having computed the weights, you could aggregate x directly — out = wei @ x — and get a data-dependent weighted average of the raw residual-stream vectors. Karpathy adds a third projection instead and aggregates v = value(x). His framing is the one to keep:

Here's what I'm interested in, here's what I have, and if you find me interesting, here's what I will communicate to you.Karpathy, 71:06

So x is private and v is public. A token's residual vector carries everything it knows — its identity, its position, whatever earlier layers wrote there — and most of that is nobody else's business. The value projection is the model's chance to decide, per head, which slice of that private state is worth broadcasting when someone attends to it. Mechanically it also decouples two things that have no business being coupled: what makes a token findable (its key) and what a token delivers once found (its value). Without the value projection those are the same vector, and a head can only route information it can also match on.

The output is (B, T, head_size) — sixteen channels, not thirty-two. This head is not a drop-in replacement for the residual stream yet; something has to get it back to C. That something is multi-head attention plus an output projection in part 6, and the fact that head_size came out smaller than C is precisely what makes concatenating several heads work out.

Carry this away: attention is a learned, directed, data-dependent weighted average, and it is the only place in the transformer where positions talk to each other. Everything else — the feed-forward layers, LayerNorm, residuals, the whole depth of the stack — operates on each position independently. If you can say what q, k and v are for and why wei gains a batch dimension the moment they exist, you have the part of this video the rest of the field is built on.

What this cell is still missing

It is useful to be explicit about the gap between the notebook cell at 71:38 and the Head class in the finished script, because four separate things get added later and it is easy to conflate them. The scaling factor * k.shape[-1]**-0.5 arrives in part 5. The tril becomes a registered buffer (a tensor that moves with .to(device) and gets saved with the model, but is not a parameter and receives no gradient) sliced as self.tril[:T, :T], so that generation can run with fewer than block_size tokens of context. Dropout on the attention weights arrives in part 7. And the head has to be multiplied and projected back to n_embd in part 6. None of that changes the mechanism; all of it is required to train.

SymbolNotebook at 62:00gpt.py, finalMeaning
B4batch_size = 64independent sequences per step
T8block_size = 256maximum context length
C32n_embd = 384residual-stream width
head_size16n_embd // n_head = 64width of one head's q/k/v space
n_head1 (implicit)6heads run in parallel per block

The code at the end of this part

Below is the Head class as it stands in the finished gpt.py. Lines 83–89 are the version-4 cell of this part, verbatim in substance; the scale factor on line 83, the buffer on line 72 and the dropout on line 86 are the later additions listed above.

class Head(nn.Module):
    """ one head of self-attention """

    def __init__(self, head_size):
        super().__init__()
        self.key = nn.Linear(n_embd, head_size, bias=False)
        self.query = nn.Linear(n_embd, head_size, bias=False)
        self.value = nn.Linear(n_embd, head_size, bias=False)
        self.register_buffer('tril', torch.tril(torch.ones(block_size, block_size)))

        self.dropout = nn.Dropout(dropout)

    def forward(self, x):
        # input of size (batch, time-step, channels)
        # output of size (batch, time-step, head size)
        B,T,C = x.shape
        k = self.key(x)   # (B,T,hs)
        q = self.query(x) # (B,T,hs)
        # compute attention scores ("affinities")
        wei = q @ k.transpose(-2,-1) * k.shape[-1]**-0.5 # (B, T, hs) @ (B, hs, T) -> (B, T, T)
        wei = wei.masked_fill(self.tril[:T, :T] == 0, float('-inf')) # (B, T, T)
        wei = F.softmax(wei, dim=-1) # (B, T, T)
        wei = self.dropout(wei)
        # perform the weighted aggregation of the values
        v = self.value(x) # (B,T,hs)
        out = wei @ v # (B, T, T) @ (B, T, hs) -> (B, T, hs)
        return out

Reading it line by line against the blob:

Where people get stuck

Go deeper, verified

Exercises

  1. Shapes and parameter count by hand — before running anything, write out for B,T,C = 4,8,32 and head_size = 16: the shape after each of the eight steps in the table above, the number of learnable parameters in one Head, and how many of the 64 entries of one batch element's wei are nonzero after the mask and softmax. A good answer gets 1,536 parameters (three bias-free 32×16 matrices) and 36 nonzero entries (1+2+…+8), and can say why the parameter count does not depend on T.
  2. Delete the mask, and watch the loss lie to you code — in the Colab, comment out the masked_fill line so every position sees the whole sequence (this is the encoder block of part 5's note 4). Retrain briefly and record: (1) what happens to the training loss, (2) what happens to the validation loss, (3) what the generated text looks like. A good answer explains that the loss drops implausibly because each position can now read the token it is being asked to predict, and notes that generation is nonsense because at sample time there is no future to read.
  3. Ablate the value projection code — change the last line of Head.forward to out = wei @ x and adjust the surrounding code so the widths still line up. Train it for the same number of steps as the unmodified model and compare validation loss. Then, on the trained original, take one batch, capture wei for a single head, and print the attention row for a position in the middle of a word. Steps: register a forward hook on the head, run one batch, index wei[0, t], and map the column indices back to characters with decode. A good answer reports a measurable gap in favour of the value projection and shows at least one row where the mass concentrates on something interpretable — the previous character, or the last newline.
Next: P05 Six notes on attention · Back to the map.