The crux: self-attention
Transcript: this part, with timestamps
The lecture has spent twenty minutes building an elaborate way to compute a running average — for-loops, then a lower-triangular matrix, then a softmax over a matrix of zeros — and Karpathy has been careful to say each time that the numbers in that matrix are placeholders. This part is where the placeholders get filled in. Everything structural was already in place at the end of part 3; what arrives here is the one idea that turns the scaffolding into a transformer, and it is small enough to fit in fifteen lines. Karpathy flags it as the single most important stretch of the video, and he is right: parts 5 through 7 are notes, plumbing and scale on top of what happens in these ten minutes.
Outline, with timestamps
- 62:00 — Setting the stage: the toy tensor is now B,T,C = 4,8,32, and the current code does a plain average of the past.
- 63:22 — Why uniform is wrong: different tokens should find different earlier tokens interesting, and that has to depend on the data.
- 63:52 — The mechanism stated: every node emits a query and a key; the affinity is their dot product.
- 64:55 — head_size = 16; three nn.Linear(C, head_size, bias=False) projections, and why the bias is dropped.
- 65:26 — Producing k and q: (B,T,16) each, computed for all positions in parallel, with no communication yet.
- 65:57 — The batched matmul: q @ k.transpose(-2,-1), and why it is -2,-1 and not 0,1.
- 66:27 — Shapes: (B,T,16) @ (B,16,T) → (B,T,T), one square affinity matrix per batch element.
- 67:29 — Inspecting the result: the weights are no longer shared across the batch, and no longer uniform within a row.
- 68:00 — The worked intuition: the eighth token's query hunting for a matching key in a particular channel.
- 69:01 — Peeling the mask and softmax off to show the raw dot products, then putting them back.
- 70:04 — The third projection: aggregate v, not x. Output is (B,T,head_size).
- 71:06 — Private versus public: x is what a token is, v is what it says.
Where part 3 leaves us, and why that is not enough
At the start of this stretch the toy tensor has been widened from two channels to thirty-two, so we are looking at B,T,C = 4,8,32: four independent sequences, eight positions each, a thirty-two-dimensional vector per position. The aggregation code from part 3 builds an 8×8 matrix of zeros, masks its upper triangle to -inf, softmaxes each row, and matrix-multiplies the result against x. Because every unmasked entry in a row was the same number, softmax hands back the uniform distribution over the allowed positions, and the matmul computes, for position t, the mean of positions 0…t.
That is a legitimate way to move information backwards in time, and it is also a terrible prior. Averaging says every earlier token is equally relevant to me, always, regardless of what any of us are. A character in the middle of a word cares enormously about the two characters before it and essentially not at all about a character forty positions back — and which forty-positions-back token it cares about changes with the content. Karpathy's example is deliberately linguistic and deliberately vague:
If I'm a vowel then maybe I'm looking for consonants in my past, and maybe I want to know what those consonants are, and I want that information to flow to me.Karpathy, 63:22
Note the shape of the wish. It is not "let me learn a fixed 8×8 weight matrix" — that would be a learned but still positional prior, the same for every sentence, and it would break the moment the context got longer. It is "let the weight from me to you be computed from what you and I currently are." The weights have to be a function of the activations, not parameters in their own right. That constraint is what forces the whole query/key construction; everything else follows from it.
Query, key, value: three questions each token answers
The trick is to give every position a small vocabulary for advertising itself and for asking about others, and then to define compatibility as a dot product in that space. Concretely, three learned linear maps run over the channel dimension, applied identically and independently at every position:
- query = nn.Linear(n_embd, head_size, bias=False) — "what am I looking for?"
- key = nn.Linear(n_embd, head_size, bias=False) — "what do I contain?"
- value = nn.Linear(n_embd, head_size, bias=False) — "if you decide I'm relevant, here is what you get."
Three points about this that are easy to skate past. First, they are bias-free, which for a single head is close to a convention: a constant added to every key shifts every score in a column by the same amount, and softmax over the row can partially absorb that, so the bias buys little and later gets subsumed by LayerNorm anyway. Second, they are applied at every position in parallel with no interaction — after the projections, position 5 still knows nothing about position 3. All the communication in a transformer happens in exactly one place, the matmul that comes next. Third, head_size is a genuinely new hyperparameter, unrelated to n_embd: in the notebook it is 16 while C is 32. That asymmetry is the seed of multi-head attention in part 6 — six heads of size 64 concatenating back up to 384 channels.
The asymmetry between query and key is worth dwelling on, because it is the part people quietly assume away. q and k are produced by different weight matrices, so the affinity from i to j is q_i · k_j, and the affinity from j to i is q_j · k_i. There is no reason for these to be equal. Attention is a directed relation, not a similarity — "I find you interesting" and "you find me interesting" are separate claims, and the model can and does learn them separately. If the two projections were tied, the score matrix would be symmetric and the head would lose half its expressive range.
The matmul that does the talking
With k and q both (B, T, head_size), the affinity matrix is every query dotted against every key. In tensor form that is q @ k.transpose(-2, -1). Two details in that line trip people up. The transpose is by negative index, -2 and -1, because dimension 0 is the batch and must be left alone — k.transpose(0, 1) or a bare k.T would scramble batch with time. And @ on three-dimensional tensors is a batched matmul: PyTorch treats the leading dimension as a stack of independent matrix multiplies, so what actually happens is four separate 8×16 by 16×8 products, one per sequence, never mixing them.
| Step | Expression | Shape | What it is |
|---|---|---|---|
| input | x | (4, 8, 32) | (B, T, C) — token embedding plus position embedding, from part 3 |
| keys | k = key(x) | (4, 8, 16) | one head_size advertisement per position |
| queries | q = query(x) | (4, 8, 16) | one request per position |
| scores | q @ k.transpose(-2,-1) | (4, 8, 8) | (B,T,T) — row t, column s is how much t wants s |
| mask | wei.masked_fill(tril == 0, -inf) | (4, 8, 8) | strictly-upper triangle set to -inf: the future is deleted |
| weights | F.softmax(wei, dim=-1) | (4, 8, 8) | each row a distribution; row t has t+1 nonzero entries |
| values | v = value(x) | (4, 8, 16) | what each position is willing to broadcast |
| output | out = wei @ v | (4, 8, 16) | (B, T, head_size) — not C; part 6 fixes that |
The single most important line in that table is the fourth, because of what it does not change: the shape. Part 3's wei was (T, T) — one matrix, broadcast identically across all four sequences, since it was built from constants. Now it is (B, T, T): every sequence in the batch gets its own affinity matrix, because every sequence contains different characters at different positions. Karpathy pauses on exactly this when he prints the tensor at 67:29 — the rows are no longer uniform, and batch element 0 no longer looks like batch element 1. (The auto-captions garble his conclusion there into "so this is not data dependent"; he is saying the opposite.) That extra B is the whole difference between a hand-designed smoothing filter and a learned one.
Mask, then softmax — in that order
The mask and the softmax are inherited unchanged from part 3, which is a nice piece of engineering economy, but their interaction is worth spelling out because it is the piece students most often reimplement wrongly. tril is a lower-triangular matrix of ones; masked_fill(tril == 0, float('-inf')) writes -inf into every position where tril is zero, which is the strictly-upper triangle — the future. Then softmax exponentiates: exp(-inf) = 0, so those entries contribute exactly zero to the row sum and exactly zero to the normalized weights. The masked positions are not merely down-weighted, they are gone, and no gradient flows back through them.
Order matters and so does the axis. Masking before the softmax is what makes the surviving weights renormalize to sum to one; masking after would leave rows summing to less than one and shrink later positions' outputs. And dim=-1 normalizes across the key axis, so each query's weights over the tokens it can see form a distribution. Softmaxing over dim=-2 instead would normalize each key's outgoing attention — a coherent-looking tensor of the same shape that trains to nothing useful, and one of the more painful silent bugs in this code.
Karpathy makes this visible by temporarily deleting the mask and the softmax at 69:01, so you can see the raw dot products, which at initialization sit roughly in the range −2 to +2 and are as happy to be negative as positive. Then he puts the two lines back: clamp the future, exponentiate, normalize. Those raw scores are also where the missing 1/√head_size factor belongs — part 5's sixth note — and its absence here is deliberate. Without the scale, the variance of the scores grows with head_size (about 16 in this toy), softmax starts to saturate toward one-hot, and the head degenerates into copying a single earlier token. The version-4 cell is correct in structure and slightly wrong in calibration, and Karpathy fixes it eight minutes later.
Why we aggregate v and not x
The last move is the least obvious and the easiest to drop when reimplementing from memory. Having computed the weights, you could aggregate x directly — out = wei @ x — and get a data-dependent weighted average of the raw residual-stream vectors. Karpathy adds a third projection instead and aggregates v = value(x). His framing is the one to keep:
Here's what I'm interested in, here's what I have, and if you find me interesting, here's what I will communicate to you.Karpathy, 71:06
So x is private and v is public. A token's residual vector carries everything it knows — its identity, its position, whatever earlier layers wrote there — and most of that is nobody else's business. The value projection is the model's chance to decide, per head, which slice of that private state is worth broadcasting when someone attends to it. Mechanically it also decouples two things that have no business being coupled: what makes a token findable (its key) and what a token delivers once found (its value). Without the value projection those are the same vector, and a head can only route information it can also match on.
The output is (B, T, head_size) — sixteen channels, not thirty-two. This head is not a drop-in replacement for the residual stream yet; something has to get it back to C. That something is multi-head attention plus an output projection in part 6, and the fact that head_size came out smaller than C is precisely what makes concatenating several heads work out.
What this cell is still missing
It is useful to be explicit about the gap between the notebook cell at 71:38 and the Head class in the finished script, because four separate things get added later and it is easy to conflate them. The scaling factor * k.shape[-1]**-0.5 arrives in part 5. The tril becomes a registered buffer (a tensor that moves with .to(device) and gets saved with the model, but is not a parameter and receives no gradient) sliced as self.tril[:T, :T], so that generation can run with fewer than block_size tokens of context. Dropout on the attention weights arrives in part 7. And the head has to be multiplied and projected back to n_embd in part 6. None of that changes the mechanism; all of it is required to train.
| Symbol | Notebook at 62:00 | gpt.py, final | Meaning |
|---|---|---|---|
| B | 4 | batch_size = 64 | independent sequences per step |
| T | 8 | block_size = 256 | maximum context length |
| C | 32 | n_embd = 384 | residual-stream width |
| head_size | 16 | n_embd // n_head = 64 | width of one head's q/k/v space |
| n_head | 1 (implicit) | 6 | heads run in parallel per block |
The code at the end of this part
Below is the Head class as it stands in the finished gpt.py. Lines 83–89 are the version-4 cell of this part, verbatim in substance; the scale factor on line 83, the buffer on line 72 and the dropout on line 86 are the later additions listed above.
class Head(nn.Module):
""" one head of self-attention """
def __init__(self, head_size):
super().__init__()
self.key = nn.Linear(n_embd, head_size, bias=False)
self.query = nn.Linear(n_embd, head_size, bias=False)
self.value = nn.Linear(n_embd, head_size, bias=False)
self.register_buffer('tril', torch.tril(torch.ones(block_size, block_size)))
self.dropout = nn.Dropout(dropout)
def forward(self, x):
# input of size (batch, time-step, channels)
# output of size (batch, time-step, head size)
B,T,C = x.shape
k = self.key(x) # (B,T,hs)
q = self.query(x) # (B,T,hs)
# compute attention scores ("affinities")
wei = q @ k.transpose(-2,-1) * k.shape[-1]**-0.5 # (B, T, hs) @ (B, hs, T) -> (B, T, T)
wei = wei.masked_fill(self.tril[:T, :T] == 0, float('-inf')) # (B, T, T)
wei = F.softmax(wei, dim=-1) # (B, T, T)
wei = self.dropout(wei)
# perform the weighted aggregation of the values
v = self.value(x) # (B,T,hs)
out = wei @ v # (B, T, T) @ (B, T, hs) -> (B, T, hs)
return out
Reading it line by line against the blob:
- L69–L71 — the three projections, all n_embd → head_size, all bias-free. These are the only parameters a head has: 3 × n_embd × head_size weights, which for the final config is 3 × 384 × 64 = 73,728 per head.
- L72 — register_buffer puts the block_size × block_size triangle in the module's state without making it a parameter, so it follows the model onto the GPU and gets no gradient.
- L79–L81 — unpack the shape, project. Nothing has communicated yet.
- L83 — the only line where positions interact. k.shape[-1] is head_size, so **-0.5 is 1/√head_size. Note the video's own correction: at 01:20:05 Karpathy normalizes by C rather than head_size on screen; the script here is the correct version.
- L84 — self.tril[:T, :T] crops to the current sequence length, which matters during generation when the context is shorter than block_size. Broadcasting handles the missing batch dimension: a (T,T) boolean mask applies to every element of a (B,T,T) tensor.
- L85–L89 — normalize across keys, drop out some of the attention edges during training, project the values, and aggregate. The output leaves the head at head_size width.
Where people get stuck
- "Why transpose(-2, -1) instead of .T?" — because k is three-dimensional. .T means "reverse every dimension" (current PyTorch deprecates it on 3-D tensors precisely because that is rarely what anyone wants), and transpose(0, 1) swaps batch with time — which produces a tensor of a perfectly plausible shape while silently mixing sequences together. Negative indices name the last two axes regardless of how many batch-like dimensions sit in front, which is why nanoGPT can reuse the identical expression on a four-dimensional (B, nh, T, hs) tensor.
- The mask direction, and the video's own erratum. tril keeps the lower triangle, so tril == 0 selects the strictly-upper triangle and that is what gets -inf. Row t therefore attends to columns 0…t inclusive — the past and itself. Karpathy misspeaks about this at 57:00 and posts a correction in the description: it is tokens from the future that cannot communicate, not the past.
- Softmax over the wrong axis. dim=-1 makes each query's row a distribution over the keys it can see. dim=-2 or dim=1 produces a tensor of exactly the same shape that normalizes columns instead — no error, no warning, just a model that trains badly. If you are debugging a from-scratch head, print wei.sum(-1) and confirm it is all ones.
- Expecting the attention matrix to be symmetric. It is not, and it should not be. Different weight matrices produce q and k, so wei[i, j] ≠ wei[j, i] in general — and under the causal mask only one of the two is ever visible anyway. "Attention" here means directed routing, not distance.
- Forgetting that out is head_size wide, not C. A single head cannot be dropped into a residual stream, because x + head(x) will not broadcast. That mismatch is not a bug in this part; it is the setup for multi-head attention and the output projection in part 6.
Go deeper, verified
- Attention Is All You Need — Vaswani et al. (2017) · §3.2.1 is the equation this part implements, scaling factor included. Read it after watching, not before; the notation is much easier once you have seen the shapes move.
- ng-video-lecture · gpt.py, class Head — Karpathy (2023) · the twenty-seven lines this part produces, with the shape comments he types on screen.
- nanoGPT · CausalSelfAttention — Karpathy (2022–) · the same math with all heads batched into one (B, nh, T, hs) tensor and one fused c_attn projection producing q, k and v at once. This is the answer to Karpathy's EX1 and the version you would actually ship.
- The Illustrated Transformer — Jay Alammar (2018) · the picture-first account of the same mechanism. Useful as a second pass if the matrix shapes are clear but the intuition is not.
- torch.nn.functional.scaled_dot_product_attention — PyTorch docs · field map extra. Modern PyTorch collapses lines 83–89 into a single call with is_causal=True. Worth reading the signature to see which of the pieces in this part became arguments.
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Dao et al. (2022) · field map extra. The (B, T, T) tensor in the middle of this part is quadratic in context length and is the reason long contexts were expensive; FlashAttention computes the same result without ever materializing it.
- RoFormer: Enhanced Transformer with Rotary Position Embedding — Su et al. (2021) · field map extra. Part 3 added positions by summing a learned table into x; RoPE instead rotates q and k so that the dot product itself becomes a function of relative distance. Nearly every open model since 2023 does it this way.
Exercises
- Shapes and parameter count by hand — before running anything, write out for B,T,C = 4,8,32 and head_size = 16: the shape after each of the eight steps in the table above, the number of learnable parameters in one Head, and how many of the 64 entries of one batch element's wei are nonzero after the mask and softmax. A good answer gets 1,536 parameters (three bias-free 32×16 matrices) and 36 nonzero entries (1+2+…+8), and can say why the parameter count does not depend on T.
- Delete the mask, and watch the loss lie to you code — in the Colab, comment out the masked_fill line so every position sees the whole sequence (this is the encoder block of part 5's note 4). Retrain briefly and record: (1) what happens to the training loss, (2) what happens to the validation loss, (3) what the generated text looks like. A good answer explains that the loss drops implausibly because each position can now read the token it is being asked to predict, and notes that generation is nonsense because at sample time there is no future to read.
- Ablate the value projection code — change the last line of Head.forward to out = wei @ x and adjust the surrounding code so the widths still line up. Train it for the same number of steps as the unmodified model and compare validation loss. Then, on the trained original, take one batch, capture wei for a single head, and print the attention row for a position in the middle of a word. Steps: register a forward hook on the head, run one batch, index wei[0, t], and map the column indices back to characters with decode. A good answer reports a measurable gap in favour of the value projection and shows at least one row where the mass concentrates on something interpretable — the previous character, or the last newline.