Scaling up: hyperparameters, dropout, the final gpt.py, nanoGPT
Transcript: this part, with timestamps
Everything up to 97:49 was construction: each segment added a component and the loss ticked down a hundredth at a time on a toy-sized model. This part is the payoff and the debrief. It answers the question a reader has been holding since the first attention block — does this actually work if you make it real? — and then, having answered yes, spends the remaining eight minutes placing what was built on the map: which part of the original paper it is, which part it deliberately isn't, and what the same code looks like when someone writes it for production rather than for a lecture.
Outline, with timestamps
- 97:49 — Cosmetic refactor: n_layer and n_head become variables, blocks become an nn.Sequential, the final LayerNorm is pulled out.
- 98:17 — Where dropout goes: on the attention matrix after softmax, and on each residual branch just before it rejoins the stream.
- 98:49 — What dropout is: random masking each forward/backward pass, read as training an ensemble of subnetworks that merge at test time.
- 99:20 — The new hyperparameters: batch 64, context 256, lr 3e-4, 384 channels, 6 heads, 6 layers, dropout 0.2.
- 100:21 — The result: validation loss 1.48, from roughly 2.07 before scaling.
- 100:52 — Hardware reality: ~15 minutes on an A100; what to shrink if you have no GPU.
- 101:24 — Reading a 10,000-character sample: correct Shakespeare shape, no Shakespeare meaning.
- 102:39 — Encoder vs decoder vs both: why half the famous figure was never implemented.
- 103:28 — The mask is the definition: translation needs an encoder because the generation is conditioned on something.
- 105:00 — Cross-attention: queries from the decoder, keys and values from the encoder's output.
- 106:22 — nanoGPT: train.py is the messy part, model.py is nearly what you just wrote.
- 107:08 — CausalSelfAttention: all heads as a fourth tensor dimension (this is EX1), then GELU, weight-decay groups, generate.
The refactor comes first, and it is not cosmetic
Karpathy calls the first change cosmetic, and in one sense it is: no math moves. But it is the change that makes the rest possible. Up to now the model was written with its shape baked in — one hand-placed block after another, a head size computed inline. He replaces that with three numbers you can turn: n_layer, n_head, n_embd. The blocks become a single nn.Sequential built from a list comprehension, and the final LayerNorm — the one he added as an afterthought at 97:17, sitting between the last block and the vocabulary projection — gets its own name, ln_f.
The head size is now derived rather than chosen: inside Block.__init__, head_size = n_embd // n_head. With 384 channels and 6 heads that is 64 per head, which is the same per-head width GPT-2 and GPT-3 use at every scale — as you widen a model you add heads, not head size. That single line also silently imposes a constraint the lecture code never checks: n_head must divide n_embd exactly, or the concatenated heads come out narrower than the residual stream and the addition x + self.sa(...) fails. nanoGPT asserts it; gpt.py lets you find out the hard way.
Dropout, and why it arrives exactly here
Dropout is not part of the architecture argument. It appears at this moment for one reason, which Karpathy is explicit about:
I added it because I'm about to scale up the model quite a bit and I was concerned about overfitting.Andrej Karpathy, 99:20
The arithmetic behind that worry is worth doing. Tiny Shakespeare is a little over a million characters. The model about to be trained on it has 10.8 million parameters — roughly ten parameters per character of training data. Nothing prevents a network with that much capacity from memorising the corpus outright, and the train/validation gap is where you would see it happen.
Dropout, from Srivastava et al. (2014), zeroes a random subset of activations on every forward/backward pass. Karpathy gives the standard reading — because the mask is redrawn each pass, you are effectively training an exponentially large family of thinned subnetworks that share weights, and at test time the full network approximates their ensemble — and then declines to go further, pointing at the paper for the detail. That is a fair place to stop for this lecture, but two mechanical facts matter for reading the code:
- PyTorch uses inverted dropout. nn.Dropout(p) zeroes a fraction p of entries and scales the survivors by 1/(1-p) during training, so the expected value of the tensor is unchanged and evaluation needs no rescaling at all. At dropout=0.2 the survivors are multiplied by 1.25.
- It is the first module in this codebase whose behaviour differs between train and eval. The model.eval() / model.train() pair inside estimate_loss was written back in the bigram section and has been a no-op ever since — LayerNorm behaves identically in both modes. From this commit on, deleting those two lines would quietly corrupt every reported validation number.
Karpathy puts dropout in three places, and the placement is the interesting part. Two of them sit at the end of a residual branch, immediately before the branch is added back into the stream: the output projection of multi-head attention (out = self.dropout(self.proj(out))) and the last layer of the feed-forward nn.Sequential. That is the canonical position — you are randomly deleting whole contributions to the residual stream, which is what forces the model to spread a computation across redundant paths rather than betting it all on one.
The third is more unusual and more specific to attention: wei = self.dropout(wei) sits after the softmax inside Head.forward. It randomly severs individual token-to-token communication links. A token that has learned to lean entirely on one earlier token will, one pass in five, find that edge gone. Note the side effect: the attention rows are a probability distribution when they leave the softmax, and they are not one after dropout — mass is deleted and the remainder scaled up. The expectation survives; the per-row sum does not. This is exactly what the 2017 paper does too (§5.4, with Pdrop = 0.1).
The numbers, and what they bought
Everything below the model definition is unchanged from the toy run. Only the constants move.
| Hyperparameter | Final value | What it controls, and why this number |
|---|---|---|
| batch_size | 64 | Independent sequences per step. Doubled from the toy runs — a GPU is idle at 32, and a bigger batch steadies the gradient. |
| block_size | 256 | Context length. Up from 8: 256 characters of history to predict the 257th. This is the axis that costs quadratically. |
| learning_rate | 3e-4 | Lowered, because the network is much bigger. Stated as judgement, not derived — and left constant, with no warmup or decay. |
| n_embd | 384 | Residual-stream width. Up from 32. Divided by 6 heads gives 64 channels per head. |
| n_head | 6 | Heads per block. Must divide n_embd; nothing in the script checks that it does. |
| n_layer | 6 | Blocks stacked in nn.Sequential. Residuals plus LayerNorm are what make six trainable at all. |
| dropout | 0.2 | 20% of the affected activations zeroed on every forward/backward pass, at three sites. |
| — | 10.788929 M | Resulting parameter count, printed by the script. |
The reported result is a validation loss of 1.48. The comparison point Karpathy quotes is ~2.07, which is the small model with LayerNorm at the end of P6; he gives it as 2.06 a minute earlier and 2.07 here, so read it as "about 2.07" rather than a precise figure. In per-character terms 1.48 nats is roughly 2.14 bits per character — down from about 2.99 bits. Note what did not change: not one line of the model class, the optimiser, the data loader or the loss. The entire improvement came from width, depth and context.
The parameter count printed by the script is 10.788929 M, which he rounds to "about 10 million". If you want to check your own understanding of the architecture, that number is fully derivable: each block is 1,773,312 parameters (442,368 of key/query/value with no bias, 147,840 for the output projection, 1,181,568 for the two feed-forward matrices, 1,536 for the two LayerNorms), times six, plus 24,960 of token embeddings, 98,304 of position embeddings, 768 for ln_f and 25,025 for the language-model head.
The sample he prints is 10,000 characters written to a file rather than the 300-character dribble from earlier runs, and it is worth looking at for what it gets right. Speaker names on their own lines, colons, act-like line lengths, plausible early-modern morphology, correctly balanced apostrophes. It has learned the form of the document with no supervision of any kind about form. It has learned nothing about meaning, because at one million characters and character-level granularity there is not enough signal to learn any. That gap — perfect surface, absent content — is the single most useful intuition to take out of this run, and it is the thing that scale, in the next part, closes.
Which half of the paper this is
Karpathy now points at Figure 1 of Attention Is All You Need and accounts for the parts he never implemented. The figure has two towers. The right-hand one is what you built; the left-hand one, the encoder, does not exist in this codebase, and neither does the middle attention block inside the decoder that connects them.
The distinction is not architectural in any deep sense — encoder and decoder blocks are the same code — it is one line:
what makes it a decoder is that we are using the triangular mask.Andrej Karpathy, 103:28
Delete wei.masked_fill(self.tril[:T, :T] == 0, float('-inf')) and every token attends to every other token in both directions; you have an encoder, useful for producing representations, useless for generation because position t has already seen its own answer. Keep the mask and you have the autoregressive property that makes sampling well-defined.
The 2017 paper needed both towers because it was a machine-translation paper. The generation there is conditional: emit English, given a French sentence. So the French is run through an unmasked encoder — all tokens free to talk — and its output is fed into every decoder block through a third attention module. In that module the queries come from the decoder's own residual stream while the keys and values come from the encoder's output. That asymmetry is the whole of cross-attention: the decoder asks the questions, the encoded source supplies the content. Special <start> and <end> tokens bracket the generation.
This lecture's model is unconditional — there is nothing to condition on, just a corpus to imitate — so the encoder and the cross-attention have nothing to do and are simply absent. That is what "decoder-only" means, and it is the configuration GPT uses. Worth knowing for reading modern code: an instruction-tuned chat model is still decoder-only. The prompt is not encoded separately; it is just earlier tokens in the same masked stream.
nanoGPT: the same model, with the heads batched
The last three minutes are a tour of nanoGPT, and the framing is useful: train.py is where all the complexity you have been spared lives — checkpointing, learning-rate decay, torch.compile, distributed training across GPUs — while model.py is recognisably the file you just wrote. The Block is line-for-line identical in structure. GPT.forward does token embeddings, position embeddings, the blocks, a final LayerNorm, a linear head.
The one real difference is CausalSelfAttention, and it is the answer to Karpathy's first suggested exercise. In gpt.py, a Head is a module and MultiHeadAttention holds six of them in an nn.ModuleList, running a Python loop and concatenating. Mathematically fine, computationally wasteful: six small matmuls that the GPU launches one at a time. nanoGPT does the identical arithmetic in one pass by promoting the head index to a tensor dimension:
- One fused projection. c_attn is a single nn.Linear(n_embd, 3 * n_embd) producing queries, keys and values for all heads at once; .split(n_embd, dim=2) cuts the result into three (B, T, C) tensors. Eighteen small matrices in gpt.py become one.
- Reshape, then transpose. view(B, T, nh, hs).transpose(1, 2) gives (B, nh, T, hs). The last two dimensions are now exactly the matrix a single head works on, and B and nh are both just batch dimensions that @ broadcasts over. q @ k.transpose(-2, -1) yields (B, nh, T, T) — every head's attention matrix, computed together.
- The mask gains two dummy axes. tril is registered as (1, 1, block_size, block_size) so it broadcasts across batch and head. Same mask, same direction, same -inf.
- Reassembly is the concatenate. y.transpose(1, 2).contiguous().view(B, T, C) puts head outputs side by side in the channel dimension — which is precisely what torch.cat(..., dim=-1) did in the lecture version, just without materialising six tensors. The .contiguous() is required: after a transpose the memory layout no longer matches the logical order and view refuses.
- Flash attention if available. When PyTorch ≥ 2.0 is present, all four of the manual lines collapse into F.scaled_dot_product_attention(q, k, v, is_causal=True), which fuses scale/mask/softmax/dropout/matmul into one kernel that never writes the (B, nh, T, T) matrix to memory. Same numbers, far less memory traffic.
Karpathy flags two other cosmetic differences: the MLP uses GELU rather than ReLU purely so OpenAI's GPT-2 checkpoints load cleanly, and the optimiser splits parameters into weight-decayed and non-decayed groups. One more, which he does not mention and which is easy to miss: nanoGPT ties the token embedding and the output head (wte.weight = lm_head.weight), and its get_num_params subtracts position embeddings by default — so its printed counts are not comparable to gpt.py's.
The code at the end of this part
This is gpt.py — the final script. The excerpt below is the part this section produced: the scaled hyperparameters, the three dropout sites, and the assembled model.
# hyperparameters
batch_size = 64 # how many independent sequences will we process in parallel?
block_size = 256 # what is the maximum context length for predictions?
max_iters = 5000
eval_interval = 500
learning_rate = 3e-4
device = 'cuda' if torch.cuda.is_available() else 'cpu'
eval_iters = 200
n_embd = 384
n_head = 6
n_layer = 6
dropout = 0.2
# ------------
class Head(nn.Module):
""" one head of self-attention """
def __init__(self, head_size):
super().__init__()
self.key = nn.Linear(n_embd, head_size, bias=False)
self.query = nn.Linear(n_embd, head_size, bias=False)
self.value = nn.Linear(n_embd, head_size, bias=False)
self.register_buffer('tril', torch.tril(torch.ones(block_size, block_size)))
self.dropout = nn.Dropout(dropout)
def forward(self, x):
# input of size (batch, time-step, channels)
# output of size (batch, time-step, head size)
B,T,C = x.shape
k = self.key(x) # (B,T,hs)
q = self.query(x) # (B,T,hs)
# compute attention scores ("affinities")
wei = q @ k.transpose(-2,-1) * k.shape[-1]**-0.5 # (B, T, hs) @ (B, hs, T) -> (B, T, T)
wei = wei.masked_fill(self.tril[:T, :T] == 0, float('-inf')) # (B, T, T)
wei = F.softmax(wei, dim=-1) # (B, T, T)
wei = self.dropout(wei)
# perform the weighted aggregation of the values
v = self.value(x) # (B,T,hs)
out = wei @ v # (B, T, T) @ (B, T, hs) -> (B, T, hs)
return out
class MultiHeadAttention(nn.Module):
""" multiple heads of self-attention in parallel """
def __init__(self, num_heads, head_size):
super().__init__()
self.heads = nn.ModuleList([Head(head_size) for _ in range(num_heads)])
self.proj = nn.Linear(head_size * num_heads, n_embd)
self.dropout = nn.Dropout(dropout)
def forward(self, x):
out = torch.cat([h(x) for h in self.heads], dim=-1)
out = self.dropout(self.proj(out))
return out
class FeedFoward(nn.Module):
""" a simple linear layer followed by a non-linearity """
def __init__(self, n_embd):
super().__init__()
self.net = nn.Sequential(
nn.Linear(n_embd, 4 * n_embd),
nn.ReLU(),
nn.Linear(4 * n_embd, n_embd),
nn.Dropout(dropout),
)
def forward(self, x):
return self.net(x)
class Block(nn.Module):
""" Transformer block: communication followed by computation """
def __init__(self, n_embd, n_head):
# n_embd: embedding dimension, n_head: the number of heads we'd like
super().__init__()
head_size = n_embd // n_head
self.sa = MultiHeadAttention(n_head, head_size)
self.ffwd = FeedFoward(n_embd)
self.ln1 = nn.LayerNorm(n_embd)
self.ln2 = nn.LayerNorm(n_embd)
def forward(self, x):
x = x + self.sa(self.ln1(x))
x = x + self.ffwd(self.ln2(x))
return x
class GPTLanguageModel(nn.Module):
def __init__(self):
super().__init__()
# each token directly reads off the logits for the next token from a lookup table
self.token_embedding_table = nn.Embedding(vocab_size, n_embd)
self.position_embedding_table = nn.Embedding(block_size, n_embd)
self.blocks = nn.Sequential(*[Block(n_embd, n_head=n_head) for _ in range(n_layer)])
self.ln_f = nn.LayerNorm(n_embd) # final layer norm
self.lm_head = nn.Linear(n_embd, vocab_size)
def forward(self, idx, targets=None):
B, T = idx.shape
# idx and targets are both (B,T) tensor of integers
tok_emb = self.token_embedding_table(idx) # (B,T,C)
pos_emb = self.position_embedding_table(torch.arange(T, device=device)) # (T,C)
x = tok_emb + pos_emb # (B,T,C)
x = self.blocks(x) # (B,T,C)
x = self.ln_f(x) # (B,T,C)
logits = self.lm_head(x) # (B,T,vocab_size)
...
def generate(self, idx, max_new_tokens):
# idx is (B, T) array of indices in the current context
for _ in range(max_new_tokens):
# crop idx to the last block_size tokens
idx_cond = idx[:, -block_size:]
...
model = GPTLanguageModel()
m = model.to(device)
# print the number of parameters in the model
print(sum(p.numel() for p in m.parameters())/1e6, 'M parameters')
Reading it block by block. L5–L16 is the only thing that changed to get from 2.07 to 1.48. L86 is attention dropout, applied to a (B, T, T) tensor after the softmax has already normalised each row. L103 and L115 are the two residual-branch dropouts, each the last operation before its branch is added back into the stream at L134–L135. L145 is the depth knob — six Blocks wrapped in nn.Sequential so self.blocks(x) runs the whole stack. Shapes through forward: idx is (64, 256) integers; tok_emb is (64, 256, 384); pos_emb is (256, 384) and broadcasts across the batch; every block preserves (64, 256, 384); lm_head produces (64, 256, 65), which is flattened to (16384, 65) against 16,384 targets for the cross-entropy.
One drift to be aware of. The gpt.py in the repo today contains an _init_weights method and a self.apply(self._init_weights) call at L149–L158, added after the recording — its own comment says "not covered in the original GPT video". If you are following along keystroke by keystroke and your file does not have it, you have not made a mistake. It initialises Linear and Embedding weights from a normal with std 0.02, which is what GPT-2 does, and it matters more than the lecture lets on: the default PyTorch initialisation trains fine here but converges from a worse starting point.
For contrast, the batched form from nanoGPT's model.py — the same arithmetic with nh as a fourth dimension:
def forward(self, x):
B, T, C = x.size() # batch size, sequence length, embedding dimensionality (n_embd)
# calculate query, key, values for all heads in batch and move head forward to be the batch dim
q, k, v = self.c_attn(x).split(self.n_embd, dim=2)
k = k.view(B, T, self.n_head, C // self.n_head).transpose(1, 2) # (B, nh, T, hs)
q = q.view(B, T, self.n_head, C // self.n_head).transpose(1, 2) # (B, nh, T, hs)
v = v.view(B, T, self.n_head, C // self.n_head).transpose(1, 2) # (B, nh, T, hs)
# causal self-attention; Self-attend: (B, nh, T, hs) x (B, nh, hs, T) -> (B, nh, T, T)
att = (q @ k.transpose(-2, -1)) * (1.0 / math.sqrt(k.size(-1)))
att = att.masked_fill(self.bias[:,:,:T,:T] == 0, float('-inf'))
att = F.softmax(att, dim=-1)
att = self.attn_dropout(att)
y = att @ v # (B, nh, T, T) x (B, nh, T, hs) -> (B, nh, T, hs)
y = y.transpose(1, 2).contiguous().view(B, T, C) # re-assemble all head outputs side by side
# output projection
y = self.resid_dropout(self.c_proj(y))
return y
Line for line this maps onto Head.forward plus MultiHeadAttention.forward, with two tensor dimensions where the lecture had one and a Python loop. The self.bias buffer is nanoGPT's name for tril — an unfortunate collision with the usual meaning of "bias", and a reliable source of confusion when reading the file cold.
Where people get stuck
- "My generate crashes with an index error at 256 tokens." The bigram version of generate fed the whole running sequence back in, which was harmless because that model had no positional embeddings. Now position_embedding_table has exactly block_size rows, and torch.arange(T) indexes past the end the moment the context exceeds 256. The fix is the one line in gpt.py that the earlier script lacks: idx_cond = idx[:, -block_size:]. Cropping is not an optimisation, it is a correctness requirement.
- "Validation loss looks different every time I evaluate, and worse than training." Two separate things. The noise is because estimate_loss averages a random sample of eval_iters batches, not the whole split. The gap is real and is what dropout was added to control — but only if model.eval() is actually being called. Dropout is the first layer in this codebase for which train and eval mode differ; before this part those calls did nothing, so they are easy to have dropped while refactoring.
- "Why doesn't my batched attention match the loop version?" Almost always the reshape. From (B, T, C) you must view(B, T, nh, hs) and then transpose(1, 2) — reshaping straight to (B, nh, T, hs) interleaves the wrong elements, because the head index is the slower-varying part of the channel axis, not of the time axis. Coming back, transpose(1, 2) then .contiguous().view(B, T, C); without .contiguous() PyTorch raises rather than silently reordering. Sanity-check by running both implementations on the same input with dropout at 0 and asserting torch.allclose.
- "Is a decoder-only model unable to see the prompt?" A common misreading of the encoder/decoder discussion. The mask stops a token seeing the future, not the prompt — prompt tokens are simply earlier positions in the same sequence, and every generated token attends to all of them. Encoders are for when the conditioning input is a genuinely separate sequence that should be read bidirectionally, as a source sentence is in translation.
Go deeper, verified
- Attention Is All You Need — Vaswani et al. (2017) · Figure 1 is the diagram being accounted for here; §3.2.3 enumerates the three places attention is used (encoder self-attention, masked decoder self-attention, encoder–decoder cross-attention) and §5.4 is the dropout recipe.
- nanoGPT model.py — Karpathy · CausalSelfAttention is the published answer to EX1; also the place to read weight tying, the GPT-2 scaled residual init, and from_pretrained.
- nanoGPT train.py — Karpathy · everything the lecture's twenty-line training loop omits: cosine learning-rate decay with warmup, gradient accumulation, mixed precision, checkpointing, DDP.
- Dropout: A Simple Way to Prevent Neural Networks from Overfitting — Srivastava, Hinton, Krizhevsky, Sutskever & Salakhutdinov (2014) · the paper Karpathy points at; §7.5 covers the inverted-dropout test-time scaling that PyTorch implements for you.
- Gaussian Error Linear Units (GELUs) — Hendrycks & Gimpel (2016) · the nonlinearity nanoGPT uses in place of ReLU, kept for GPT-2 checkpoint compatibility.
- FlashAttention — Dao, Fu, Ermon, Rudra & Ré (2022) · field map extra · what F.scaled_dot_product_attention dispatches to; it computes exactly the same result without ever materialising the T×T matrix, which is why long contexts became affordable after 2022.
- Using the Output Embedding to Improve Language Models — Press & Wolf (2016) · field map extra · weight tying, the one nanoGPT line with no counterpart in gpt.py; at character-level vocabularies it barely matters, at 50k tokens it is a large fraction of the parameters.
Exercises
- EX1 — the n-dimensional tensor mastery challengecode — This is Karpathy's own first suggested exercise, and this part is where it belongs. Collapse Head and MultiHeadAttention into a single module that processes all heads in parallel by treating the head index as another batch dimension. Steps: (1) replace the three per-head nn.Linear(n_embd, head_size) layers with one nn.Linear(n_embd, 3 * n_embd, bias=False) and .split(n_embd, dim=2); (2) reshape each of q, k, v with view(B, T, n_head, hs).transpose(1, 2) and confirm the shape is (B, nh, T, hs); (3) register tril with shape (1, 1, block_size, block_size) so it broadcasts; (4) reassemble with transpose(1, 2).contiguous().view(B, T, C) and keep the output projection. A good answer includes a test: build both modules, copy the loop version's weights into the batched one, run the same input with dropout=0 and model.eval(), and assert torch.allclose(a, b, atol=1e-6) — getting shapes to line up is easy, getting the weight-copy permutation right is the actual exercise. The published answer is nanoGPT's CausalSelfAttention; write yours before reading it.
- Fit the run to your machinecode — The 384/6/6/256 configuration assumes a datacentre GPU. Find the largest configuration that trains in under ten minutes on whatever you have, and record the trade. Sweep three or four points (e.g. 128/4/4/64, 192/6/4/128, 256/8/6/128), and for each log parameter count, wall-clock time and final validation loss. Plot loss against parameters, and separately against wall-clock. A good answer notes which knob bought the most loss per second — and observes that block_size costs quadratically in time while n_embd and n_layer cost linearly, so the cheap and expensive axes are not the ones intuition suggests.
- Ablate dropout, and locate the overfittingcode — Karpathy adds dropout preemptively and never shows the counterfactual. Run your configuration twice, at dropout=0.0 and at 0.2, logging train and validation loss at every eval_interval, and plot all four curves on one chart. Then run a third at 0.2 with the attention dropout removed but both residual dropouts kept, to see which of the three sites is doing the work. A good answer states the iteration at which train and validation diverge in each run, whether dropout delays that point or only flattens it, and whether at this data-to-parameter ratio the 0.2 setting is actually paying for itself or is just costing convergence speed.