The bigram baseline: model, loss, generation, training
Transcript: this part, with timestamps
Part 1 ended with a batch: an integer tensor x of shape (4, 8) and a target tensor y of the same shape, holding 32 independent next-character prediction problems. This part answers the obvious next question — what is the smallest neural network you can feed that into and get a trained, sampling language model out the other end? — and the answer is deliberately, almost insultingly small. That is the point. A baseline you fully understand is a fixed reference for everything that follows: when self-attention arrives at 62:00 and drops validation loss from 2.5 to 2.0, the only reason that number means anything is that you know exactly what the 2.5 model was doing.
Outline, with timestamps
- 22:11 — The bigram model as a lookup table: nn.Embedding(vocab_size, vocab_size), and why an embedding table can be read directly as logits.
- 23:42 — (B, T, C): indexing a table with a (4, 8) integer tensor returns (4, 8, 65). The naming convention for the whole lecture is set here.
- 24:42 — Cross-entropy as the score: negative log likelihood, and what "good logits" means numerically.
- 26:14 — The shape argument: PyTorch wants channels second, so the batch and time axes get flattened into one with .view().
- 28:17 — Sanity-checking the loss: uniform over 65 symbols is ln 65 = 4.17; the real initial loss is 4.87, and the gap is informative.
- 28:47 — generate(): the autoregressive loop — forward, take the last step, softmax, sample, concatenate.
- 31:24 — targets=None: the forward pass has to work without labels, so the loss becomes optional.
- 32:26 — Kicking off from token 0: a (1, 1) tensor holding a newline is the seed context.
- 33:58 — Why generate() is overbuilt: it feeds the whole history in and uses only the last position. Wasteful now, correct later.
- 34:53 — Training: AdamW instead of SGD, a learning rate far above the usual 3e-4, batch size bumped from 4 to 32.
- 37:05 — Loss ≈ 2.5: the first output that has word-shaped things in it, and the ceiling of what a bigram can do.
- 38:00 — Port to a script: hyperparameters hoisted to the top, a device flag, and estimate_loss() under @torch.no_grad().
A language model that is one table
A bigram language model says: the probability of the next character depends on the current character and nothing else. That is a claim about a 65 × 65 grid of numbers — one row per current character, 65 scores per row. Karpathy's implementation is exactly that grid, wrapped in nn.Embedding, which is a thin convenience layer over a learnable tensor with integer indexing: table(idx) is table.weight[idx] with gradients wired up.
The line that does the work is bigram.py#L66: nn.Embedding(vocab_size, vocab_size). Both arguments are 65, and it is worth being clear that they are 65 for two unrelated reasons. The first 65 is the number of rows — one per possible input token, because the model looks itself up by identity. The second 65 is the width of each row, and it is 65 because the row is being interpreted as logits over the vocabulary. In this model those two numbers coincide, and the coincidence is what makes the model a bigram model rather than something more expressive. In the final gpt.py the two are pulled apart: the table becomes nn.Embedding(vocab_size, n_embd) with n_embd = 384, and a separate nn.Linear(n_embd, vocab_size) projects back out at the end. Everything between those two layers is the Transformer. Watching that seam open up is one useful way to read the rest of the lecture.
The whole model has 65 × 65 = 4,225 parameters. For comparison, the model at 97:49 has about 10 million.
(B, T, C), and the shape argument you will hit again
Indexing the table with the batch produces the tensor naming convention that governs the next ninety minutes. Feed in idx of shape (B, T) — batch by time — and PyTorch returns one row per element, giving (B, T, C): batch, time, channels. In the notebook that is (4, 8, 65); in bigram.py, with batch_size = 32, it is (32, 8, 65). Karpathy uses B, T and C constantly and rarely re-explains them, so it is worth over-learning them here.
Then comes the first genuine friction of the video, and it is not conceptual — it is an API convention. F.cross_entropy accepts multi-dimensional input, but it expects the class dimension second: (N, C, d1, ..., dk). Our logits are (B, T, C), with C last. Calling it directly fails.
There are two ways out. You can transpose to (B, C, T) and hand PyTorch its preferred layout. Karpathy takes the other route: collapse batch and time into a single flat axis, so the tensor becomes an ordinary 2-D (N, C) classification problem with N = B·T examples. That reshape is the four lines at bigram.py#L76-L79. It is the better choice for a teaching codebase, because it makes explicit what is actually going on: the 32 sequences × 8 positions are 256 completely independent supervised examples, and the loss is their mean. Nothing in a decoder-only language model's loss couples one position to another — the coupling all lives in the forward pass.
| Tensor | Notebook shape | bigram.py shape | What it holds |
|---|---|---|---|
| idx / targets | (4, 8) | (32, 8) | integer token ids, 0–64 |
| logits out of the table | (4, 8, 65) | (32, 8, 65) | unnormalised next-token scores |
| logits.view(B*T, C) | (32, 65) | (256, 65) | one row per training example |
| targets.view(B*T) | (32,) | (256,) | the correct class index for each row |
| loss | scalar | scalar | mean negative log likelihood |
Note the shadowing at L77: logits is rebound to the flattened version, so the tensor the function returns in the training path is 2-D, while in the generation path (no targets) it is 3-D. Nothing downstream cares, because training only reads loss and generation never passes targets — but it is a real inconsistency in the return type, and it survives unchanged into gpt.py.
What the loss should be, and what it is
Before training, the table is random, so the model's predictions ought to be roughly uniform over 65 symbols. The negative log likelihood of a uniform guess is −ln(1/65) = ln 65 ≈ 4.174. The measured initial loss is 4.87.
That gap is worth sitting with, because it is the first instance of a diagnostic habit that runs through the whole Zero-to-Hero series: always know what number a broken model should produce. 4.87 > 4.17 means the random initialisation is not uniform — PyTorch initialises nn.Embedding from a unit normal, so each row has logits spread over roughly ±2, which after softmax is a peaked, confidently-wrong distribution. Confident and wrong scores worse than agnostic. At this scale it costs a few dozen steps of training and nothing more, which is why Karpathy notes it and moves on. At depth it becomes a real problem, and it is exactly why gpt.py later grows an _init_weights method scaling everything to std 0.02 — a method he flags in a comment as "not covered in the original GPT video."
The other number worth holding: a bigram model on tiny Shakespeare bottoms out around 2.5. That is not a training failure, it is the model class's ceiling. Knowing one character genuinely does not tell you much about the next one.
Generation: one character at a time, deliberately overbuilt
Sampling is a loop that runs max_new_tokens times, and each pass does five things: run the forward pass on the current context; slice out the final time step with logits[:, -1, :], turning (B, T, C) into (B, C); softmax over the last axis to get probabilities; draw one sample per batch row with torch.multinomial, giving (B, 1); concatenate it onto the running sequence along dim=1, giving (B, T+1). The context grows by one every iteration, which is what "autoregressive" means in code.
Two design choices in that loop deserve calling out. The first is that it samples rather than taking the argmax. Greedy decoding from a bigram table would produce a short cycle almost immediately — the highest-probability successor of the highest-probability successor loops. Sampling is what makes the output look like text-shaped noise instead of "e the the the". It is also why the model gives a different answer every run, which Karpathy uses at the very start of the video to explain why ChatGPT does too.
The second is that generate() feeds the model the entire accumulated context and then throws away all but the last position's prediction. For a bigram model this is pure waste: only the final character can possibly matter, and by token 300 the model is doing 300 table lookups to use one. Karpathy writes it this way on purpose — he wants the function to be correct without modification once the model does look at history.
obviously it's garbage and the reason it's garbage is because this is a totally random model Karpathy on the untrained sample, 33:26
The training loop that never changes
Four lines, in a fixed order: sample a batch, compute the loss, zero the gradients, backpropagate, step. Karpathy switches from the SGD used in the earlier makemore videos to AdamW, which is what everyone actually trains Transformers with, and picks learning_rate = 1e-2 — roughly thirty times the 3e-4 he calls a good default. A 4,225-parameter model can take a learning rate that would destroy a real network; a well-conditioned lookup table is about the easiest optimisation problem there is.
Which raises a point Karpathy skips past. Gradient descent on this model is unnecessary. The optimum of cross-entropy for a table of independent rows has a closed form: count every bigram in the training text, normalise each row, take the log. That is precisely how the first makemore video does it, and it reaches the same loss in one pass over the data with no optimiser at all. He uses AdamW here because the loop needs to exist for the next ninety minutes, not because the bigram model needs it — the fastest possible route to the correct answer would have taught you nothing reusable. Worth knowing so that "the bigram model was trained by gradient descent" doesn't get filed as a fact about bigram models.
obviously this is a very simple model because the tokens are not talking to each other the pivot into self-attention, 37:36
What porting to a script actually added
The last four minutes convert the notebook into bigram.py, 122 lines, and three things appear that were not in the notebook.
A device flag. device = 'cuda' if torch.cuda.is_available() else 'cpu', and then three places that have to respect it: the batches after they are stacked (L43), the model's parameters via model.to(device) (L99), and the seed context for generation (L121). Miss any one and you get a device-mismatch error rather than silently wrong results, which is the merciful failure mode. Note that the full dataset stays on the CPU and only each batch is moved — for a 1 MB corpus that is arbitrary, but it is the pattern that scales.
A real evaluation. Printing loss.item() inside the training loop reports the loss on one batch of 32 sequences, which bounces around by a tenth or more depending on which chunks got sampled. estimate_loss() averages 200 batches per split and returns both train and validation numbers, so the printed figure moves smoothly and the train/val gap is visible. It costs 400 extra forward passes every 300 steps — about 11 evaluations over the 3,000-step run.
Two correctness wrappers around that evaluation. The functional loss doesn't care, but the decorator @torch.no_grad() tells PyTorch not to build the autograd graph, so no intermediate activations are retained — a pure memory and speed win, with no effect on the numbers. The model.eval() / model.train() pair around the loop is a no-op for this model, since it contains only an embedding, and Karpathy says so explicitly. It matters from 97:49 onward, when dropout arrives: dropout must be off during evaluation, and the mode flag is what turns it off. He writes it now so that it is already correct later. This is the same discipline as the overbuilt generate() — build the scaffolding once, in its final form, while the model is small enough that you can see it is right.
| Hyperparameter | bigram.py | gpt.py (for contrast) | Why |
|---|---|---|---|
| batch_size | 32 | 64 | notebook used 4; bumped when training starts |
| block_size | 8 | 256 | irrelevant to a bigram — it only sets how many examples per row |
| max_iters | 3000 | 5000 | enough for 4,225 parameters to converge |
| eval_interval / eval_iters | 300 / 200 | 500 / 200 | 11 evaluations, each averaging 200 batches per split |
| learning_rate | 1e-2 | 3e-4 | tiny model tolerates 30x the usual Transformer LR |
| final loss | ≈ 2.5 | ≈ 1.48 val | the gap self-attention buys you |
The code at the end of this part
The model class, plus the training loop and the generation call that surround it. This is bigram.py lines 60–122, verbatim.
# super simple bigram model
class BigramLanguageModel(nn.Module):
def __init__(self, vocab_size):
super().__init__()
# each token directly reads off the logits for the next token from a lookup table
self.token_embedding_table = nn.Embedding(vocab_size, vocab_size)
def forward(self, idx, targets=None):
# idx and targets are both (B,T) tensor of integers
logits = self.token_embedding_table(idx) # (B,T,C)
if targets is None:
loss = None
else:
B, T, C = logits.shape
logits = logits.view(B*T, C)
targets = targets.view(B*T)
loss = F.cross_entropy(logits, targets)
return logits, loss
def generate(self, idx, max_new_tokens):
# idx is (B, T) array of indices in the current context
for _ in range(max_new_tokens):
# get the predictions
logits, loss = self(idx)
# focus only on the last time step
logits = logits[:, -1, :] # becomes (B, C)
# apply softmax to get probabilities
probs = F.softmax(logits, dim=-1) # (B, C)
# sample from the distribution
idx_next = torch.multinomial(probs, num_samples=1) # (B, 1)
# append sampled index to the running sequence
idx = torch.cat((idx, idx_next), dim=1) # (B, T+1)
return idx
model = BigramLanguageModel(vocab_size)
m = model.to(device)
# create a PyTorch optimizer
optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate)
for iter in range(max_iters):
# every once in a while evaluate the loss on train and val sets
if iter % eval_interval == 0:
losses = estimate_loss()
print(f"step {iter}: train loss {losses['train']:.4f}, val loss {losses['val']:.4f}")
# sample a batch of data
xb, yb = get_batch('train')
# evaluate the loss
logits, loss = model(xb, yb)
optimizer.zero_grad(set_to_none=True)
loss.backward()
optimizer.step()
# generate from the model
context = torch.zeros((1, 1), dtype=torch.long, device=device)
print(decode(m.generate(context, max_new_tokens=500)[0].tolist()))
Line by line, the parts that are easy to skim past:
- L63 — __init__ takes vocab_size as an argument even though it is a module-level global. In gpt.py the argument is dropped and the global is read directly, which is more consistent with how n_embd and friends are handled.
- L73–79 — the branch on targets is None. This exists solely so that generate() can call self(idx) with no labels. It reads as defensive coding but it is load-bearing.
- L89 — logits[:, -1, :]. Negative-one indexing on the time axis: the prediction made at the last position is the prediction for the position after it. Off-by-one here is the single most common bug when people reimplement this from memory.
- L91 — F.softmax(logits, dim=-1). The dim is not optional and not guessable; softmaxing the batch axis by accident produces plausible-looking garbage that trains to nothing.
- L98–99 — model and m are the same object; .to(device) mutates modules in place and returns self. The two names are an artefact of the notebook, and they persist into gpt.py.
- L116 — zero_grad(set_to_none=True) frees the gradient tensors instead of filling them with zeros. Slightly faster, slightly less memory, and it makes "I forgot to call backward" fail loudly with None rather than silently with stale zeros.
- L121–122 — the seed context is a (1, 1) tensor holding 0, which decodes to a newline because chars is sorted and newline is the lowest codepoint in tiny Shakespeare. Starting mid-line would work too; newline is just the most natural "beginning of a document" token available when you have no dedicated BOS symbol.
Where people get stuck
- "Why is the loss averaged over all 8 positions, not just the last one?" — because each of the 8 positions is a separate labelled example. y is x shifted one step left, so position t in a row predicts position t+1, and a batch of 32 sequences of length 8 is 256 examples. This is the free lunch of causal language modelling: one forward pass over T tokens yields T supervised gradients, not one. It is also why the mask in part 4 has to be exactly triangular — if position 3 could see position 5, its label would be leaking into its input.
- "Why .view(B*T, C) and not a transpose?" — both work. F.cross_entropy accepts (B, C, T) directly, so logits.transpose(1, 2) is a valid one-liner. Karpathy flattens instead because the flat form is the honest description of the problem: 256 independent 65-way classifications. Also note targets.view(B*T) could be targets.view(-1) — he writes the explicit form, and says so, to keep the shapes readable.
- "Does generate() break when the context exceeds block_size?" — not here, and that is a trap. block_size = 8, yet the script generates 500 tokens and feeds all of them back in. The bigram model has no positional table and no attention mask, so a context of length 500 is harmless — it just wastes compute. From 60:18 onward, once a position_embedding_table of exactly block_size rows exists, the same call would index out of range. That is what idx_cond = idx[:, -block_size:] is for. If you write your own version and skip that line, you get a crash roughly 250 tokens into your first sample.
- "4.87 versus 4.17 — is something wrong?" — no. Uniform-over-65 gives 4.174; a randomly initialised embedding gives peaked distributions that are confidently wrong, which scores worse. The number to be alarmed by is one that is much lower than 4.17 at step zero, since that usually means labels have leaked into the input.
Go deeper, verified
- bigram.py — Andrej Karpathy (2023) · the 122-line script this part produces, in full. Diff it against gpt.py to see exactly how little outside the model class changes.
- The lecture's Colab notebook — Andrej Karpathy (2023) · the notebook being screen-shared. The right place to run the exercises below; it already has tiny Shakespeare loaded.
- torch.nn.functional.cross_entropy — PyTorch docs · the shape contract that forces the .view(). Read the Shape section; it also documents ignore_index, which Karpathy's EX2 asks you to use.
- torch.multinomial — PyTorch docs · row-wise categorical sampling. Confirms it takes unnormalised weights and samples per row, which is why one call handles all B batch elements.
- The spelled-out intro to language modeling: building makemore — Andrej Karpathy (2022) · the bigram model at four times the depth, including the count-and-normalise closed form and why it gives the same answer as gradient descent. This is the video he defers to at 22:11.
- Decoupled Weight Decay Regularization — Loshchilov & Hutter (2017) · the paper that AdamW is. Field map extra — Karpathy just says "much more advanced and popular"; this is why the W is there and why it is the default for Transformers.
- The Curious Case of Neural Text Degeneration — Holtzman et al. (2019) · field map extra. Why plain multinomial sampling is not what production decoders do, and where top-k and nucleus sampling come from. nanoGPT's generate has temperature and top-k for exactly this reason; the lecture's version has neither.
Exercises
None of Karpathy's EX1–EX4 land in this part — they all need the finished Transformer. These are local to the bigram baseline, and all three run in the Colab in a couple of minutes each.
- Beat gradient descent by countingcode — build the closed-form optimum and check the optimiser found it. (1) Allocate N = torch.zeros(65, 65) and, in one pass over train_data, increment N[data[i], data[i+1]]. (2) Add 1 to every cell for smoothing, then normalise rows to sum to 1. (3) Take torch.log and copy the result into model.token_embedding_table.weight.data. (4) Call estimate_loss(). A good answer reports a validation loss within about 0.02 of the trained model's ~2.5, and explains why: cross-entropy on a table of independent rows is minimised by the empirical conditional distribution, so 3,000 AdamW steps are an expensive way to compute a histogram. State what the smoothing constant changes and why zero-count bigrams need it.
- Break generate(), then fix it the way gpt.py doescode — (1) Inside the sampling loop, print idx.shape every 50 tokens and confirm it grows to (1, 501) while block_size is 8. (2) Add pos = nn.Embedding(block_size, vocab_size) to the model and add pos(torch.arange(T, device=device)) to the logits in forward. (3) Re-run generation and record the exact token count at which it raises. (4) Add idx_cond = idx[:, -block_size:] and confirm it survives 500 tokens. A good answer names the error (an index out of range in the positional embedding, at T = 9) and states the invariant the crop enforces: the model can only ever be handed a context it has position vectors for.
- Calibrate the loss — no code strictly required, but a two-line experiment settles it. Predict, before running, the initial loss for (a) an embedding table initialised to all zeros, (b) the default unit-normal init, (c) an init scaled to std 0.02 as gpt.py uses. Then measure all three. A good answer gets 4.174 exactly for the zero case, explains why the normal init overshoots to ~4.87, shows that std 0.02 lands very close to 4.174, and connects that to why the initialisation trick matters far more for a 6-layer network than for a lookup table.