LET'S BUILD GPT // FIELD MAP
← all maps
ANDREJ KARPATHY · 20231 h 56 min · 31 chapters · 8 parts here

Let's build GPT, mapped

Karpathy writes a GPT from an empty file to a 10-million-parameter Transformer that babbles Shakespeare, and explains every line. This map splits the lecture into eight parts you can read before, during or after watching: what each stretch answers, the tensor shapes, the code it arrives at, where people get stuck, and exercises. Every timestamp opens the video at that moment.

Tick a part when you've worked through it; this browser remembers.

01 · 00:00–42:13

Baseline: the problem and the simplest model

Predict the next character of Shakespeare. Set up the data, the tokens, the batches, and a bigram model that already trains, evaluates and generates. Everything after this is making that model look further back.

P1
▶ 00:0022 minintro · reading the data · tokenization, train/val split · data loader
Why a character-level model, what the 65-symbol vocabulary is, how a block of text becomes 8 training examples at once, and what B and T mean for the rest of the lecture.
P2
▶ 22:1120 minbigram model · loss · generate() · training loop · port to a script
A lookup table as a language model, cross-entropy as the score, sampling one character at a time, and the training loop that every later version reuses unchanged.
02 · 42:13–79:11

Self-attention, built up in four versions

The heart of the lecture. Karpathy arrives at attention by making "average the past" progressively smarter: a for-loop, then a matrix multiply, then softmax, then learned queries and keys. Then six notes on what the thing actually is.

P3
▶ 42:1320 minversion 1 · matmul as weighted aggregation · version 2 · version 3 softmax · cleanup · positional encoding
The lower-triangular matrix that turns "average everything before me" into one matrix multiply, why softmax of a masked zero matrix is the same thing, and where positions enter.
P4
▶ 62:0010 minversion 4: queries, keys, values
Every token emits a query and a key; affinities are their dot products; the mask keeps the future out; values are what gets aggregated. Ten minutes that the whole field rests on.
P5
▶ 71:388 mincommunication · no notion of space · no cross-batch talk · encoder vs decoder · self vs cross · why divide by √head_size
Attention as message passing on a directed graph, why it needs positional encodings, what changes when you drop the mask, and the variance argument behind the scaling.
03 · 79:11–108:53

Assembling the Transformer

One head becomes many, a feed-forward layer gives tokens time to think, residual connections and LayerNorm make it deep, dropout makes it big. Then the same thing as nanoGPT writes it.

P6
▶ 79:1119 minsingle head in the network · multi-head · feed-forward · residual connections · LayerNorm
Why heads are concatenated, why the MLP is 4× wide, the gradient superhighway that residuals create, and the pre-norm placement that differs from the 2017 paper.
P7
▶ 97:4911 minscaling up + dropout · encoder vs decoder vs both · nanoGPT walkthrough
The 10M-parameter run (val loss 1.48), what dropout does at this size, which half of the original Transformer this is, and how nanoGPT batches all heads into one matmul.
04 · 108:53–116:20

Back to ChatGPT

The distance between a Shakespeare babbler and ChatGPT is pretraining at scale plus fine-tuning and RLHF. Karpathy sketches it; this part fills in what has changed since 2023.

P8
▶ 108:537 minGPT-3 · pretraining vs fine-tuning · RLHF · conclusions · the four exercises
Document completer to assistant in three stages, why a 10M model and a 175B model are the same code, and the four exercises Karpathy sets, expanded.
05

Karpathy's four exercises

From the video description. Each part also sets its own smaller exercises; these are the big ones.

  1. EX1 — the n-dimensional tensor mastery challenge. Combine Head and MultiHeadAttention into one class that processes all heads in parallel, treating heads as another batch dimension. The answer is in nanoGPT's model.py. Covered in P7.
  2. EX2 — train it on your own data. Any corpus you like; the advanced version is a GPT that adds two numbers (predict the digits of the sum in reverse, serve random problems instead of train.bin, mask the loss on the prompt with ignore_index=-1). Then the "swole doge" version: a calculator for + − × ÷, which may need chain-of-thought traces. Covered in P8.
  3. EX3 — pretrain, then fine-tune. Find a dataset so large the train/val gap disappears, pretrain on it, then fine-tune on tiny Shakespeare with fewer steps and a lower learning rate. Does the pretrained start reach a lower validation loss? Covered in P8.
  4. EX4 — implement one more feature from the papers. Read some Transformer papers and add one change people use. Does it help your GPT? The parts suggest candidates: RoPE (P3), FlashAttention (P4), grouped-query attention (P6), RMSNorm and SwiGLU (P6), weight tying (P7).
06

Sources