LET'S BUILD GPT // FIELD MAP
← field map
PART 08 · BACK TO CHATGPT108:53–116:20 · 7 min

Back to ChatGPT: pretraining, fine-tuning, RLHF, and what to do next

Andrej Karpathy · Let's build GPT (2023) · part 08 of 8

Transcript: this part, with timestamps

TL;DR — The model you just built and GPT-3 are the same code with different constants: 10.8 million parameters against 175 billion, 300 thousand tokens against 300 billion. Nothing in gpt.py would have to change architecturally to be the big one — what changes is infrastructure and data. But scale alone gives you a document completer, not an assistant: ask it a question and it may answer with more questions, because completing a document is literally all it was trained to do. Turning that into ChatGPT is a second stage — supervised fine-tuning on demonstration data, then a reward model trained on human rankings, then reinforcement learning against that reward model. Karpathy sketches this in four minutes and says the data is internal to OpenAI and hard to replicate; that last claim is the one part of the lecture that has genuinely expired.

This is the bookend. The lecture opened on ChatGPT writing a haiku and asked what is under the hood; 108 minutes later there is a working answer, and Karpathy walks back up the ladder to show how far the thing on your screen is from the thing in the browser tab. The value of these seven minutes is calibration: which of the gap is architecture (almost none), which is scale (a factor of roughly 10,000 to 1,000,000 depending on what you count), and which is a training objective you have not written yet (all of alignment). Get that split right and you know exactly which parts of the field are downstream of what you just implemented.

Outline, with timestamps

Stage one is what you already have

The framing to hold onto: training a system like ChatGPT is pretraining followed by fine-tuning, and the entire lecture — 116 minutes, every version of attention, every residual connection — was an implementation of stage one on a toy corpus. Not an analogy for stage one. The actual objective, the actual architecture, the actual loss. Predict the next token, average cross-entropy over every position in the batch, backprop, AdamW. If you swapped tiny Shakespeare for a scrape of the web and turned the constants up, you would be running a pretraining job.

That is a stronger claim than it sounds, and it is worth checking rather than taking on faith. The forward pass in GPTLanguageModel.forward is: embed tokens, add position embeddings, run n_layer blocks of (LayerNorm → multi-head causal self-attention → residual, LayerNorm → MLP → residual), a final LayerNorm, a linear projection to vocabulary logits, cross-entropy against the inputs shifted by one. That list is also an accurate description of GPT-3's forward pass, and of most decoder-only models shipped since. The differences that exist at frontier scale are real but small in kind: a different tokenizer, learned positions replaced by rotary embeddings, ReLU replaced by GELU or SwiGLU, LayerNorm replaced by RMSNorm, attention kernels fused for memory bandwidth. None of them changes what the model computes in any way you would call architectural.

The same code with bigger constants

Karpathy makes the point by putting his printed parameter count next to Table 2.1 of the GPT-3 paper — the table listing every model size they trained. Its columns are, in order, parameters, layers, model dimension, heads, head size, batch size, learning rate. Those are precisely n_layer, n_embd, n_head, head_size, batch_size and learning_rate from the top of gpt.py. The paper is describing your script's hyperparameter block with different numbers in it.

Knobthis lecture's GPTGPT-3 175Bratio
parameters10,788,929 (prints as 10.788929 M)175 B~16,000×
n_layer69616×
n_embd / d_model38412,28832×
n_head69616×
head size64 (384 ÷ 6)1282×
block_size (context)256 characters2,048 BPE tokens8× in tokens, far more in text
batch_size64 sequences = 16,384 tokens3.2 M tokens~200×
learning_rate3e-40.6e-4smaller model, larger LR
vocabulary65 characters50,257 BPE merges—
training data~1 M characters ≈ 300 K BPE tokens300 B tokens~1,000,000×
hardware / timeone A100, ~15 min (100:21)thousands of GPUs, weeks—

Two of those rows deserve a second look. First, the token count. Tiny Shakespeare is about a million characters, and because this model is character-level, a million characters is a million training tokens. But OpenAI's models are not character-level — they use subword chunks, roughly 50 K of them, so the same file is only around 300 K tokens in their vocabulary. That is the honest apples-to-apples number, and it is what makes the ratio to GPT-3's 300 B tokens come out at almost exactly a millionfold. Karpathy then adds the line that has aged best in the whole part: 300 B would not be considered large today, you would be going to a trillion and above. Written in January 2023, and by 2024 open models were routinely pretrained on ten to fifteen trillion tokens.

Second, the batch size. The captions garble the number here — the on-screen comparison is our batch_size = 64 against GPT-3's 3.2 M, and 64 is easy to mishear as 65, which is also a number in this script (the vocabulary size). They are unrelated. Note too that the two batch sizes are not measured in the same unit: ours is 64 sequences, which at block_size = 256 is 16,384 tokens per step; GPT-3's is quoted directly in tokens. Comparing 64 to 3,200,000 overstates the gap by a factor of 256.

The thing to carry away from the comparison: scale is a data and infrastructure problem, not a modelling one. The reason you cannot train GPT-3 tonight is not that you are missing a layer type. It is that 300 B tokens have to be collected and cleaned, and that thousands of GPUs have to exchange gradients without falling over for weeks. Every hard part of that lives in a training harness, not in a model definition — which is exactly why nanoGPT's model.py looks like what you wrote while its train.py looks nothing like your training loop (P7).

What pretraining actually gives you

Here is the part that surprises people who have only ever used a chat interface. Run pretraining to completion and you do not get something that answers questions. You get a model that continues documents, because continuing documents is the only thing the loss ever rewarded. Yours babbles Shakespeare; one trained on the web babbles the web — plausible news articles, plausible forum posts, plausible boilerplate.

Give such a model a question and its behaviour is undefined in a specific, mechanical sense. It samples the continuation that is most likely given the training distribution, and on the internet a question is very often followed by more questions — an FAQ, a quiz, a list of related searches. So it may answer your question with five more questions, or ignore it and write the rest of the article the question seemed to be embedded in. That is not a bug or a failure of intelligence. It is the model doing its job perfectly against an objective that has no concept of "helpful".

The mental correction to make: helpfulness is not an emergent property of next-token prediction, it is a separate thing you train for. Pretraining buys capability — grammar, facts, code, reasoning patterns, the shape of arguments. It buys no interface to that capability. Prompt engineering in the GPT-3 era was largely the craft of tricking a document completer into a document format whose completion happened to be the answer you wanted ("Q: … A:"). Fine-tuning replaces the trick with training.

The three-step alignment recipe

Karpathy walks through the diagram from OpenAI's ChatGPT announcement — the same three steps written up properly in the InstructGPT paper. Each step optimizes something different, and knowing which is which is most of the literacy this section is worth:

Notice the shape of the whole pipeline: it converts a model that can produce good answers into one that reliably does, using two very different data types — demonstrations for step 1, preferences for steps 2 and 3. And notice what it does not do: it adds no parameters, changes no architecture, and touches nothing you built in parts 1 through 7. The Transformer is the substrate; alignment is a training regime applied to it.

Karpathy hedges honestly at 113:53: there are more steps in between, a lot of the data is internal to OpenAI, and this stage is much harder to replicate than pretraining — nanoGPT deliberately covers only the pretraining half. In January 2023 that was simply true. It is the one claim in the lecture that has since been overtaken by events.

What changed after January 2023

Everything above still describes the shape of post-training. The specific algorithm in step 3, and the "you cannot replicate this" caveat, are what moved.

The reward model and the RL loop got optional. Direct Preference Optimization (Rafailov et al., May 2023) showed that the RLHF objective has a closed-form reparameterization: you can fold the reward model into the policy and train directly on preference pairs with an ordinary supervised-looking loss, comparing the model's log-probabilities against a frozen copy of itself. No reward network, no sampling loop, no PPO. It is a few dozen lines on top of an SFT script, which is why nearly every open-weights chat model between 2023 and 2025 used DPO or one of its variants. The three-box diagram in the video collapses to two.

Then RL came back, cheaper and with better rewards. GRPO (introduced in DeepSeekMath, Feb 2024) is PPO with the value network deleted: sample a group of completions for one prompt, and use the group's own mean and spread of rewards as the baseline for the advantage. That removes a second full-sized model from GPU memory and made policy-gradient post-training practical for labs without OpenAI's budget. Meanwhile Tulu 3 (Nov 2024) named and popularized RLVR — reinforcement learning with verifiable rewards. Where step 2 learns a reward model to imitate human taste, RLVR skips it entirely on tasks where correctness is checkable by a program: does the final answer match, do the unit tests pass. The reward is a boolean from a checker rather than a float from a network, so there is nothing to over-optimize against.

And that turned out to produce reasoning. DeepSeek-R1 (Jan 2025) ran large-scale RL with verifiable rewards on math and code, and reported that long chain-of-thought behaviour — self-checking, backtracking, spending more tokens on harder problems — emerged from the RL rather than being demonstrated in SFT data. Its R1-Zero variant reached this from a base model with no supervised fine-tuning at all. This is the lineage of the "reasoning models" that appeared across the industry through 2025, and it is a direct descendant of the third box in Karpathy's diagram.

The replication caveat expired. Open preference datasets, open post-training recipes (Tulu 3 published data, code and evaluations), and libraries like TRL that implement SFT, DPO and GRPO as drop-in trainers mean the second stage is now something a motivated individual can run on rented GPUs. Karpathy himself closed the loop with nanochat (2025), which is nanoGPT plus tokenizer, midtraining, SFT, optional RL and a chat UI — the full pipeline of this section, in the same readable style, for about a hundred dollars of compute. If you want the updated version of these seven minutes in the author's own voice, his 2025 talk Deep Dive into LLMs like ChatGPT spends three hours on what this part covers in four minutes.

Go forth and transform.— Karpathy, signing off, 115:55

The code at the end of this part

No new code is written in this part — the script is finished. What the part does is reinterpret two blocks you already have as the dials that set scale, so those are the ones worth re-reading with GPT-3's table next to them.

# hyperparameters
batch_size = 64 # how many independent sequences will we process in parallel?
block_size = 256 # what is the maximum context length for predictions?
max_iters = 5000
eval_interval = 500
learning_rate = 3e-4
device = 'cuda' if torch.cuda.is_available() else 'cpu'
eval_iters = 200
n_embd = 384
n_head = 6
n_layer = 6
dropout = 0.2
# ------------

# ... 180 lines later ...

model = GPTLanguageModel()
m = model.to(device)
# print the number of parameters in the model
print(sum(p.numel() for p in m.parameters())/1e6, 'M parameters')

Where people get stuck

Go deeper, verified

Exercises

These are Karpathy's own EX2, EX3 and EX4 from the video description, expanded into steps. (EX1, the batched-heads challenge, belongs to P7.) All three run on top of the finished gpt.py; the first two want a GPU, the third does not.

  1. EX2 — train it on your own data, then teach it to addcode — Karpathy's exercise, in two escalating versions.
    (a) The easy half: replace input.txt with any corpus you like — your own writing, a codebase, song lyrics — and rerun. Nothing else changes; vocab_size recomputes itself from sorted(list(set(text))). Watch the train/val gap: a small corpus will overfit long before 5,000 iterations, and you will see val loss turn back upward while train loss keeps falling.
    (b) The advanced half, which is the interesting one: train a GPT to compute a+b=c. Steps: (1) replace get_batch with a generator that emits random problems as character strings — no train.bin, no held-out file, just fresh samples each step, so there is no train/val gap to speak of. (2) Write the digits of c in reverse order, because the carry propagates right-to-left and a causal model can only condition on what it has already emitted; left-to-right output would require the model to know the high digit before computing the low ones. (3) Mask the loss over the prompt positions — the a+b= part is given, not predicted — by setting those targets to -1 and passing ignore_index=-1 to F.cross_entropy. (4) Train and evaluate on held-out digit lengths. A good answer reports exact-match accuracy on unseen operands, and notes where it fails: usually longer numbers than were ever trained on. The "swole doge" extension is a full calculator for + − × ÷, which likely needs chain-of-thought traces in the training data — the model must be allowed to write intermediate steps, because a single forward pass has a fixed depth and long multiplication does not.
  2. EX3 — pretrain on something big, then fine-tune on Shakespearecode — Karpathy's exercise, and a hands-on version of this part's whole argument. Steps: (1) find a corpus large enough that train and val loss stay glued together for the whole run — a sample of FineWeb, all of Project Gutenberg, a large code dump. No visible gap means you are not memorizing and every step is buying real generalization. (2) Pretrain gpt.py on it, keeping the character vocabulary — but build the vocabulary from the union of both corpora so the fine-tuning stage does not meet unknown characters. (3) Save the weights with torch.save(model.state_dict(), ...). (4) Fine-tune: load that checkpoint, swap in tiny Shakespeare, drop the learning rate by roughly 10× (3e-5 rather than 3e-4) and run far fewer iterations. (5) Compare final validation loss against the from-scratch run's 1.48. A good answer reports both curves on the same axes and says how many fine-tuning steps the pretrained model needed to pass the from-scratch model's best loss — that number, not the final loss, is the measure of what pretraining bought.
  3. EX4 — implement one thing from the paperscode — Karpathy's exercise: read some Transformer papers, add one feature people actually use, and measure whether it helps. Steps: (1) pick one change and only one — candidates the earlier parts flag are rotary position embeddings (P3), FlashAttention or PyTorch's scaled_dot_product_attention (P4), grouped-query attention, RMSNorm or SwiGLU (P6), and weight tying between the embedding table and lm_head (P7). (2) Before touching anything, fix a baseline: same seed, same max_iters, record val loss at every eval_interval, and note wall-clock time and peak torch.cuda.max_memory_allocated(). (3) Make the change. (4) Rerun identically and plot both curves. A good answer is honest about which axis moved: several of these (FlashAttention, GQA) are not meant to improve loss at all — they buy speed or memory at equal loss, and reporting "no improvement" for them is the correct result. State the comparison you are making before you run it, or you will find a way to call any outcome a win.
Next: Karpathy's four exercises, collected on the map · Back to the map.