Back to ChatGPT: pretraining, fine-tuning, RLHF, and what to do next
Transcript: this part, with timestamps
This is the bookend. The lecture opened on ChatGPT writing a haiku and asked what is under the hood; 108 minutes later there is a working answer, and Karpathy walks back up the ladder to show how far the thing on your screen is from the thing in the browser tab. The value of these seven minutes is calibration: which of the gap is architecture (almost none), which is scale (a factor of roughly 10,000 to 1,000,000 depending on what you count), and which is a training objective you have not written yet (all of alignment). Get that split right and you know exactly which parts of the field are downstream of what you just implemented.
Outline, with timestamps
- 108:53 — Two stages: everything so far was pretraining; ChatGPT needs a fine-tuning stage on top.
- 109:12 — Our pretraining in miniature: print(sum(p.numel() ...)) says ~10 million parameters.
- 109:42 — Character tokens vs. BPE: 1 M characters is only ~300 K tokens in OpenAI's 50 K vocabulary.
- 110:14 — The GPT-3 table: 175 B parameters, and the columns are the variables at the top of gpt.py.
- 110:46 — 300 billion training tokens, a millionfold increase, and "not even large by today's standards".
- 111:17 — Why pretraining alone leaves you with a document completer, not an answerer.
- 112:20 — Alignment step 1: supervised fine-tuning on maybe thousands of question/answer demonstrations.
- 113:20 — Steps 2 and 3: raters rank responses, a reward model learns the ranking, PPO optimizes against it.
- 113:53 — The hedge: this data is internal, this stage is much harder to replicate, nanoGPT covers pretraining only.
- 114:32 — Conclusions: a decoder-only Transformer, ~200 lines of code, the repo with its git log.
- 115:24 — What was not covered: every flavour of fine-tuning, from plain SFT to reward-model rounds.
- 115:55 — "Go forth and transform."
Stage one is what you already have
The framing to hold onto: training a system like ChatGPT is pretraining followed by fine-tuning, and the entire lecture — 116 minutes, every version of attention, every residual connection — was an implementation of stage one on a toy corpus. Not an analogy for stage one. The actual objective, the actual architecture, the actual loss. Predict the next token, average cross-entropy over every position in the batch, backprop, AdamW. If you swapped tiny Shakespeare for a scrape of the web and turned the constants up, you would be running a pretraining job.
That is a stronger claim than it sounds, and it is worth checking rather than taking on faith. The forward pass in GPTLanguageModel.forward is: embed tokens, add position embeddings, run n_layer blocks of (LayerNorm → multi-head causal self-attention → residual, LayerNorm → MLP → residual), a final LayerNorm, a linear projection to vocabulary logits, cross-entropy against the inputs shifted by one. That list is also an accurate description of GPT-3's forward pass, and of most decoder-only models shipped since. The differences that exist at frontier scale are real but small in kind: a different tokenizer, learned positions replaced by rotary embeddings, ReLU replaced by GELU or SwiGLU, LayerNorm replaced by RMSNorm, attention kernels fused for memory bandwidth. None of them changes what the model computes in any way you would call architectural.
The same code with bigger constants
Karpathy makes the point by putting his printed parameter count next to Table 2.1 of the GPT-3 paper — the table listing every model size they trained. Its columns are, in order, parameters, layers, model dimension, heads, head size, batch size, learning rate. Those are precisely n_layer, n_embd, n_head, head_size, batch_size and learning_rate from the top of gpt.py. The paper is describing your script's hyperparameter block with different numbers in it.
| Knob | this lecture's GPT | GPT-3 175B | ratio |
|---|---|---|---|
| parameters | 10,788,929 (prints as 10.788929 M) | 175 B | ~16,000× |
| n_layer | 6 | 96 | 16× |
| n_embd / d_model | 384 | 12,288 | 32× |
| n_head | 6 | 96 | 16× |
| head size | 64 (384 ÷ 6) | 128 | 2× |
| block_size (context) | 256 characters | 2,048 BPE tokens | 8× in tokens, far more in text |
| batch_size | 64 sequences = 16,384 tokens | 3.2 M tokens | ~200× |
| learning_rate | 3e-4 | 0.6e-4 | smaller model, larger LR |
| vocabulary | 65 characters | 50,257 BPE merges | — |
| training data | ~1 M characters ≈ 300 K BPE tokens | 300 B tokens | ~1,000,000× |
| hardware / time | one A100, ~15 min (100:21) | thousands of GPUs, weeks | — |
Two of those rows deserve a second look. First, the token count. Tiny Shakespeare is about a million characters, and because this model is character-level, a million characters is a million training tokens. But OpenAI's models are not character-level — they use subword chunks, roughly 50 K of them, so the same file is only around 300 K tokens in their vocabulary. That is the honest apples-to-apples number, and it is what makes the ratio to GPT-3's 300 B tokens come out at almost exactly a millionfold. Karpathy then adds the line that has aged best in the whole part: 300 B would not be considered large today, you would be going to a trillion and above. Written in January 2023, and by 2024 open models were routinely pretrained on ten to fifteen trillion tokens.
Second, the batch size. The captions garble the number here — the on-screen comparison is our batch_size = 64 against GPT-3's 3.2 M, and 64 is easy to mishear as 65, which is also a number in this script (the vocabulary size). They are unrelated. Note too that the two batch sizes are not measured in the same unit: ours is 64 sequences, which at block_size = 256 is 16,384 tokens per step; GPT-3's is quoted directly in tokens. Comparing 64 to 3,200,000 overstates the gap by a factor of 256.
What pretraining actually gives you
Here is the part that surprises people who have only ever used a chat interface. Run pretraining to completion and you do not get something that answers questions. You get a model that continues documents, because continuing documents is the only thing the loss ever rewarded. Yours babbles Shakespeare; one trained on the web babbles the web — plausible news articles, plausible forum posts, plausible boilerplate.
Give such a model a question and its behaviour is undefined in a specific, mechanical sense. It samples the continuation that is most likely given the training distribution, and on the internet a question is very often followed by more questions — an FAQ, a quiz, a list of related searches. So it may answer your question with five more questions, or ignore it and write the rest of the article the question seemed to be embedded in. That is not a bug or a failure of intelligence. It is the model doing its job perfectly against an objective that has no concept of "helpful".
The mental correction to make: helpfulness is not an emergent property of next-token prediction, it is a separate thing you train for. Pretraining buys capability — grammar, facts, code, reasoning patterns, the shape of arguments. It buys no interface to that capability. Prompt engineering in the GPT-3 era was largely the craft of tricking a document completer into a document format whose completion happened to be the answer you wanted ("Q: … A:"). Fine-tuning replaces the trick with training.
The three-step alignment recipe
Karpathy walks through the diagram from OpenAI's ChatGPT announcement — the same three steps written up properly in the InstructGPT paper. Each step optimizes something different, and knowing which is which is most of the literacy this section is worth:
- Step 1 — supervised fine-tuning (SFT). Collect documents in a fixed shape: a prompt on top, an ideal response below, written by humans. Then keep training with exactly the same next-token loss you already implemented, only on this data. Nothing changes in the code — same F.cross_entropy, same optimizer, fewer steps and a lower learning rate. The dataset is small, maybe thousands of examples rather than internet-scale, and Karpathy flags the reason it works anyway: models this large are extremely sample-efficient during fine-tuning. They are not learning to write. They are learning which of the many document formats already in their head is now the relevant one.
- Step 2 — the reward model. Sample several responses to the same prompt, have human raters rank them best to worst, and train a second network to predict those rankings — a scalar score for how desirable any candidate response is. The interesting design choice is that raters compare rather than score. People are far more consistent at "A is better than B" than at "this is a 7/10", and a ranking loss turns the comparisons into a usable numeric reward.
- Step 3 — reinforcement learning against that reward. With a reward model in hand you no longer need a human in the loop, so you can optimize the language model's sampling policy to produce responses the reward model scores highly. The algorithm named in the video is PPO — Proximal Policy Optimization, a policy-gradient method (the auto-captions render it as "po", which is not a thing). Steps 2 and 3 together are what people mean by RLHF, reinforcement learning from human feedback.
Notice the shape of the whole pipeline: it converts a model that can produce good answers into one that reliably does, using two very different data types — demonstrations for step 1, preferences for steps 2 and 3. And notice what it does not do: it adds no parameters, changes no architecture, and touches nothing you built in parts 1 through 7. The Transformer is the substrate; alignment is a training regime applied to it.
Karpathy hedges honestly at 113:53: there are more steps in between, a lot of the data is internal to OpenAI, and this stage is much harder to replicate than pretraining — nanoGPT deliberately covers only the pretraining half. In January 2023 that was simply true. It is the one claim in the lecture that has since been overtaken by events.
What changed after January 2023
Everything above still describes the shape of post-training. The specific algorithm in step 3, and the "you cannot replicate this" caveat, are what moved.
The reward model and the RL loop got optional. Direct Preference Optimization (Rafailov et al., May 2023) showed that the RLHF objective has a closed-form reparameterization: you can fold the reward model into the policy and train directly on preference pairs with an ordinary supervised-looking loss, comparing the model's log-probabilities against a frozen copy of itself. No reward network, no sampling loop, no PPO. It is a few dozen lines on top of an SFT script, which is why nearly every open-weights chat model between 2023 and 2025 used DPO or one of its variants. The three-box diagram in the video collapses to two.
Then RL came back, cheaper and with better rewards. GRPO (introduced in DeepSeekMath, Feb 2024) is PPO with the value network deleted: sample a group of completions for one prompt, and use the group's own mean and spread of rewards as the baseline for the advantage. That removes a second full-sized model from GPU memory and made policy-gradient post-training practical for labs without OpenAI's budget. Meanwhile Tulu 3 (Nov 2024) named and popularized RLVR — reinforcement learning with verifiable rewards. Where step 2 learns a reward model to imitate human taste, RLVR skips it entirely on tasks where correctness is checkable by a program: does the final answer match, do the unit tests pass. The reward is a boolean from a checker rather than a float from a network, so there is nothing to over-optimize against.
And that turned out to produce reasoning. DeepSeek-R1 (Jan 2025) ran large-scale RL with verifiable rewards on math and code, and reported that long chain-of-thought behaviour — self-checking, backtracking, spending more tokens on harder problems — emerged from the RL rather than being demonstrated in SFT data. Its R1-Zero variant reached this from a base model with no supervised fine-tuning at all. This is the lineage of the "reasoning models" that appeared across the industry through 2025, and it is a direct descendant of the third box in Karpathy's diagram.
The replication caveat expired. Open preference datasets, open post-training recipes (Tulu 3 published data, code and evaluations), and libraries like TRL that implement SFT, DPO and GRPO as drop-in trainers mean the second stage is now something a motivated individual can run on rented GPUs. Karpathy himself closed the loop with nanochat (2025), which is nanoGPT plus tokenizer, midtraining, SFT, optional RL and a chat UI — the full pipeline of this section, in the same readable style, for about a hundred dollars of compute. If you want the updated version of these seven minutes in the author's own voice, his 2025 talk Deep Dive into LLMs like ChatGPT spends three hours on what this part covers in four minutes.
Go forth and transform.— Karpathy, signing off, 115:55
The code at the end of this part
No new code is written in this part — the script is finished. What the part does is reinterpret two blocks you already have as the dials that set scale, so those are the ones worth re-reading with GPT-3's table next to them.
# hyperparameters
batch_size = 64 # how many independent sequences will we process in parallel?
block_size = 256 # what is the maximum context length for predictions?
max_iters = 5000
eval_interval = 500
learning_rate = 3e-4
device = 'cuda' if torch.cuda.is_available() else 'cpu'
eval_iters = 200
n_embd = 384
n_head = 6
n_layer = 6
dropout = 0.2
# ------------
# ... 180 lines later ...
model = GPTLanguageModel()
m = model.to(device)
# print the number of parameters in the model
print(sum(p.numel() for p in m.parameters())/1e6, 'M parameters')
- batch_size and block_size — the two shape constants. Every activation in the network is (B, T, C) = (64, 256, 384); the attention matrix inside each head is (64, 256, 256) per head, and there are 36 of them (6 heads × 6 layers). This is the memory wall: attention is quadratic in T, so quadrupling the context multiplies that tensor by sixteen.
- n_embd, n_head, n_layer — the three that set parameter count. Roughly, parameters scale as n_layer × n_embd² (each block holds 4·n_embd² in attention projections plus 8·n_embd² in the 4×-wide MLP, so ≈ 12·n_embd² per layer). Check it: 12 × 384² × 6 ≈ 10.6 M, which is nearly all of the 10,788,929 the print statement reports; the embeddings and the final head make up the rest. The same formula applied to GPT-3's 96 layers of width 12,288 gives ≈ 174 B — the paper's 175 B.
- print(sum(p.numel() ...)) — the line that produces the number Karpathy reads out. p.numel() is the element count of each parameter tensor; summing over m.parameters() walks every registered submodule. Worth running before every training job you ever start: it is the cheapest possible sanity check that the model you built is the size you think it is.
- the training loop — 15 lines, unchanged since P2's bigram model. This is the pretraining loop. A real one adds checkpointing, learning-rate decay, gradient accumulation, mixed precision and multi-GPU synchronization, and none of those change what is being optimized.
Where people get stuck
- "So if I train this on more data, do I get ChatGPT?" No — you get a better document completer. More data and more parameters improve the pretraining objective, and the pretraining objective has no term for being helpful. The assistant behaviour comes from stage two, on data that looks nothing like web text. Scale and alignment are orthogonal axes; a 175 B base model with no fine-tuning is still, in the video's phrasing, undefined behaviour when you ask it a question.
- "Ours is 10 M and GPT-3 is 175 B, so it's 16,000× bigger — but he says up to a millionfold?" Two different ratios, both correct. Parameters differ by roughly 16,000×; training tokens differ by roughly 1,000,000× (300 K vs 300 B). Karpathy's "10,000 to 1 million times bigger, depending on how you count" at 115:24 is deliberately spanning both. If you want the single number that predicts capability, it is neither: it is total training compute, ≈ 6 × parameters × tokens, where the gap is the product of the two — around ten billionfold.
- "Is the reward model the same network as the language model?" Separate network, usually initialized from the same SFT checkpoint with the language-modelling head swapped for a scalar head. It reads a (prompt, response) pair and emits one number. It is trained once from the ranking data and then frozen while step 3 optimizes against it — which is exactly why step 3 can drift into reward hacking: the policy finds responses the frozen reward model loves and humans do not. Practical RLHF adds a KL penalty against the SFT model to keep the policy from wandering; DPO's use of a frozen reference model is the same guard rail in a different costume.
- "Character-level was a toy choice — does anything I learned transfer?" Everything except the tokenizer transfers unchanged. Tokenization only changes what an integer in idx stands for and how big the embedding table and the final linear layer are; vocab_size is the only line of gpt.py that would move. It matters for efficiency (BPE packs ~4× more text into the same 256-slot context) and it is the source of a family of well-known model quirks around spelling and arithmetic, but it is not part of the Transformer.
Go deeper, verified
- Language Models are Few-Shot Learners — Brown et al. (2020) · the GPT-3 paper. Table 2.1 is the table on screen at 110:14; §2.1 and §2.2 are the scaled-up version of your hyperparameter block and data loader.
- Training language models to follow instructions with human feedback — Ouyang et al. (2022) · InstructGPT: the three-step SFT → reward model → PPO recipe Karpathy sketches, written down properly, including the finding that a 1.3 B aligned model was preferred to the 175 B base model.
- Proximal Policy Optimization Algorithms — Schulman et al. (2017) · the "PPO" in step 3, from well before anyone applied it to language. Read it for what the clipped objective is actually protecting against: too large a policy update from one batch.
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafailov et al. (2023) · field map extra. The derivation that removes the reward model and the RL loop from the picture. The single biggest change to this part of the lecture.
- DeepSeekMath (GRPO) — Shao et al. (2024) · field map extra. §4 introduces GRPO: PPO with the critic replaced by group-relative normalization. The default RL algorithm for open post-training since.
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training — Lambert et al. (2024) · field map extra. Names RLVR and publishes the whole stage-two pipeline — data, code, evaluations — which is the direct refutation of "this stage is hard to replicate".
- DeepSeek-R1 — DeepSeek-AI et al. (2025) · field map extra. Large-scale RL with verifiable rewards produces long chain-of-thought reasoning; R1-Zero does it from a base model with no SFT at all.
Exercises
These are Karpathy's own EX2, EX3 and EX4 from the video description, expanded into steps. (EX1, the batched-heads challenge, belongs to P7.) All three run on top of the finished gpt.py; the first two want a GPU, the third does not.
- EX2 — train it on your own data, then teach it to addcode — Karpathy's exercise, in two escalating versions.
(a) The easy half: replace input.txt with any corpus you like — your own writing, a codebase, song lyrics — and rerun. Nothing else changes; vocab_size recomputes itself from sorted(list(set(text))). Watch the train/val gap: a small corpus will overfit long before 5,000 iterations, and you will see val loss turn back upward while train loss keeps falling.
(b) The advanced half, which is the interesting one: train a GPT to compute a+b=c. Steps: (1) replace get_batch with a generator that emits random problems as character strings — no train.bin, no held-out file, just fresh samples each step, so there is no train/val gap to speak of. (2) Write the digits of c in reverse order, because the carry propagates right-to-left and a causal model can only condition on what it has already emitted; left-to-right output would require the model to know the high digit before computing the low ones. (3) Mask the loss over the prompt positions — the a+b= part is given, not predicted — by setting those targets to -1 and passing ignore_index=-1 to F.cross_entropy. (4) Train and evaluate on held-out digit lengths. A good answer reports exact-match accuracy on unseen operands, and notes where it fails: usually longer numbers than were ever trained on. The "swole doge" extension is a full calculator for + − × ÷, which likely needs chain-of-thought traces in the training data — the model must be allowed to write intermediate steps, because a single forward pass has a fixed depth and long multiplication does not. - EX3 — pretrain on something big, then fine-tune on Shakespearecode — Karpathy's exercise, and a hands-on version of this part's whole argument. Steps: (1) find a corpus large enough that train and val loss stay glued together for the whole run — a sample of FineWeb, all of Project Gutenberg, a large code dump. No visible gap means you are not memorizing and every step is buying real generalization. (2) Pretrain gpt.py on it, keeping the character vocabulary — but build the vocabulary from the union of both corpora so the fine-tuning stage does not meet unknown characters. (3) Save the weights with torch.save(model.state_dict(), ...). (4) Fine-tune: load that checkpoint, swap in tiny Shakespeare, drop the learning rate by roughly 10× (3e-5 rather than 3e-4) and run far fewer iterations. (5) Compare final validation loss against the from-scratch run's 1.48. A good answer reports both curves on the same axes and says how many fine-tuning steps the pretrained model needed to pass the from-scratch model's best loss — that number, not the final loss, is the measure of what pretraining bought.
- EX4 — implement one thing from the paperscode — Karpathy's exercise: read some Transformer papers, add one feature people actually use, and measure whether it helps. Steps: (1) pick one change and only one — candidates the earlier parts flag are rotary position embeddings (P3), FlashAttention or PyTorch's scaled_dot_product_attention (P4), grouped-query attention, RMSNorm or SwiGLU (P6), and weight tying between the embedding table and lm_head (P7). (2) Before touching anything, fix a baseline: same seed, same max_iters, record val loss at every eval_interval, and note wall-clock time and peak torch.cuda.max_memory_allocated(). (3) Make the change. (4) Rerun identically and plot both curves. A good answer is honest about which axis moved: several of these (FlashAttention, GQA) are not meant to improve loss at all — they buy speed or memory at equal loss, and reporting "no improvement" for them is the correct result. State the comparison you are making before you run it, or you will find a way to call any outcome a win.