Overview and tokenization
Transcript: cleaned auto-captions with timestamps
This is the load-bearing lecture for the whole course: it sets up the frame ("efficiency given fixed resources") that every later lecture is an instance of, previews all five assignments so you know what you are signing up for, and then does one full unit of real content — tokenization — end to end, from Unicode code points to a working BPE trainer. Everything after this is either a component of the pipeline sketched here (architecture, kernels, parallelism, data, alignment) or a way to spend your compute budget better.
Outline, with timestamps
- 00:05 — Staff, ethos, and the case for building from scratch: researchers have drifted above the stack, and the abstraction is leaky.
- 04:02 — Industrialization: frontier scale is out of reach, and small models can mislead you.
- 06:47 — Mechanics, mindset, intuitions — and the bitter lesson restated as accuracy = efficiency × resources.
- 12:24 — A compressed history, Shannon to DeepSeek, and the three levels of openness.
- 18:07 — What an executable lecture is; then course logistics, the workload warning, and the cluster.
- 26:39 — Unit 1 · Basics: tokenizer, Transformer variants, optimizer, training loop.
- 33:23 — Unit 2 · Systems: kernels and the memory hierarchy, multi-GPU parallelism, inference.
- 40:45 — Unit 3 · Scaling laws: trade model size against tokens, and the D* ≈ 20N* rule of thumb.
- 44:39 — Unit 4 · Data: evaluation, then curation — the data does not fall from the sky.
- 50:14 — Unit 5 · Alignment: SFT first, then learning from preferences and verifiers.
- 55:55 — Efficiency as the through-line, and what changes when you stop being compute-constrained.
- 60:23 — Tokenization: what the interface actually is, and what a compression ratio buys you.
- 65:23 — Characters, bytes, words: three attempts that fail, each for a different reason.
- 71:05 — BPE: train the vocabulary from corpus statistics, merge by merge, and what Assignment 1 adds.
The argument for building it yourself
Percy opens with a diagnosis rather than a syllabus (02:19). Eight years ago an AI researcher implemented and trained their own model. Six years ago they downloaded BERT and fine-tuned it. Today a large fraction of published work never leaves the prompt. That progression is not a moral failure — abstraction is how any field scales its productivity, and Percy is explicit that he prompts models too. The problem is the kind of abstraction. A programming language or an operating system gives you a contract you can reason about and, when it leaks, a stack you can descend into. A frontier model gives you a string in and a string out. There is no descent.
The consequence is that a whole class of research questions — the ones that require co-designing the data, the systems layer and the model together — becomes inaccessible to anyone who only has the string interface. That is the gap the course is built to close.
"Our philosophy is: to understand it, you have to build it."— Percy Liang, 03:27
Then the obvious objection, which he raises himself (04:02): you cannot build what the frontier builds. GPT-4 is reported at ~1.8T parameters and ~$100M to train; xAI has talked about a 200,000-H100 cluster; the Stargate announcement is $500B over four years. And the technical reports have stopped telling you anything — the GPT-4 report says outright that competitive and safety considerations mean no architecture, dataset, or training-compute details. So the class trains sub-1B models, and the honest question is whether anything learned there is even true at scale.
Two counterexamples make the risk concrete. The first is cheap: the fraction of FLOPs spent in attention versus the MLP shifts dramatically with model size — roughly comparable at small scale, MLP-dominated by 175B. Optimize attention at 100M parameters and you may be tuning something that gets washed out three orders of magnitude later. Percy's point is that this one you can catch with napkin math, no compute required. The second is harder: emergence. Plot accuracy against training FLOPs across many tasks and you see long flat stretches followed by sharp turn-ons. Sit below the threshold and you would conclude, wrongly, that the capability does not exist.
What actually transfers: mechanics, mindset, intuitions
The resolution (06:47) is the most useful piece of scaffolding in the lecture, and it is worth carrying into any "does this small-scale result hold?" argument you have later. Three categories of knowledge:
| Mechanics | How things work: what a Transformer is, how model parallelism maps onto GPUs, what a fused kernel does. Fully teachable. Scale-invariant. |
| Mindset | Squeeze the hardware; take scaling seriously; measure before you optimize. Teachable, and Percy argues this is the piece OpenAI actually contributed — the ingredients were mostly lying around by 2019. |
| Intuitions | Which architectural and data choices produce good models. Only partly teachable — this is exactly the category that can flip sign across scales. |
Two and a half out of three, as he puts it. And the class is candid that even the intuitions we have are often post-hoc. The example is Noam Shazeer's GLU-variants paper, which introduced SwiGLU and whose conclusion offers "divine benevolence" in place of an explanation for why it works (08:59). SwiGLU is now near-universal anyway. That is the epistemic state of the field, stated plainly rather than papered over.
This leads into a restatement of the bitter lesson (09:33) that is worth internalizing because the popular reading is wrong. The popular reading is "scale is all that matters, algorithms don't." Percy's reading is that algorithms that scale are what matter, and he formalizes it as:
accuracy = efficiency × resources
Efficiency is not the consolation prize for the GPU-poor; it matters more as budgets grow, because at $100M a run you cannot afford to relaunch after a crash the way you can on a lab cluster. And the empirical support is direct: OpenAI's algorithmic-efficiency study found a 44× reduction in the compute needed to hit AlexNet-level ImageNet accuracy between 2012 and 2019 — faster than Moore's law over the same window. Without that 44×, the same result costs 44× more. Algorithms plainly matter.
The framing that falls out — what is the best model you can build given a fixed compute and data budget? — is scale-free. It is a sensible question with 8 GPUs and a sensible question with 200,000, which is precisely why it can organize a course taught at the former for people who may one day work at the latter.
The five units, and why efficiency links them
The middle of the lecture is a preview of the course, framed by a single worked question: given a Common Crawl dump and 32 H100s for two weeks, what should you do? (26:39)
Basics gets the whole pipeline standing: tokenizer, architecture, training. The architecture section is a useful map of what has actually changed since 2017 — activation (ReLU → SwiGLU), positional encoding (sinusoidal → RoPE), normalization (LayerNorm → RMSNorm, and pre-norm rather than post-norm), MLP (dense → mixture of experts), attention (full → sliding-window, linear, GQA, MLA), plus the non-Transformer alternatives such as Hyena and the hybrids that mix them in. Each individually is a small delta; Percy's claim is that summed they are worth an order of magnitude against a vanilla Transformer.
Systems (33:23) is where the abstraction gets torn open. The mental model is one analogy: DRAM is a warehouse, SRAM is the factory, and the bottleneck is trucking. Compute is free relative to data movement, so kernel work — fusion, tiling, written in Triton for this course — is about restructuring computation so bytes cross the gap less often. Scale that up to 8 GPUs and the same principle holds with a slower interconnect, which is what data / tensor / pipeline / sequence parallelism are all negotiating. Inference gets its own treatment (new this year), split into prefill, which is compute-bound and looks like training, and decode, which is autoregressive, memory-bound, and hard to saturate a GPU with. The speedups follow from that split: cheaper models, KV caching and batching, and speculative decoding — a draft model scouts ahead and the full model scores its guesses in parallel, which is exact, not approximate.
Scaling laws (40:45) answers the size-versus-tokens trade-off. For a fixed FLOPs budget you can sweep parameter count, measure loss, and take the minimum of each isoFLOP curve; plot those minima and they line up remarkably straight, which is what makes extrapolation possible. The Chinchilla rule of thumb is D* ≈ 20 N* — a 1.4B model wants about 28B tokens. Percy immediately flags the limitation: it optimizes training cost only and ignores inference entirely, which is why production models are routinely trained well past the Chinchilla point.
Data (44:39) is where he is most insistent, and the live demo is the argument: he samples ten random Common Crawl documents on stage and most of them are spam or boilerplate.
"Data doesn't just fall from the sky."— Percy Liang, 46:53
Scraped data is not text — it is HTML, PDFs, and directory trees, and the HTML-to-text step is lossy in ways that determine what your model can learn. Then filtering (usually by trained classifiers, for both quality and harm) and deduplication (MinHash, Bloom filters). Evaluation comes first in this unit rather than last, which is the right order: perplexity, standardized benchmarks like MMLU, instruction-following evals, LM-as-judge, and whole-system evaluation for RAG and agents.
Alignment (50:14) turns the base model — raw next-token potential — into something usable, along three axes: instruction-following, style, and safety/refusal. Phase one is supervised fine-tuning on (prompt, response) pairs, and it is mechanically identical to pre-training; the LIMA result is the reason it works with so little data, since the capability is already latent and you are mostly surfacing it. Phase two is learning from cheaper signal: preference data (A vs B), or verifiers, which can be formal for math and code or learned. Algorithmically that is PPO originally, DPO when your data is purely preferences and you want to skip the RL machinery, and GRPO — which drops the value function from PPO — when you need to learn from verifier rewards.
Tokenization: the interface, and the number that matters
The second half (60:23) is a full unit of real content, and it is structured as a derivation rather than a description — three attempts fail, and BPE is what is left standing.
The interface is small. A tokenizer is a pair of functions, encode: str → list[int] and decode: list[int] → str, and it must round-trip. Vocabulary size is the number of distinct integers. That is the whole contract, and it is worth noting how much it differs from classical NLP tokenization: nothing is allowed to disappear. Whitespace is carried inside tokens, which is why in GPT-2's vocabulary "hello" and " hello" are different tokens — a fact that quietly causes real bugs when you concatenate prompts. A student asks whether the leading-space convention is deliberate or an artifact; Percy's answer is that it falls out of the pre-tokenizer regex and could have gone the other way (63:11). Numbers, similarly, get chopped left-to-right in chunks that have nothing to do with place value.
The number that carries the rest of the lecture is the compression ratio: bytes per token.
compression_ratio = len(utf8_bytes(s)) / len(encode(s))
It is the exchange rate between text and sequence length, and because attention is quadratic in sequence length, it is directly an efficiency number. Running the lecture's demo string "Hello, 🌍! 你好!" through GPT-2's tokenizer gives 20 bytes over 12 tokens ≈ 1.7 — deliberately unflattering, since the string is emoji- and CJK-heavy. On ordinary English prose GPT-2 lands closer to 4–5 bytes per token, and that gap is itself the reason multilingual tokenizer design is a live problem: the same sentence costs several times more context in some languages than others.
Three tokenizers that don't work, each for a different reason
The derivation (65:23) is clean because each failure mode is distinct, and each one motivates a specific property of BPE.
| Characters | One Unicode code point per token (ord/chr). Vocabulary ~150K, and you burn one slot on 🌍 (code point 127757) exactly as on "e". Compression ratio ≈ 1.5 on the demo string. Fails on: vocabulary allocated uniformly to wildly non-uniform frequencies. |
| Bytes | UTF-8 bytes, so the vocabulary is exactly 256 and every byte gets used. Elegant. Fails on: compression ratio is exactly 1.0 by construction — sequences are as long as the text is, and quadratic attention makes that ruinous. |
| Words | Split on a regex, map each segment to an integer. Adaptive in the right way — common words become one token. Fails on: the vocabulary is unbounded in principle and unknown in practice; unseen segments need an UNK token, which is ugly and silently corrupts perplexity comparisons. |
Read together, the three failures are a specification. You want the byte tokenizer's closed, always-round-tripping vocabulary; the word tokenizer's adaptivity, where frequent strings get short codes; and neither one's failure mode. That is exactly what BPE delivers, and it is why the derivation is more instructive than simply being handed the algorithm.
BPE: let the corpus choose the vocabulary
Byte-pair encoding (71:05) is a 1994 data-compression algorithm by Philip Gage, brought into NLP by Sennrich, Haddow and Birch in 2016 for machine translation — where it solved the rare-word problem that had forced everyone into word-level vocabularies — and then adopted by GPT-2, which is how it became the default for language modeling.
The insight is the one word train. Instead of choosing a segmentation rule up front, you learn the vocabulary from corpus statistics: start from bytes, and repeatedly merge whichever adjacent pair occurs most often, minting a new token for it. Frequent strings organically collapse into single tokens; rare strings stay decomposed. The training loop is about fifteen lines:
indices = list(s.encode("utf-8")); vocab = {i: bytes([i]) for i in range(256)}
repeat num_merges times:
count every adjacent pair in indices
pair = argmax(counts); new_index = 256 + step
merges[pair] = new_index; vocab[new_index] = vocab[a] + vocab[b]
indices = replace every occurrence of pair with new_index
Percy runs it live on "the cat in the hat" for three merges (72:44). Merge 1 takes (116, 104) — "t","h" — to token 256, which appears twice. Merge 2 takes (256, 101) to 257, giving you "the". Merge 3 takes (257, 32) to 258, folding in the trailing space. Watch the sequence shrink at each step: that shrinkage is the compression ratio improving, bought with vocabulary slots.
Encoding is the mirror image — convert to bytes, then replay the stored merges in training order. The ordering is not incidental; merges are a dependency chain, since token 257 cannot form before 256 exists. Decoding is a dictionary lookup and a concatenation.
Two production details are flagged rather than implemented. First, GPT-2 does not run BPE over raw text — it pre-tokenizes with a regex ('(?:[sdmt]|ll|ve|re)| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S)|\s+) and runs BPE inside each segment, which both bounds the work and stops merges from crossing word boundaries. That regex is also where the leading-space convention comes from. Second, special tokens like <|endoftext|> must be detected and preserved rather than merged through.
And the lecture's own implementation is deliberately naive — encode loops over every merge in the table for every input, which is the wrong complexity by a wide margin. Making it fast is your job (77:15), and Percy warns twice that this is where students lose the most time on Assignment 1.
The honest coda is that none of this is something anyone is proud of. Tokenization is a compute-efficiency workaround, not a modeling idea — it is why models can't spell, can't do digit arithmetic reliably, and cost different amounts per language. The tokenizer-free line of work (ByT5, MegaByte, BLT, T-FREE) tries to remove it, and Percy's assessment is that it is promising but has not been scaled to the frontier.
"I hope that one day I won't have to give this lecture, because we'll just have architectures that map from bytes — but until then, we'll have to deal with tokenization."— Percy Liang, 78:24
Where the lecturer hedges
Worth marking, because these are the places the field genuinely does not know:
- Intuitions may not transfer. Stated up front and repeated. Any small-scale finding in this course is provisional at frontier scale.
- SwiGLU has no explanation. "Divine benevolence" is the actual text of the paper's conclusion, and Percy quotes it as representative rather than exceptional.
- Chinchilla ignores inference. He raises this himself immediately after presenting the rule.
- Leading-space is arbitrary. Asked directly, he says it could have gone either way.
- Tokenization shouldn't exist. Ends the lecture wishing it away.
What you build with this
This lecture starts Assignment 1: Basics — the largest of the five, and the one Percy warns about with a course evaluation quote: the whole assignment is about as much work as all five CS 224N assignments plus the final project. You get unit tests and adapter interfaces, but no scaffolding — a blank file and a spec.
The tokenization half of this lecture maps directly onto the first deliverable: implement a BPE tokenizer with training, encoding, decoding, pre-tokenization via the GPT-2 regex, and special-token handling. The naive train_bpe and BPETokenizer.encode from lecture_01.py are correct and slow; the assignment is largely about making them fast (the merge loop, and parallelizing the counting pass). The rest of the assignment comes from Lecture 2's material — Transformer, cross-entropy loss, AdamW, training loop — trained on TinyStories and OpenWebText, with a leaderboard for minimum OpenWebText perplexity in 90 minutes on one H100.
- Assignment 1: Basics — repo — starter tests and adapters.
- Assignment 1 handout (PDF) — the spec, problem by problem, with point values.
- Assignment 1 leaderboard — 90 minutes, one H100, minimize OpenWebText perplexity.
Supporting materials, verified
- A New Algorithm for Data Compression — Philip Gage (1994) · The original byte-pair encoding, from The C Users Journal. Pure compression, no linguistics. (The lecture links the original pennelynn.com copy, which is now dead; this is the Internet Archive snapshot.)
- Neural Machine Translation of Rare Words with Subword Units — Sennrich, Haddow, Birch (2016) · Brought BPE into NLP and killed the UNK token. The paper Percy credits for the whole subword era.
- Language Models are Unsupervised Multitask Learners (GPT-2) — Radford et al. (2019) · Where the pre-tokenize-then-BPE recipe and the GPT-2 regex come from; the tokenizer you will benchmark against.
- tiktoken — OpenAI · The fast BPE library used live in the lecture; get_encoding("gpt2") is the reference implementation to check your round-trips against.
- tiktokenizer — interactive · The site Percy types into at 61:30. Fastest way to build intuition about spaces, numbers and multilingual cost.
- Let's build the GPT Tokenizer — Andrej Karpathy (2024) · Explicitly named as the inspiration for this unit; two hours of the same derivation at a slower pace.
- Attention Is All You Need — Vaswani et al. (2017) · The architecture baseline everything in Unit 1 is a delta against.
- GLU Variants Improve Transformer — Noam Shazeer (2020) · Introduced SwiGLU; source of the "divine benevolence" conclusion Percy uses to characterize the field's explanatory state.
- RoFormer: Enhanced Transformer with Rotary Position Embedding — Su et al. (2021) · RoPE, the positional encoding that replaced sinusoidal.
- Root Mean Square Layer Normalization — Zhang & Sennrich (2019) · RMSNorm: LayerNorm minus the mean-centering, cheaper and no worse.
- Decoupled Weight Decay Regularization — Loshchilov & Hutter (2017) · AdamW, the optimizer you implement in Assignment 1.
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Shazeer et al. (2017) · The MoE ancestor; Lecture 4 is entirely about its descendants.
- GQA: Training Generalized Multi-Query Transformer Models — Ainslie et al. (2023) · Grouped-query attention, the standard KV-cache shrink.
- DeepSeek-V2 — DeepSeek (2024) · Source of multi-head latent attention (MLA), named in the architecture preview.
- Hyena Hierarchy: Towards Larger Convolutional Language Models — Poli et al. (2023) · The non-attention alternative Percy cites when sketching hybrids.
- Scaling Laws for Neural Language Models — Kaplan et al. (2020) · The OpenAI half of the compute-optimal story.
- Training Compute-Optimal Large Language Models (Chinchilla) — Hoffmann et al. (2022) · The isoFLOP plot shown on screen and the source of D* ≈ 20N*.
- Measuring the Algorithmic Efficiency of Neural Networks — Hernandez & Brown, OpenAI (2020) · The 44× ImageNet result behind "algorithms matter."
- Emergent Abilities of Large Language Models — Wei et al. (2022) · The emergence plot; the strongest argument that small-scale conclusions can be flatly wrong.
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling — Gao et al., EleutherAI (2021) · The data-mixture pie chart used in the Data preview.
- LIMA: Less Is More for Alignment — Zhou et al. (2023) · ~1000 examples suffice for instruction-following; the evidence that SFT surfaces rather than teaches.
- Direct Preference Optimization — Rafailov et al. (2023) · DPO. (Percy says "Direct Policy Optimization" on the slide; the paper's name is Direct Preference Optimization.)
- DeepSeekMath — Shao et al. (2024) · Introduced GRPO, which drops PPO's value function. (Slide says "Group Relative Preference Optimization"; the paper's term is Group Relative Policy Optimization.)
- Proximal Policy Optimization Algorithms — Schulman et al. (2017) · PPO, the RLHF baseline both DPO and GRPO simplify away from.
- ByT5 · MegaByte · Byte Latent Transformer · T-FREE — the four tokenizer-free papers Percy links when saying he wishes this lecture were unnecessary.
- SentencePiece — Kudo & Richardson (2018) · field map extra. The other production tokenizer path (Llama, T5, Gemma). Worth reading alongside BPE because it treats the input as a raw byte stream with no pre-tokenization at all.
- Subword Regularization — Kudo (2018) · field map extra. The unigram-LM alternative to BPE's greedy merges — a probabilistic segmentation model rather than a heuristic, and the answer to "is greedy merging actually optimal?"
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in LLMs — Land & Bartolo (2024) · field map extra. What goes wrong when the tokenizer corpus and the training corpus disagree: glitch tokens that exist in the vocabulary but were never trained. The concrete failure mode behind Percy's discomfort.
- How Good is Your Tokenizer? — Rust et al. (2021) · field map extra. Measures the downstream cost of tokenizer quality across languages — the empirical case that the compression ratio is not a cosmetic number.
Exercises
- Measure the exchange rate code — Quantify what tokenization actually buys. (1) pip install tiktoken and load the gpt2 encoding. (2) Assemble five ~2KB samples: English prose, Python source, JSON, Chinese, and a line of digits. (3) For each, compute bytes/token for GPT-2 and for a pure byte tokenizer. (4) Repeat with the cl100k_base encoding. (5) Convert each ratio into an attention-cost multiplier — cost goes as the square of sequence length. A good answer names which sample is worst, states the ratio between best and worst, and explains what that means for a fixed context window and for the per-token price a user in that language pays.
- Write train_bpe, then make it fast code — (1) Reimplement the lecture's trainer from scratch: bytes in, (vocab, merges) out, and assert round-trip on held-out text. (2) Time training a 1000-merge vocabulary on ~10MB of text; the naive version recounts all pairs and rewrites the whole sequence per merge, so it will be slow. (3) Fix the counting: maintain an incremental pair-count table and update only the neighborhoods touched by each merge. (4) Fix encode: the lecture version loops over every merge in the table — index merges by pair so you only consider ones actually present. (5) Add GPT-2 regex pre-tokenization and parallelize across segments. A good answer reports before/after wall-clock, states the complexity change, and proves the fast version produces byte-identical output to the slow one.
- Break the round-trip code — BPE tokenizers are supposed to round-trip, but the naive decoder in the lecture concatenates bytes and calls .decode("utf-8"). Construct an index sequence that makes it raise UnicodeDecodeError — a merge boundary that splits a multi-byte character — then fix the decoder so partial sequences degrade gracefully. A good answer explains why this matters for streaming generation, where you must render tokens before the character they belong to is complete.
- Argue the regime — Percy claims that today's design decisions reflect being compute-constrained, and that data-constrained labs will make different ones. Pick three decisions from the lecture — single-epoch training, aggressive data filtering, and tokenization itself — and write a paragraph each on how you would expect the choice to change under a fixed 10T-token corpus and effectively unlimited compute. A good answer identifies for each decision the specific resource being economized, and names an observable that would tell you the field has flipped regimes. (No code required.)