The setup: the problem, the data, tokens, batches
Transcript: this part, with timestamps
Nothing in this part is a Transformer. That is deliberate, and it is the reason the rest of the lecture goes as fast as it does. Karpathy spends 22 minutes making the problem so concrete that the model becomes the only remaining unknown: by the end there is a tensor of integers, a function that hands you a random batch of windows out of it, and a precise statement of what a correct answer looks like at every position in that batch. Everything from 22:11 onward — the bigram lookup table, the four versions of attention, the full gpt.py — is a different function plugged into the same socket. The socket is what gets built here.
Outline, with timestamps
- 00:00 — Intro: ChatGPT is a probabilistic sequence completer; same prompt, different completion.
- 02:05 — The architecture under it: Attention Is All You Need (2017), and what the "T" in GPT stands for.
- 03:36 — Why tiny Shakespeare, and nanoGPT as the finished artifact this lecture rewrites from an empty file.
- 07:52 — Reading the data: 1.1 MB, one flat string, no document boundaries.
- 09:28 — sorted(list(set(text))) gives 65 symbols; tokenization is then two dicts, encode and decode.
- 10:51 — SentencePiece, tiktoken, and the codebook-size / sequence-length trade.
- 13:27 — The 90/10 train/val split, and what it is meant to detect.
- 14:27 — block_size: you never feed the whole corpus, you feed chunks.
- 15:31 — Nine characters are eight examples; the loop that spells them out.
- 17:02 — Why train on every context length from 1 to block_size — it is not only efficiency.
- 19:04 — The batch dimension: get_batch, random offsets, torch.stack, a (4, 8) tensor.
- 21:06 — 32 independent examples in one batch, spelled out prefix by prefix.
The problem, stated so it can be coded
The opening five minutes are a framing argument, and it is worth extracting because the rest of the lecture leans on it. ChatGPT completes sequences. Give it a prompt twice and you get two different answers, which tells you it is not looking anything up — it is sampling from a distribution over what comes next. The architecture doing that came out of a 2017 machine-translation paper, and the "GPT" acronym unpacks as generatively pretrained Transformer: the Transformer is the network, "generatively pretrained" is what you did to it.
Karpathy then makes the concession that defines the whole lecture: a production system is trained on a good chunk of the internet and then put through further stages, and none of that is reproducible in an afternoon. So he keeps the architecture and shrinks everything else. The corpus becomes one megabyte of Shakespeare. The alphabet becomes literal characters instead of subword tokens. What survives the shrinking is exactly the part he wants you to understand — the network — and the claim, which the rest of the lecture makes good on, is that nothing structural changes. The same code that babbles fake Shakespeare at 10 million parameters is the code that writes marketing copy at 175 billion. Only the numbers move.
Two places to keep your hedge-detector on. He dates GPT-2 to "2017 if I recall correctly" — GPT-2 is February 2019 (2017 is the Transformer paper, and GPT-1 is mid-2018); the point he is making, that nanoGPT reproduces the 124 M-parameter GPT-2 and can load OpenAI's released weights, is correct. And he describes the 2017 paper as reading "like a pretty random machine translation paper." That is fair as a reaction, but it is worth knowing why: the paper's pitch was parallelism and training cost for translation, and what this lecture builds is only the decoder half of the architecture it proposes. That asymmetry gets named explicitly much later, in P5's note 4 and again in P7.
The data: one flat string, 65 distinct symbols
The corpus is tiny Shakespeare, a file Karpathy has been using since char-rnn in 2015. Read it and count and you get 1,115,394 characters — about 1.1 MB, which at UTF-8 with a pure-ASCII corpus is also 1,115,394 bytes. It is a concatenation of plays with no metadata, no separators, no document boundaries. The only structure in it is typographic:
First Citizen:
Before we proceed any further, hear me speak.
All:
Speak, speak.
Speaker name, colon, newline, line of verse, blank line. Remember that shape — it is the first thing any model learns here, and in P2 you will watch a model that cannot see past one character reproduce the colon-and-newline rhythm while producing no English at all.
The vocabulary comes from three stacked calls: set(text) collapses to distinct characters, list(...) gives them an order, sorted(...) makes that order deterministic. The sort is the only part that matters and it is easy to skip past. Python's set iteration order is not something you want a trained checkpoint to depend on; sorting means the integer 46 means h today and h on any other machine. The result:
['\n', ' ', '!', '$', '&', "'", ',', '-', '.', '3', ':', ';', '?',
'A'…'Z', 'a'…'z'] # 65 symbols
Thirteen punctuation-and-whitespace symbols, then the two alphabets. Two details are worth a second: the only digit in all of Shakespeare-as-distributed is 3, and there is no uppercase/lowercase merging — A and a are unrelated integers as far as the model is concerned, and it has to learn their relationship from data like everything else. vocab_size = 65 is then wired into two places for the rest of the lecture: the number of rows in the embedding table, and the number of columns in the final logits.
Tokenization is a dial, not a design
The tokenizer here is two dictionaries and two one-liners: stoi maps character to integer, itos maps back, encode is a list comprehension and decode is a join. Encoding "hi there" gives [46, 47, 1, 58, 46, 43, 56, 43] — eight integers for eight characters, and decode(encode(s)) == s for any string over the vocabulary.
Karpathy then does something more useful than explaining his own tokenizer: he shows you what you are giving up. Google's SentencePiece and OpenAI's tiktoken are subword schemes — units bigger than a character, smaller than a word. Run the same "hi there" through GPT-2's tiktoken encoding and you get three integers instead of eight, drawn from a vocabulary of 50,257 instead of 65.
That is the whole trade, and it is a genuine dial with a cost on both ends:
| Scheme | Vocabulary | "hi there" | What it costs |
|---|---|---|---|
| character (this lecture) | 65 | 8 integers | long sequences; attention is O(T²), so context is expensive |
| subword (GPT-2 BPE) | 50,257 | 3 integers | a 50k-row embedding table and a 50k-wide output layer |
Small codebook, long sequences; large codebook, short sequences. Practice sits in the middle because both ends are quadratic in something you care about. Character level is the right choice for a lecture not because it is better but because it deletes an entire subsystem — no merge table, no training a tokenizer, no byte-fallback edge cases — while leaving every downstream line of code identical. Karpathy is explicit that this is a simplification, and he later gave the skipped subsystem its own two hours; that video is linked below.
Splitting, and what a validation loss here actually measures
The encoded corpus goes into a single 1-D torch.Tensor of dtype long — one integer per character, 1,115,394 of them — and then gets cut 90/10:
| Tensor | Length | Slice |
|---|---|---|
| data | 1,115,394 | the whole corpus |
| train_data | 1,003,854 | data[:n], first 90% |
| val_data | 111,540 | data[n:], last 10% |
The stated motivation is overfitting detection: we do not want a network that has memorised this exact text, we want one that produces text like it, so hold some back and watch whether the two losses separate. That is right, and in P7 the gap between them becomes the main diagnostic for whether the model is too big for the data.
What goes unsaid is worth knowing. The split is contiguous, not random. Contiguous is the correct choice — a randomly-held-out character sitting between two training characters is trivially predictable, so a random split would make the validation loss meaningless. But contiguity has a price: the last 10% of the file is a different stretch of plays, with different speakers and different vocabulary density. So the train/val gap you measure is partly memorisation and partly distribution shift between one end of the corpus and the other. At the scale of this lecture that is fine. At the scale where you would care about the difference, you would hold out whole documents.
Chunks, and why nine characters are eight examples
You never feed a corpus to a Transformer. You feed fixed-length windows, and the length is block_size — Karpathy's name for what other codebases call context length. Here it starts at 8.
The first thing he prints is train_data[:block_size+1], nine integers: [18, 47, 56, 57, 58, 1, 15, 47, 58], which decodes to "First Cit". The +1 is the crux of the data format. Inputs are the first eight, targets are the same window shifted one step left — so the target for position t is the character that actually followed the prefix ending at t. A nine-integer window therefore contains eight (context, target) pairs, with contexts of every length from 1 to 8:
x = train_data[:block_size] # 18 47 56 57 58 1 15 47
y = train_data[1:block_size+1] # 47 56 57 58 1 15 47 58
for t in range(block_size):
context = x[:t+1]
target = y[t]
print(f"when input is {context.tolist()} the target: {target}")
Reading it out: given [18] predict 47; given [18, 47] predict 56; given [18, 47, 56] predict 57; and so on to the full eight-character prefix. One window, eight supervised examples.
Karpathy is careful to say this is not merely a way to squeeze value out of data you already loaded:
We train on all the eight examples here with context between one all the way up to context of block_size… it's also done to make the Transformer network be used to seeing contexts all the way from as little as one all the way to block_size. Andrej Karpathy, 17:02 (cleaned auto-captions)
This is a generation requirement, not a training convenience. In P2 sampling starts from a single token and grows the context one character at a time. If the network had only ever seen full-length contexts, its first seven predictions would be out of distribution every time you sampled. Training on all prefix lengths is what makes the model well-defined on a one-character prompt.
There is a second reason he leaves implicit, and it is the one that makes Transformer language modelling economically different from most supervised learning: the prediction at every position is computed anyway. Because of the causal mask you will meet in P3, position t's output depends only on positions ≤ t, so a single forward pass over an eight-token window legitimately produces eight independent predictions from eight different amounts of history. You get T training signals for the price of one pass. Nothing is being reused improperly; the mask is what makes it honest.
The batch dimension, and the shape contract
The second dimension exists purely to keep the hardware busy. GPUs are wide; a single 8-integer sequence uses almost none of that width, so we stack several unrelated windows into one tensor and let them ride through the network together. They never interact — a fact that gets a whole numbered note in P5 — they just share a kernel launch.
get_batch does three things. It picks which split to read from; it draws batch_size random offsets into that split with torch.randint; and it torch.stacks the resulting 1-D windows into rows of a 2-D tensor. With the notebook's batch_size = 4 and block_size = 8, and torch.manual_seed(1337) so the numbers are reproducible, you get:
| Name | Shape | Contents |
|---|---|---|
| ix | (4,) | random start offsets into train_data |
| xb | (4, 8) = (B, T) | four independent 8-character windows, as integers |
| yb | (4, 8) = (B, T) | the same four windows shifted one step left |
The first row begins 24, 43, 58, 5 — decode it and you get "Let'". Karpathy walks it prefix by prefix: input 24 → target 43; input 24, 43 → target 58; input 24, 43, 58 → target 5. Four rows times eight positions is 32 supervised examples in one (4, 8) tensor, and that count is the single most useful thing to carry out of this part. The loss you will see in P2 is a mean over 32 numbers, not 4.
Those chunks are processed completely independently, they don't talk to each other. Andrej Karpathy, 18:33 (cleaned auto-captions)
The code at the end of this part
By 22:11 the notebook holds everything below. This is the excerpt of bigram.py that this part produces — the file it eventually gets written into at 38:00, in P2. The corpus, the vocabulary, the tokenizer, the split:
# wget https://raw.githubusercontent.com/karpathy/char-rnn/master/data/tinyshakespeare/input.txt
with open('input.txt', 'r', encoding='utf-8') as f:
text = f.read()
# here are all the unique characters that occur in this text
chars = sorted(list(set(text)))
vocab_size = len(chars)
# create a mapping from characters to integers
stoi = { ch:i for i,ch in enumerate(chars) }
itos = { i:ch for i,ch in enumerate(chars) }
encode = lambda s: [stoi[c] for c in s] # encoder: take a string, output a list of integers
decode = lambda l: ''.join([itos[i] for i in l]) # decoder: take a list of integers, output a string
# Train and test splits
data = torch.tensor(encode(text), dtype=torch.long)
n = int(0.9*len(data)) # first 90% will be train, rest val
train_data = data[:n]
val_data = data[n:]
Line by line (L17–L34): the encoding='utf-8' matters only because Python would otherwise guess from the locale. chars is the 65 symbols in sorted order, and vocab_size is derived from the data rather than hard-coded — swap the corpus and everything downstream re-sizes itself. stoi and itos are exact inverses by construction because both come from the same enumerate. torch.long is not optional: these integers are used as indices into an embedding table and as class labels for cross-entropy, and both APIs require 64-bit integers. And int(0.9*len(data)) is a plain truncation — n = 1,003,854.
Then the loader, which is the actual deliverable of this part (L36–L44):
batch_size = 32 # how many independent sequences will we process in parallel?
block_size = 8 # what is the maximum context length for predictions?
# data loading
def get_batch(split):
# generate a small batch of data of inputs x and targets y
data = train_data if split == 'train' else val_data
ix = torch.randint(len(data) - block_size, (batch_size,))
x = torch.stack([data[i:i+block_size] for i in ix])
y = torch.stack([data[i+1:i+block_size+1] for i in ix])
x, y = x.to(device), y.to(device)
return x, y
Five lines of body, each doing one thing. data = train_data if … deliberately shadows the module-level data inside the function, which is tidy but will bite you if you ever want the full corpus in here. torch.randint(hi, (batch_size,)) draws batch_size integers uniformly from [0, hi) — sampling with replacement, so a batch can in principle contain the same window twice, and there is no notion of an epoch anywhere in this codebase. The two list comprehensions build the input and target windows, offset by exactly one. torch.stack creates a new leading dimension, turning batch_size tensors of shape (8,) into one of shape (batch_size, 8) — that new axis is B. And the .to(device) happens here, in the loader, so the model never has to think about where its inputs live.
Two numbers to note while you are here. The script version uses batch_size = 32 where the notebook demo used 4 — the (4, 8) tensor you were just staring at is (32, 8) once the code is a script. And with max_iters = 3000 in P2, training touches 3000 × 32 × 8 = 768,000 character positions out of 1,003,854 available — less than a single pass over the training split. At bigram scale that is plenty.
Where people get stuck
- "Why is the window block_size + 1 long?" Because the targets are the inputs shifted one step. To supply a target for the eighth input position you need a ninth character. In get_batch this shows up as y = data[i+1 : i+block_size+1] — same length as x, starting one later. Both tensors are (B, T); only the alignment differs.
- "len(data) - block_size looks like an off-by-one bug." It isn't, and it has zero slack — walk it. torch.randint's bound is exclusive, so the largest offset drawn is len(data) - block_size - 1. The last element y touches is at index i + block_size = len(data) - 1, the final valid index. One character further and the slice would silently come back short — Python slicing does not raise on over-run — and you would get a confusing torch.stack size-mismatch rather than an index error.
- "Is a (4, 8) batch four examples or eight?" Neither: it is 32. Every one of the B × T positions is a separate (context, target) pair with its own loss term, and the reported loss is their mean. This is why the batch and block dimensions both show up in the cost of a step, and why P2 has to flatten (B, T, C) logits to (B*T, C) before handing them to F.cross_entropy.
- "Everyone else says tokens; this says characters. Am I learning the wrong thing?" The model code is byte-for-byte identical either way. A tokenizer is a bijection between strings and integer sequences; the network only ever sees integers in [0, vocab_size). Swapping in GPT-2's BPE changes exactly one number — vocab_size from 65 to 50,257 — plus the physical size of two matrices. Nothing about attention, residuals or LayerNorm knows or cares.
Go deeper, verified
- Attention Is All You Need — Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin (2017) · The architecture the whole lecture reconstructs. Read §3.1–§3.2 now for the shapes and come back to §3.5 (positional encoding) before P3.
- tiny Shakespeare — input.txt — Andrej Karpathy (2015, from char-rnn) · The actual corpus: 1,115,394 characters. Download it and run the numbers in this part yourself; every count here came from that file.
- Google Colab notebook for the lecture — Andrej Karpathy (2023) · The live notebook this part builds up, including the eight-examples printout and the seeded (4, 8) batch. The exercises below are meant to be done in a copy of it.
- ng-video-lecture — bigram.py, get_batch — Andrej Karpathy (2023) · The nine lines this part exists to produce, in their final form. Note that gpt.py ships the identical function with only the hyperparameters changed.
- Neural Machine Translation of Rare Words with Subword Units — Sennrich, Haddow, Birch (2016) · Field map extra. The BPE algorithm behind the "50,257 tokens" that Karpathy gestures at — merge the most frequent adjacent pair, repeat. Short and readable.
- Let's build the GPT Tokenizer — Andrej Karpathy (2024) · Field map extra. Two hours on precisely the subsystem this part skips, from the same series. Watch after you finish this lecture, not before; it is the answer to "but what if I want real tokens?"
- nanoGPT — data/openwebtext/prepare.py — Andrej Karpathy (2023) · Field map extra. What this part looks like at real scale: tokenize once, write train.bin/val.bin as raw uint16, then np.memmap them so random windows come off disk. Same shifted-by-one contract, no torch.tensor(encode(text)) of the whole corpus in RAM.
Exercises
- Swap the tokenizer for raw bytescode — In a copy of the Colab, replace encode/decode with UTF-8 byte encoding (list(s.encode('utf-8')) and bytes(l).decode('utf-8')) and set vocab_size = 256. Then: (1) report the new len(data) against 1,115,394 and explain why it barely moves for this corpus; (2) confirm decode(encode(text)) == text; (3) say what changed in the shape of the embedding table and what did not change anywhere else. A good answer names the one property tiny Shakespeare has — pure ASCII — that makes byte-level and character-level nearly the same here, and states what would break on a corpus with emoji.
- Break the loader on purpose, then prove it safecode — Set block_size = len(train_data) and call get_batch('train'). Predict the error before you run it, then run it. Next, change y to look two steps ahead (data[i+2 : i+block_size+2]) with a normal block_size, run it a few hundred times, and explain why the failure is intermittent and why it surfaces as a torch.stack error rather than an index error. Finish by writing the one-line assertion you would put at the top of get_batch to turn both failures into a clear message.
- Measure what your validation split is actually measuringcode — Compute the per-character frequency distribution of train_data and val_data and compare them (total variation distance is enough). Check how many of the 65 symbols appear in each. Then build a second, deliberately wrong split by sampling random positions instead of taking a contiguous tail, and argue in three sentences why a validation loss on that split would look far better while meaning far less. A good answer separates the two things the real gap contains: memorisation, and the fact that the last tenth of the file is different plays.