LET'S BUILD GPT // FIELD MAP
← field map
PART 01 · BASELINE00:00–22:11 · 22 min

The setup: the problem, the data, tokens, batches

Andrej Karpathy · Let's build GPT (2023) · part 01 of 8

Transcript: this part, with timestamps

TL;DR — The question this part answers is: what is the smallest honest version of the thing ChatGPT does, and what does its input look like? The answer is next-symbol prediction over a flat file of Shakespeare, with characters as symbols — 1,115,394 of them drawn from a 65-symbol vocabulary. Text becomes a 1-D tensor of integers; training data becomes random fixed-length windows out of that tensor, stacked into a batch. The one thing to remember is the shape contract that every later part inherits: inputs are (B, T) integers, targets are the same tensor shifted one step left, and a (4, 8) batch is not 4 training examples but 32.

Nothing in this part is a Transformer. That is deliberate, and it is the reason the rest of the lecture goes as fast as it does. Karpathy spends 22 minutes making the problem so concrete that the model becomes the only remaining unknown: by the end there is a tensor of integers, a function that hands you a random batch of windows out of it, and a precise statement of what a correct answer looks like at every position in that batch. Everything from 22:11 onward — the bigram lookup table, the four versions of attention, the full gpt.py — is a different function plugged into the same socket. The socket is what gets built here.

Outline, with timestamps

The problem, stated so it can be coded

The opening five minutes are a framing argument, and it is worth extracting because the rest of the lecture leans on it. ChatGPT completes sequences. Give it a prompt twice and you get two different answers, which tells you it is not looking anything up — it is sampling from a distribution over what comes next. The architecture doing that came out of a 2017 machine-translation paper, and the "GPT" acronym unpacks as generatively pretrained Transformer: the Transformer is the network, "generatively pretrained" is what you did to it.

Karpathy then makes the concession that defines the whole lecture: a production system is trained on a good chunk of the internet and then put through further stages, and none of that is reproducible in an afternoon. So he keeps the architecture and shrinks everything else. The corpus becomes one megabyte of Shakespeare. The alphabet becomes literal characters instead of subword tokens. What survives the shrinking is exactly the part he wants you to understand — the network — and the claim, which the rest of the lecture makes good on, is that nothing structural changes. The same code that babbles fake Shakespeare at 10 million parameters is the code that writes marketing copy at 175 billion. Only the numbers move.

Two places to keep your hedge-detector on. He dates GPT-2 to "2017 if I recall correctly" — GPT-2 is February 2019 (2017 is the Transformer paper, and GPT-1 is mid-2018); the point he is making, that nanoGPT reproduces the 124 M-parameter GPT-2 and can load OpenAI's released weights, is correct. And he describes the 2017 paper as reading "like a pretty random machine translation paper." That is fair as a reaction, but it is worth knowing why: the paper's pitch was parallelism and training cost for translation, and what this lecture builds is only the decoder half of the architecture it proposes. That asymmetry gets named explicitly much later, in P5's note 4 and again in P7.

The data: one flat string, 65 distinct symbols

The corpus is tiny Shakespeare, a file Karpathy has been using since char-rnn in 2015. Read it and count and you get 1,115,394 characters — about 1.1 MB, which at UTF-8 with a pure-ASCII corpus is also 1,115,394 bytes. It is a concatenation of plays with no metadata, no separators, no document boundaries. The only structure in it is typographic:

First Citizen:
Before we proceed any further, hear me speak.

All:
Speak, speak.

Speaker name, colon, newline, line of verse, blank line. Remember that shape — it is the first thing any model learns here, and in P2 you will watch a model that cannot see past one character reproduce the colon-and-newline rhythm while producing no English at all.

The vocabulary comes from three stacked calls: set(text) collapses to distinct characters, list(...) gives them an order, sorted(...) makes that order deterministic. The sort is the only part that matters and it is easy to skip past. Python's set iteration order is not something you want a trained checkpoint to depend on; sorting means the integer 46 means h today and h on any other machine. The result:

['\n', ' ', '!', '$', '&', "'", ',', '-', '.', '3', ':', ';', '?',
 'A'…'Z', 'a'…'z']        # 65 symbols

Thirteen punctuation-and-whitespace symbols, then the two alphabets. Two details are worth a second: the only digit in all of Shakespeare-as-distributed is 3, and there is no uppercase/lowercase merging — A and a are unrelated integers as far as the model is concerned, and it has to learn their relationship from data like everything else. vocab_size = 65 is then wired into two places for the rest of the lecture: the number of rows in the embedding table, and the number of columns in the final logits.

Tokenization is a dial, not a design

The tokenizer here is two dictionaries and two one-liners: stoi maps character to integer, itos maps back, encode is a list comprehension and decode is a join. Encoding "hi there" gives [46, 47, 1, 58, 46, 43, 56, 43] — eight integers for eight characters, and decode(encode(s)) == s for any string over the vocabulary.

Karpathy then does something more useful than explaining his own tokenizer: he shows you what you are giving up. Google's SentencePiece and OpenAI's tiktoken are subword schemes — units bigger than a character, smaller than a word. Run the same "hi there" through GPT-2's tiktoken encoding and you get three integers instead of eight, drawn from a vocabulary of 50,257 instead of 65.

That is the whole trade, and it is a genuine dial with a cost on both ends:

SchemeVocabulary"hi there"What it costs
character (this lecture)658 integerslong sequences; attention is O(T²), so context is expensive
subword (GPT-2 BPE)50,2573 integersa 50k-row embedding table and a 50k-wide output layer

Small codebook, long sequences; large codebook, short sequences. Practice sits in the middle because both ends are quadratic in something you care about. Character level is the right choice for a lecture not because it is better but because it deletes an entire subsystem — no merge table, no training a tokenizer, no byte-fallback edge cases — while leaving every downstream line of code identical. Karpathy is explicit that this is a simplification, and he later gave the skipped subsystem its own two hours; that video is linked below.

The consequence to hold onto: a character-level model with block_size = 8 can see about a word and a half of history. Even the final scaled-up gpt.py at block_size = 256 sees roughly fifty words. When the samples in P7 read like Shakespeare-shaped noise rather than coherent scenes, the binding constraint is context, not parameters — and the tokenizer is half of why.

Splitting, and what a validation loss here actually measures

The encoded corpus goes into a single 1-D torch.Tensor of dtype long — one integer per character, 1,115,394 of them — and then gets cut 90/10:

TensorLengthSlice
data1,115,394the whole corpus
train_data1,003,854data[:n], first 90%
val_data111,540data[n:], last 10%

The stated motivation is overfitting detection: we do not want a network that has memorised this exact text, we want one that produces text like it, so hold some back and watch whether the two losses separate. That is right, and in P7 the gap between them becomes the main diagnostic for whether the model is too big for the data.

What goes unsaid is worth knowing. The split is contiguous, not random. Contiguous is the correct choice — a randomly-held-out character sitting between two training characters is trivially predictable, so a random split would make the validation loss meaningless. But contiguity has a price: the last 10% of the file is a different stretch of plays, with different speakers and different vocabulary density. So the train/val gap you measure is partly memorisation and partly distribution shift between one end of the corpus and the other. At the scale of this lecture that is fine. At the scale where you would care about the difference, you would hold out whole documents.

Chunks, and why nine characters are eight examples

You never feed a corpus to a Transformer. You feed fixed-length windows, and the length is block_size — Karpathy's name for what other codebases call context length. Here it starts at 8.

The first thing he prints is train_data[:block_size+1], nine integers: [18, 47, 56, 57, 58, 1, 15, 47, 58], which decodes to "First Cit". The +1 is the crux of the data format. Inputs are the first eight, targets are the same window shifted one step left — so the target for position t is the character that actually followed the prefix ending at t. A nine-integer window therefore contains eight (context, target) pairs, with contexts of every length from 1 to 8:

x = train_data[:block_size]        # 18 47 56 57 58  1 15 47
y = train_data[1:block_size+1]     # 47 56 57 58  1 15 47 58

for t in range(block_size):
    context = x[:t+1]
    target  = y[t]
    print(f"when input is {context.tolist()} the target: {target}")

Reading it out: given [18] predict 47; given [18, 47] predict 56; given [18, 47, 56] predict 57; and so on to the full eight-character prefix. One window, eight supervised examples.

Karpathy is careful to say this is not merely a way to squeeze value out of data you already loaded:

We train on all the eight examples here with context between one all the way up to context of block_size… it's also done to make the Transformer network be used to seeing contexts all the way from as little as one all the way to block_size. Andrej Karpathy, 17:02 (cleaned auto-captions)

This is a generation requirement, not a training convenience. In P2 sampling starts from a single token and grows the context one character at a time. If the network had only ever seen full-length contexts, its first seven predictions would be out of distribution every time you sampled. Training on all prefix lengths is what makes the model well-defined on a one-character prompt.

There is a second reason he leaves implicit, and it is the one that makes Transformer language modelling economically different from most supervised learning: the prediction at every position is computed anyway. Because of the causal mask you will meet in P3, position t's output depends only on positions ≤ t, so a single forward pass over an eight-token window legitimately produces eight independent predictions from eight different amounts of history. You get T training signals for the price of one pass. Nothing is being reused improperly; the mask is what makes it honest.

The batch dimension, and the shape contract

The second dimension exists purely to keep the hardware busy. GPUs are wide; a single 8-integer sequence uses almost none of that width, so we stack several unrelated windows into one tensor and let them ride through the network together. They never interact — a fact that gets a whole numbered note in P5 — they just share a kernel launch.

get_batch does three things. It picks which split to read from; it draws batch_size random offsets into that split with torch.randint; and it torch.stacks the resulting 1-D windows into rows of a 2-D tensor. With the notebook's batch_size = 4 and block_size = 8, and torch.manual_seed(1337) so the numbers are reproducible, you get:

NameShapeContents
ix(4,)random start offsets into train_data
xb(4, 8) = (B, T)four independent 8-character windows, as integers
yb(4, 8) = (B, T)the same four windows shifted one step left

The first row begins 24, 43, 58, 5 — decode it and you get "Let'". Karpathy walks it prefix by prefix: input 24 → target 43; input 24, 43 → target 58; input 24, 43, 58 → target 5. Four rows times eight positions is 32 supervised examples in one (4, 8) tensor, and that count is the single most useful thing to carry out of this part. The loss you will see in P2 is a mean over 32 numbers, not 4.

Those chunks are processed completely independently, they don't talk to each other. Andrej Karpathy, 18:33 (cleaned auto-captions)
Carry this forward: B, T, C. Batch, time, channels. The input to every model in this lecture is (B, T) integers; the output is (B, T, C) floats where C = vocab_size at the very end; the targets are (B, T) integers, shifted. From here to the final gpt.py, that socket never changes — only the function between the two ends, and the value of C in the middle layers.

The code at the end of this part

By 22:11 the notebook holds everything below. This is the excerpt of bigram.py that this part produces — the file it eventually gets written into at 38:00, in P2. The corpus, the vocabulary, the tokenizer, the split:

# wget https://raw.githubusercontent.com/karpathy/char-rnn/master/data/tinyshakespeare/input.txt
with open('input.txt', 'r', encoding='utf-8') as f:
    text = f.read()

# here are all the unique characters that occur in this text
chars = sorted(list(set(text)))
vocab_size = len(chars)
# create a mapping from characters to integers
stoi = { ch:i for i,ch in enumerate(chars) }
itos = { i:ch for i,ch in enumerate(chars) }
encode = lambda s: [stoi[c] for c in s] # encoder: take a string, output a list of integers
decode = lambda l: ''.join([itos[i] for i in l]) # decoder: take a list of integers, output a string

# Train and test splits
data = torch.tensor(encode(text), dtype=torch.long)
n = int(0.9*len(data)) # first 90% will be train, rest val
train_data = data[:n]
val_data = data[n:]

Line by line (L17–L34): the encoding='utf-8' matters only because Python would otherwise guess from the locale. chars is the 65 symbols in sorted order, and vocab_size is derived from the data rather than hard-coded — swap the corpus and everything downstream re-sizes itself. stoi and itos are exact inverses by construction because both come from the same enumerate. torch.long is not optional: these integers are used as indices into an embedding table and as class labels for cross-entropy, and both APIs require 64-bit integers. And int(0.9*len(data)) is a plain truncation — n = 1,003,854.

Then the loader, which is the actual deliverable of this part (L36–L44):

batch_size = 32 # how many independent sequences will we process in parallel?
block_size = 8  # what is the maximum context length for predictions?

# data loading
def get_batch(split):
    # generate a small batch of data of inputs x and targets y
    data = train_data if split == 'train' else val_data
    ix = torch.randint(len(data) - block_size, (batch_size,))
    x = torch.stack([data[i:i+block_size] for i in ix])
    y = torch.stack([data[i+1:i+block_size+1] for i in ix])
    x, y = x.to(device), y.to(device)
    return x, y

Five lines of body, each doing one thing. data = train_data if … deliberately shadows the module-level data inside the function, which is tidy but will bite you if you ever want the full corpus in here. torch.randint(hi, (batch_size,)) draws batch_size integers uniformly from [0, hi) — sampling with replacement, so a batch can in principle contain the same window twice, and there is no notion of an epoch anywhere in this codebase. The two list comprehensions build the input and target windows, offset by exactly one. torch.stack creates a new leading dimension, turning batch_size tensors of shape (8,) into one of shape (batch_size, 8) — that new axis is B. And the .to(device) happens here, in the loader, so the model never has to think about where its inputs live.

Two numbers to note while you are here. The script version uses batch_size = 32 where the notebook demo used 4 — the (4, 8) tensor you were just staring at is (32, 8) once the code is a script. And with max_iters = 3000 in P2, training touches 3000 × 32 × 8 = 768,000 character positions out of 1,003,854 available — less than a single pass over the training split. At bigram scale that is plenty.

Where people get stuck

Go deeper, verified

Exercises

  1. Swap the tokenizer for raw bytescode — In a copy of the Colab, replace encode/decode with UTF-8 byte encoding (list(s.encode('utf-8')) and bytes(l).decode('utf-8')) and set vocab_size = 256. Then: (1) report the new len(data) against 1,115,394 and explain why it barely moves for this corpus; (2) confirm decode(encode(text)) == text; (3) say what changed in the shape of the embedding table and what did not change anywhere else. A good answer names the one property tiny Shakespeare has — pure ASCII — that makes byte-level and character-level nearly the same here, and states what would break on a corpus with emoji.
  2. Break the loader on purpose, then prove it safecode — Set block_size = len(train_data) and call get_batch('train'). Predict the error before you run it, then run it. Next, change y to look two steps ahead (data[i+2 : i+block_size+2]) with a normal block_size, run it a few hundred times, and explain why the failure is intermittent and why it surfaces as a torch.stack error rather than an index error. Finish by writing the one-line assertion you would put at the top of get_batch to turn both failures into a clear message.
  3. Measure what your validation split is actually measuringcode — Compute the per-character frequency distribution of train_data and val_data and compare them (total variation distance is enough). Check how many of the 65 symbols appear in each. Then build a second, deliberately wrong split by sampling random positions instead of taking a contiguous tail, and argue in three sentences why a validation loss on that split would look far better while meaning far less. A good answer separates the two things the real gap contains: memorisation, and the fact that the last tenth of the file is different plays.
Next: P02 The bigram baseline: model, loss, generation, training · Back to the map.