Stanford CS336 · Spring 2025 · Percy Liang & Tatsunori Hashimoto · seventeen lectures

Build the whole thing. Then make it good.

CS336 is the operating-systems course of language models: tokenizer, transformer, kernels, parallelism, scaling laws, data, evaluation and alignment, each built by hand. This map gives every lecture a deep-dive explainer written from the talk with timestamped links back into the video, the cleaned transcript on its own page, the slides or executable lecture, the assignment it feeds, and supporting materials that were checked before they were linked.

$stanford-cs336.github.io/spring2025
▲ amber = build the model — tokens, architecture, systems ▼ cyan = make it good — scaling, data, evaluation, alignment
00 the map

Five arcs, five assignments

The course runs the pipeline in order and hands you an assignment at each hinge. The deep dives follow the same order; the assignments section says which lectures each one needs.

ARC 1 · lectures 1–4

Build it

Tokenization, PyTorch and resource accounting, the architecture and hyperparameter choices that survived, mixture of experts.

ARC 2 · lectures 5–8

Make it fast

What a GPU actually is, kernels and Triton, then parallelism across devices in two lectures.

ARC 3 · lectures 9–12

Scale it and serve it

Scaling laws twice, inference, and how to evaluate a model without fooling yourself.

ARC 4 · lectures 13–14

Feed it

Where pretraining data comes from, how it is filtered, deduplicated and mixed, and why it matters more than the architecture.

ARC 5 · lectures 15–17

Align it

SFT and RLHF, then reinforcement learning with verifiable rewards in two lectures.

HOW TO USE THIS

Watch, read, build

Each lecture page opens with a TL;DR and a timestamped outline so you can jump into the video where it matters; the transcript page is for search and quoting; the assignment card says what to build.

01 arc 1 · amber

Build it

lectures 1–4 · ~5.5 h
L01 · Apr 1 · Percy · 78 min

Overview and tokenization

What the course is for, why "from scratch", and byte-pair encoding built up from bytes.

L02 · Apr 3 · Percy · 79 min

PyTorch, resource accounting

Tensors, memory, FLOPs and the arithmetic that tells you whether a training run is even possible.

L03 · Apr 8 · Tatsu · 87 min

Architectures, hyperparameters

The transformer variants that stuck: norms, activations, positional encodings, and the hyperparameters people copy.

L04 · Apr 10 · Tatsu · 82 min

Mixture of experts

Sparse experts, routing, load balancing, and why the frontier models went this way.

02 arc 2 · amber

Make it fast

lectures 5–8 · ~5.2 h
L05 · Apr 15 · Tatsu · 74 min

GPUs

The memory hierarchy, arithmetic intensity, and the roofline that explains most performance surprises.

L06 · Apr 17 · Tatsu · 80 min

Kernels, Triton

Writing your own kernels, fusion, and when the compiler beats you.

L07 · Apr 22 · Tatsu · 84 min

Parallelism 1

Data, tensor and pipeline parallelism, and the communication costs that bound them.

L08 · Apr 24 · Percy · 75 min

Parallelism 2

The executable version: collective ops in code, sharding in practice, and what breaks.

03 arc 3 · cyan

Scale it and serve it

lectures 9–12 · ~5.1 h
L09 · Apr 29 · Tatsu · 65 min

Scaling laws 1

Power laws, Chinchilla, and how to spend a compute budget before you have it.

L10 · May 1 · Percy · 82 min

Inference

KV caches, batching, speculative decoding, and why serving is a different engineering problem from training.

L11 · May 6 · Tatsu · 78 min

Scaling laws 2

The details that decide whether a scaling study means anything: fits, hyperparameter transfer, muP and its cousins.

L12 · May 8 · Percy · 80 min

Evaluation

Perplexity, benchmarks, contamination, and what a number on a leaderboard does and does not tell you.

04 arc 4 · cyan

Feed it

lectures 13–14 · ~2.6 h
L13 · May 13 · Percy · 79 min

Data 1

Common Crawl to a training set: extraction, filtering, deduplication, and the datasets that shaped the models.

L14 · May 15 · Percy · 79 min

Data 2

Mixtures, quality classifiers, synthetic data, copyright and the legal edges of the corpus.

05 arc 5 · cyan

Align it

lectures 15–17 · ~3.8 h
L15 · May 20 · Tatsu · 74 min

Alignment: SFT and RLHF

Instruction tuning, preference data, reward models, and DPO as the shortcut.

L16 · May 22 · Tatsu · 80 min

Alignment: RL 1

Reinforcement learning with verifiable rewards: policy gradients, GRPO, and the reasoning-model recipe.

L17 · May 27 · Percy · 76 min

Alignment: RL 2

The executable version of RL for reasoning: the loop in code, and what the recent papers actually changed.

Two guest lectures (Junyang Lin on Qwen, May 29; Mike Lewis, June 3) closed the quarter but are not in the public playlist, so they have no pages here.
06 the assignments

Five things you build

Each assignment has a walkthrough page here: every problem with its points, the adapter and test that grade it, the data, the papers and the lecture behind it, all linked. The handouts and repos are pinned to their Spring 2025 versions.

A1 · Basicsafter L1–L4 · repo · handout · leaderboard
BPE tokenizer, the transformer, the optimizer and the training loop, from nothing, then train on TinyStories and OpenWebText.
A2 · Systemsafter L5–L8 · repo · handout · leaderboard
Profile it, write a FlashAttention-style Triton kernel, then distribute training with data parallel and optimizer sharding.
A3 · Scalingafter L9, L11 · repo · handout
Spend a fixed budget of training runs through an API to fit a scaling law and predict the best model at a bigger budget.
A4 · Dataafter L13–L14 · repo · handout · leaderboard
Turn raw Common Crawl into a training set: extraction, language ID, quality and toxicity filters, deduplication, then train on it.
SFT, expert iteration and GRPO on math reasoning; the supplement covers safety RLHF and DPO.
07 progress

Tick them off

Seventeen lectures and five assignments. The checks live in this browser only.

PROGRESS
0%
08 refs

Where this came from

The lectures are Stanford's, taught by Percy Liang and Tatsunori Hashimoto with CAs Neil Band, Marcel Rød and Rohith Kuditipudi. Every page links the video it explains and the course materials for that day. The explanations are original; the transcripts are cleaned auto-captions and will contain caption errors.