MICROGRAD // FIELD MAP
← all maps
ANDREJ KARPATHY · 2022 · ZERO TO HERO #12 h 25 min · 22 chapters · 8 parts here

micrograd, mapped

Karpathy builds a scalar autograd engine from an empty notebook, backpropagates through it by hand until the rule is obvious, then automates it, wraps it in Neuron, Layer and MLP, and trains a network by hand. This map splits the lecture into eight parts you can read before, during or after watching: what each stretch answers, the worked numbers, the code it arrives at, where people get stuck, and exercises. Every timestamp opens the video at that moment.

Tick a part when you've worked through it; this browser remembers.

01 · 00:00–32:10

Derivatives, and a data structure that remembers

What a derivative measures, numerically, before any calculus; then a Value object that records how every number was made, so the graph of an expression can be drawn and, later, walked backwards.

P1
▶ 00:0019 minintro · micrograd overview · derivative of one input · derivative with several inputs
Why a 100-line engine is enough to train a neural net, what the slope of f(x) = 3x² − 4x + 5 at a point means, and the bump-it-by-h experiment that defines the derivative for the whole lecture.
P2
▶ 19:0913 minValue with data, _prev, _op · __add__ and __mul__ · draw_dot with graphviz
A number that knows its children and the operation that produced it, and the picture of an expression as a directed graph that the rest of the lecture is drawn on.
02 · 32:10–103:55

Backpropagation, by hand and then automated

The heart of the lecture. Gradients computed manually node by node until the chain rule is obvious; then a _backward closure per operation, a topological sort, a bug when a value is used twice, and the same thing checked against PyTorch.

P3
▶ 32:1037 minmanual backprop #1 · one optimization step, previewed · manual backprop #2: a neuron with tanh
Filling in grad from the output backwards, the local-derivative-times-upstream rule stated and checked numerically, and one nudge of the inputs that makes the loss move the way the gradients promised.
P4
▶ 69:0218 min_backward for + × tanh · backward() over a topological sort · the b = a + a bug
Each operation stores how to push a gradient to its inputs, backward() visits nodes in reverse topological order, and gradients must accumulate with += or a value used twice gets the wrong answer.
P5
▶ 87:0517 minexp, division as pow, subtraction, radd/rmul · tanh rebuilt from exp · PyTorch tensors give the same gradients
The level of abstraction is a choice: tanh can be one node or a composition of exp and division and the gradients agree; PyTorch does the same with tensors and requires_grad.
03 · 103:55–136:46

A neural net, trained by hand

Neuron, Layer and MLP on top of Value; a four-example dataset; a mean-squared-error loss; the parameters collected; and gradient descent stepped manually, including the classic mistake of not zeroing the gradients.

P6
▶ 103:5517 minNeuron · Layer · MLP([3,4,4,1]) · four examples · mean squared error · parameters()
Forty-one parameters in three lines of classes, a loss that is itself a Value so backward() reaches every weight, and the moment the graph of the whole network is drawn.
P7
▶ 121:1216 minp.data += −lr × p.grad · overshooting · forgetting to zero grads · the loop · summary
Forward, backward, update, repeat; why the loss can go up; the bug where stale gradients accumulate across steps; and Karpathy's summary of how this scales to modern networks.
04 · 136:46–145:52

The real code

The finished micrograd repository read top to bottom, then a dig into PyTorch's own backward for tanh to show it is the same rule in C++.

P8
▶ 136:469 minengine.py · nn.py · the demo notebook · PyTorch's tanh_backward · conclusion
Where every lecture cell lives in the repo, what the library adds (ReLU, pow, the moons demo), and the line in PyTorch that computes 1 − tanh² just like the notebook did.
05

The exercise

Karpathy sets one, in the video description: a Colab notebook you should be able to complete after the lecture. Each part above also sets its own smaller exercises.

  1. The official exercise notebook. Open the Colab: derive the gradient of a function analytically and numerically, extend Value with the operations needed for a softmax-and-negative-log-likelihood loss, and verify against PyTorch. Sections map to P1, P5 and P7.
  2. Next in the series. makemore part 1 builds a bigram character model on this foundation; the seventh video is mapped at Let's build GPT, mapped.
06

Sources