MICROGRAD // FIELD MAP
← field map
PART 03 · BACKPROP32:10–69:02 · 37 min

Backpropagation by hand: an expression, then a neuron

Andrej Karpathy · building micrograd (2022) · part 03 of 8

Transcript: this part, with timestamps

TL;DR — The graph from P2 can be evaluated forwards; this part evaluates it backwards, by hand, one node at a time. The whole procedure is a single rule applied repeatedly: the gradient arriving at a node, times the node's own local derivative, is the gradient going to each of its inputs. On the toy expression L = (a·b + c)·f that gives a.grad = 6, b.grad = −4, c.grad = −2, f.grad = 4, every one of them confirmed by bumping the input and re-running the forward pass. Nudging all four inputs along their gradients by 0.01 moves L from −8.0 to −7.286496 — one step of training, done with nothing but arithmetic. The same procedure then runs through a two-input tanh neuron, where the bias is chosen so the answers land on 0.5 and 1.0 and you can check them in your head. Remember the rule, not the numbers: local derivative × upstream gradient, applied recursively from the output backwards.

By 32:10 the lecture has everything it needs and nothing it wants. P1 established what a derivative is — bump the input by h, see how much the output moves, divide. P2 built a Value that remembers its children and the operation that produced it, so any expression is now a graph you can draw. What is missing is the connection: a derivative is defined at the leaves of that graph, but the graph is built from the leaves toward the output. Backpropagation is the algorithm that closes that gap, and Karpathy's choice here is to run it manually — no code, no automation, just typing numbers into .grad fields and redrawing the picture — for a full thirty-seven minutes. The automation in P4 is fifteen lines. It is only obvious after you have done it by hand, which is exactly why this part is the longest single stretch of the lecture.

Outline, with timestamps

Start at the end, and check everything

The first gradient is free. Backpropagation computes dL/dx for every node x in the graph, and the graph's own output is L, so the first question is dL/dL — how much does L change when you change L by h? It changes by h. The answer is 1.0, and Karpathy types it into L.grad directly. That is not a trick; it is the boundary condition that makes everything downstream well-posed. Every other gradient in the graph is eventually derived from this one.

Before deriving anything, though, he builds a way to be wrong out loud. The lol() function at 32:39 is a deliberately disposable scope: it constructs the expression twice, bumps exactly one input by h = 0.001 in the second copy, and prints (L2 − L1)/h. Everything is local to the function, so the global notebook state stays clean — a small hygiene move that matters when you are going to run this a dozen times in a row with a different variable bumped each time. What it produces is the numerical gradient: the same quantity backprop claims to compute, obtained by brute force. Karpathy calls the practice a gradient check, and he uses it after every single claim in this section. It is the single most transferable habit in the part. When your hand-derived backward pass and your finite-difference estimate agree to five decimal places, you are done arguing.

One detail worth noticing: he has to write L.data, not L, in the subtraction, because L is now a Value object and there is no __sub__ on it yet. Subtraction and division do not arrive until P5. The numerical check runs on raw Python floats living underneath the graph.

One node at a time: local derivative × upstream gradient

The expression is a = 2.0, b = −3.0, c = 10.0, f = −2.0, with e = a·b, d = e + c, L = d·f. Forward: e = −6.0, d = 4.0, L = −8.0.

The first real step is the multiply at the top. L = d·f, so dL/dd = f = −2.0 and dL/df = d = 4.0. Karpathy does not just quote the product rule; he expands the limit definition on screen at 35:14: ((d+h)·f − d·f)/h = (d·f + h·f − d·f)/h = f. The d·f terms cancel, the h divides out, and what survives is the other input. That is the whole content of a multiply node's backward pass, and it is worth internalizing in that form: a multiply node hands each input the value of the other one.

Then comes the node the rest of the lecture rests on. c and e do not touch L directly — they only reach it through d. So the plus node is in a strange epistemic position: it knows it added two numbers to make d, and it knows nothing whatsoever about the graph it is embedded in. All it can compute is its own local derivative, dd/dc, which the same limit expansion shows is exactly 1.0 (((c+h) + e − c − e)/h = h/h = 1). What it needs from the outside world is a single number: how much L cares about d. That number, −2.0, was computed one step earlier and is sitting in d.grad.

now we're getting to the crux of backpropagation … this will be the most important node to understand, because if you understand the gradient for this node you understand all of backpropagation and all of training of neural nets, basically.Karpathy, 37:47

The chain rule is what glues the local fact to the global one. Karpathy pulls up the Wikipedia statement at 41:21 in its readable form — if z depends on y and y depends on x, then dz/dx = (dz/dy)·(dy/dx) — and then reaches for the intuition that makes multiplication feel inevitable rather than arbitrary: a car is twice as fast as a bicycle, a bicycle is four times as fast as a walking man, so the car is eight times as fast as the man. Rates of change compose by multiplying. Substituting z = L, y = d, x = c: dL/dc = (dL/dd)·(dd/dc) = −2.0 × 1.0 = −2.0.

Because the local derivative is 1.0, the plus node changes nothing. It copies the incoming gradient to each of its inputs unchanged — it routes. c.grad = −2.0 and e.grad = −2.0. That is why every experienced practitioner's mental picture of a sum in a computation graph is a wire, not an operation.

One more step reaches the leaves. e = a·b, so the local derivatives are de/da = b = −3.0 and de/db = a = 2.0, and the arriving gradient is e.grad = −2.0. Multiply: a.grad = −2.0 × −3.0 = 6.0, b.grad = −2.0 × 2.0 = −4.0.

NodedatagradWhere the grad comes fromNumerical check
L−8.01.0Base case: dL/dL1.0
f−2.04.0Multiply node hands over the other input, d3.9999999999995595
d4.0−2.0Multiply node hands over the other input, f−1.9999999999953388
c10.0−2.0Plus node routes: 1.0 × d.grad−1.9999999999953388
e−6.0−2.0Plus node routes: 1.0 × d.grad−2.0
a2.06.0b.data × e.grad = −3.0 × −2.06.000000000021544
b−3.0−4.0a.data × e.grad = 2.0 × −2.0−3.9999999999995595

The numerical column is what he actually prints in the video, and the trailing digits are not noise to be embarrassed about — they are the fingerprint of finite differences. The estimate is biased by O(h) and corrupted by floating-point cancellation in the numerator, so agreement to ten digits is as good as it gets. That mismatch pattern is itself diagnostic: an error in the eleventh decimal is float behaviour, an error in the first is a wrong derivative.

that's what backpropagation is — just a recursive application of chain rule backwards through the computation graph.Karpathy, 50:37
The rule, in the form that survives the rest of the series: at every node, gradient-out × local-derivative = gradient-in. The node knows only its own operation; the graph supplies everything else through the single number sitting in out.grad. Plus routes it, multiply swaps the inputs and scales, tanh scales by 1 − out². Everything after this — the closures in P4, the training loop in P7, PyTorch itself — is bookkeeping around that one line.

One optimization step, previewed

At 51:10 Karpathy stops to show what the gradients are for, and it takes four lines. Each leaf moves a small distance in the direction of its own gradient, x.data += 0.01 * x.grad, and then the forward pass is re-run. Note the sign: he is deliberately making L go up here, because L is not yet a loss — it is just an output. In P7, once L is a loss to be minimized, the same line gets a minus sign in front of the step and nothing else changes.

The arithmetic: a goes 2.0 → 2.06, b goes −3.0 → −3.04, c goes 10.0 → 9.98, f goes −2.0 → −1.96. Forward again: e = −6.2624, d = 3.7176, L = −7.286496. He guesses "maybe negative six or so" before running it and lands on −7.29, which is worth noticing as an honest miss — the first-order prediction is that L rises by 0.01 × (6² + 4² + 2² + 4²) = 0.72, and it rises by 0.713. The gradient tells you the direction reliably and the magnitude only for small steps. That gap is the entire subject of learning rates in P7.

The neuron: the same procedure, on a shape that matters

The second worked example is chosen so the machinery lands on something recognizable. A neuron in the standard cartoon takes inputs x, multiplies each by a synaptic weight w, sums them, adds a bias b that Karpathy glosses as the neuron's innate "trigger-happiness", and passes the total through a squashing activation function. Written out with every intermediate given its own labelled Value — which is the point, since you need pointers to those intermediates to inspect their gradients — it becomes n = x1·w1 + x2·w2 + b and o = tanh(n).

tanh forces a decision. It cannot be built from the + and * that Value currently supports; it needs exponentiation. Karpathy notes he could implement exp and compose — and in P5 he goes back and does exactly that, as a demonstration — but here he implements tanh directly as one node, and makes the general point explicit.

we can create functions at arbitrary points of abstraction … the only thing that matters is that we know how to differentiate through any one function.Karpathy, 58:15

This is the load-bearing design principle of every autodiff library. A node can be as coarse or as fine as you like — a single addition, a whole tanh, an entire fused attention kernel — provided you can state its local derivative. Granularity is a performance and convenience decision, not a correctness one.

Then the bias. At 61:19 Karpathy sets b = 6.8813735870195432 and says plainly that he picked it so the numbers come out nice. They do, and it is worth seeing why. With x1 = 2.0, w1 = −3.0, x2 = 0.0, w2 = 1.0, the weighted sum is −6.0, so n = 0.8813735870195432 — and that is precisely artanh(1/√2). Therefore o = tanh(n) = 0.7071067811865476 = 1/√2, and the tanh derivative 1 − o² is exactly 1 − ½ = 0.5. Every gradient in the backward pass is then a small multiple of one half. It is a staged demo, and he says so; the honesty is the point, because it means you can check his arithmetic mentally instead of trusting the printout.

The gradients through the neuron

Same procedure, five steps. Start at o.grad = 1.0. Through the tanh: the derivative of tanh(x) is 1 − tanh²(x), which Karpathy looks up on Wikipedia rather than deriving, and since tanh(n) is already sitting in o.data the local derivative is 1 − o.data² = 0.5. So n.grad = 0.5. Then two plus nodes in a row, each routing that 0.5 unchanged to both children: the bias b, the partial sum, and then x1·w1 and x2·w2 all get 0.5. Finally two multiply nodes, each handing over the other input scaled by the 0.5 that arrived.

NodedatagradRule applied
o0.70710678118654761.0base case
n0.88137358701954320.5(1 − o.data²) × o.grad
b6.88137358701954320.5plus routes
x1*w1 + x2*w2−6.00.5plus routes
x1*w1−6.00.5plus routes
x2*w20.00.5plus routes
x12.0−1.5w1.data × 0.5 = −3.0 × 0.5
w1−3.01.0x1.data × 0.5 = 2.0 × 0.5
x20.00.5w2.data × 0.5 = 1.0 × 0.5
w21.00.0x2.data × 0.5 = 0.0 × 0.5

The interesting entry is the zero. w2.grad = 0 because x2 = 0, and Karpathy pauses at 67:01 to make sure the intuition lands: wiggling w2 cannot change the output, because whatever you set it to gets multiplied by zero. A gradient of zero is not a failure of the method, it is a true statement about this network at this input — that weight is invisible to this example, and gradient descent will leave it exactly where it is. The mirror-image fact is w1.grad = 1.0, the largest gradient in the graph: raise w1 and the neuron fires harder, at a rate of one to one.

Note also the practical asymmetry. In training you want the gradients on w1 and w2, because weights are what you change. The gradients on x1 and x2 fall out of the same pass for free, and in a single neuron they are useless — data is fixed. In a deeper network they are not useless at all: the x of one layer is the o of the layer below, so those input gradients are exactly the signal that keeps propagating. That is why backprop is worth automating rather than special-casing.

The code at the end of this part

The gradient-check harness, verbatim from the notebook. It bumps b; in the video the bumped line moves around as he checks each variable in turn.

def lol():

  h = 0.001

  a = Value(2.0, label='a')
  b = Value(-3.0, label='b')
  c = Value(10.0, label='c')
  e = a*b; e.label = 'e'
  d = e + c; d.label = 'd'
  f = Value(-2.0, label='f')
  L = d * f; L.label = 'L'
  L1 = L.data

  a = Value(2.0, label='a')
  b = Value(-3.0, label='b')
  b.data += h
  c = Value(10.0, label='c')
  e = a*b; e.label = 'e'
  d = e + c; d.label = 'd'
  f = Value(-2.0, label='f')
  L = d * f; L.label = 'L'
  L2 = L.data

  print((L2 - L1)/h)

lol()

Prints -3.9999999999995595, matching the hand-derived b.grad = −4. The whole expression is rebuilt from scratch in both halves because Value nodes cache their data at construction time — there is no forward-pass replay yet, so mutating b.data after the fact would not update e, d or L.

One optimization step, previewed:

a.data += 0.01 * a.grad
b.data += 0.01 * b.grad
c.data += 0.01 * c.grad
f.data += 0.01 * f.grad

e = a * b
d = e + c
L = d * f

print(L.data)

Four update lines and a forward pass. This is the entire skeleton of the training loop that arrives in P7; the only differences there are that the parameters come from n.parameters() instead of being named by hand, and the step is negated because the objective becomes a loss.

The neuron, verbatim, with every intermediate labelled so it can be read off the graph:

# inputs x1,x2
x1 = Value(2.0, label='x1')
x2 = Value(0.0, label='x2')
# weights w1,w2
w1 = Value(-3.0, label='w1')
w2 = Value(1.0, label='w2')
# bias of the neuron
b = Value(6.8813735870195432, label='b')
# x1*w1 + x2*w2 + b
x1w1 = x1*w1; x1w1.label = 'x1*w1'
x2w2 = x2*w2; x2w2.label = 'x2*w2'
x1w1x2w2 = x1w1 + x2w2; x1w1x2w2.label = 'x1*w1 + x2*w2'
n = x1w1x2w2 + b; n.label = 'n'
o = n.tanh(); o.label = 'o'

And the tanh method as it stands at the end of this part — the exponential form of tanh, one child in the tuple, 'tanh' as the op label so draw_dot prints it:

def tanh(self):
    x = self.data
    t = (math.exp(2*x) - 1)/(math.exp(2*x) + 1)
    out = Value(t, (self, ), 'tanh')
    return out

Two notes on that cell. First, (self, ) needs the trailing comma — it is a one-element tuple, not a parenthesised expression, and dropping the comma is a real bug that silently makes _prev iterate over the Value instead of containing it. Second, in the saved notebook this method has a _backward closure attached; that is added in P4, not here, which is why the version above stops at return out. The finished library takes a different route again — engine.py ships relu, not tanh.

Finally, the by-hand backward pass through the neuron. In the notebook these are separate cells typed bottom-up as he worked backwards; reordered here into the order the gradients actually flow:

o.grad = 1.0
n.grad = 0.5                      # 1 - o.data**2, looked up on Wikipedia
x1w1x2w2.grad = 0.5               # plus routes
b.grad = 0.5                      # plus routes
x1w1.grad = 0.5                   # plus routes
x2w2.grad = 0.5                   # plus routes
x2.grad = w2.data * x2w2.grad     # 1.0 * 0.5  ->  0.5
w2.grad = x2.data * x2w2.grad     # 0.0 * 0.5  ->  0.0
x1.grad = w1.data * x1w1.grad     # -3.0 * 0.5 -> -1.5
w1.grad = x1.data * x1w1.grad     # 2.0 * 0.5  ->  1.0

Ten assignments for a neuron with two inputs. This is the "obviously ridiculous" that P4 opens by putting an end to: each of those right-hand sides is a fixed function of the node's operation and its children, so each belongs on the node itself. Compare the last two lines with the closure inside __mul__ in engine.py — self.grad += other.data * out.grad — and the correspondence is exact except for the +=, which is the subject of the accumulation bug in P4.

Where people get stuck

Go deeper, verified

Exercises

  1. Redo backprop #1 with one constant changedcode — Set f = 3.0 instead of −2.0 and, before running anything, write down all six gradients on paper. Then rebuild the expression, fill in .grad by hand exactly as in the video, and confirm each one with a lol()-style bump. A good answer explains why a.grad flips sign from 6 to −9 in terms of the two multiply nodes it passes through, and notes that c.grad and e.grad stay equal to each other no matter what you change, because the plus node in between routes.
  2. Break the numerical check on purposecode — Take dL/da = 6 and estimate it with h = 1, 1e-2, 1e-4, 1e-6, 1e-8, 1e-10, 1e-14. Plot |estimate − 6| against h on log axes. You should see error fall as h shrinks, bottom out somewhere around 1e-6 to 1e-8, and then climb again as catastrophic cancellation in L2 − L1 takes over. Then repeat with the central difference (f(x+h) − f(x−h))/(2h) and note that the floor drops by several orders of magnitude. This is section 1 of the official exercise Colab, whose third part asks for exactly this symmetric-difference improvement — the backprop half of that notebook belongs to P5, but the gradient-checking half is this part's material.
  3. The neuron with x2 switched on — Re-run the manual neuron backward pass with x2 = 1.0 instead of 0.0, pencil first. Note that n is no longer 0.8813735870195432, so 1 − o² is no longer a clean 0.5 and every downstream number gets messy — which is the whole reason Karpathy staged the bias. A good answer states what w2.grad becomes and why it is no longer zero, and checks the largest and smallest gradients against a numerical bump.
Next: P04 Automating backward: per-op closures, topological order, and the accumulation bug · Back to the map.