Backpropagation by hand: an expression, then a neuron
Transcript: this part, with timestamps
By 32:10 the lecture has everything it needs and nothing it wants. P1 established what a derivative is — bump the input by h, see how much the output moves, divide. P2 built a Value that remembers its children and the operation that produced it, so any expression is now a graph you can draw. What is missing is the connection: a derivative is defined at the leaves of that graph, but the graph is built from the leaves toward the output. Backpropagation is the algorithm that closes that gap, and Karpathy's choice here is to run it manually — no code, no automation, just typing numbers into .grad fields and redrawing the picture — for a full thirty-seven minutes. The automation in P4 is fifteen lines. It is only obvious after you have done it by hand, which is exactly why this part is the longest single stretch of the lecture.
Outline, with timestamps
- 32:10 — Manual backprop #1 begins: the base case, dL/dL = 1, typed into L.grad by hand.
- 32:39 — The lol() staging function: a throwaway scope for numerically checking every gradient claim before believing it.
- 34:42 — One step back: for L = d·f, dL/dd = f and dL/df = d, derived from the limit definition rather than recalled.
- 38:17 — The crux node: c and e only reach L through d. The local derivative of a plus node is 1.0.
- 41:21 — The chain rule, stated in the useful form, with the car/bicycle/walking-man intuition for why the rates multiply.
- 45:29 — A plus node routes: c.grad = e.grad = −2, unchanged from what arrived. Verified numerically.
- 47:01 — Through the times node to the leaves: a.grad = 6, b.grad = −4. Multiply nodes swap the inputs.
- 51:10 — Preview of one optimization step: nudge every leaf by 0.01 × grad, re-run the forward pass, watch L go from −8.0 to −7.286496.
- 52:52 — Manual backprop #2: the mathematical model of a neuron, weights and bias and a squashing activation.
- 57:14 — Why tanh is implemented as a single node: the level of abstraction is a free choice, bounded only by what you can differentiate.
- 62:52 — Backprop through tanh: do/dn = 1 − o², which with the chosen bias comes out to exactly 0.5.
- 66:28 — The last two multiply nodes: x1.grad = −1.5, w1.grad = 1.0, x2.grad = 0.5, and w2.grad = 0 — and why the zero is correct.
Start at the end, and check everything
The first gradient is free. Backpropagation computes dL/dx for every node x in the graph, and the graph's own output is L, so the first question is dL/dL — how much does L change when you change L by h? It changes by h. The answer is 1.0, and Karpathy types it into L.grad directly. That is not a trick; it is the boundary condition that makes everything downstream well-posed. Every other gradient in the graph is eventually derived from this one.
Before deriving anything, though, he builds a way to be wrong out loud. The lol() function at 32:39 is a deliberately disposable scope: it constructs the expression twice, bumps exactly one input by h = 0.001 in the second copy, and prints (L2 − L1)/h. Everything is local to the function, so the global notebook state stays clean — a small hygiene move that matters when you are going to run this a dozen times in a row with a different variable bumped each time. What it produces is the numerical gradient: the same quantity backprop claims to compute, obtained by brute force. Karpathy calls the practice a gradient check, and he uses it after every single claim in this section. It is the single most transferable habit in the part. When your hand-derived backward pass and your finite-difference estimate agree to five decimal places, you are done arguing.
One detail worth noticing: he has to write L.data, not L, in the subtraction, because L is now a Value object and there is no __sub__ on it yet. Subtraction and division do not arrive until P5. The numerical check runs on raw Python floats living underneath the graph.
One node at a time: local derivative × upstream gradient
The expression is a = 2.0, b = −3.0, c = 10.0, f = −2.0, with e = a·b, d = e + c, L = d·f. Forward: e = −6.0, d = 4.0, L = −8.0.
The first real step is the multiply at the top. L = d·f, so dL/dd = f = −2.0 and dL/df = d = 4.0. Karpathy does not just quote the product rule; he expands the limit definition on screen at 35:14: ((d+h)·f − d·f)/h = (d·f + h·f − d·f)/h = f. The d·f terms cancel, the h divides out, and what survives is the other input. That is the whole content of a multiply node's backward pass, and it is worth internalizing in that form: a multiply node hands each input the value of the other one.
Then comes the node the rest of the lecture rests on. c and e do not touch L directly — they only reach it through d. So the plus node is in a strange epistemic position: it knows it added two numbers to make d, and it knows nothing whatsoever about the graph it is embedded in. All it can compute is its own local derivative, dd/dc, which the same limit expansion shows is exactly 1.0 (((c+h) + e − c − e)/h = h/h = 1). What it needs from the outside world is a single number: how much L cares about d. That number, −2.0, was computed one step earlier and is sitting in d.grad.
now we're getting to the crux of backpropagation … this will be the most important node to understand, because if you understand the gradient for this node you understand all of backpropagation and all of training of neural nets, basically.Karpathy, 37:47
The chain rule is what glues the local fact to the global one. Karpathy pulls up the Wikipedia statement at 41:21 in its readable form — if z depends on y and y depends on x, then dz/dx = (dz/dy)·(dy/dx) — and then reaches for the intuition that makes multiplication feel inevitable rather than arbitrary: a car is twice as fast as a bicycle, a bicycle is four times as fast as a walking man, so the car is eight times as fast as the man. Rates of change compose by multiplying. Substituting z = L, y = d, x = c: dL/dc = (dL/dd)·(dd/dc) = −2.0 × 1.0 = −2.0.
Because the local derivative is 1.0, the plus node changes nothing. It copies the incoming gradient to each of its inputs unchanged — it routes. c.grad = −2.0 and e.grad = −2.0. That is why every experienced practitioner's mental picture of a sum in a computation graph is a wire, not an operation.
One more step reaches the leaves. e = a·b, so the local derivatives are de/da = b = −3.0 and de/db = a = 2.0, and the arriving gradient is e.grad = −2.0. Multiply: a.grad = −2.0 × −3.0 = 6.0, b.grad = −2.0 × 2.0 = −4.0.
| Node | data | grad | Where the grad comes from | Numerical check |
|---|---|---|---|---|
| L | −8.0 | 1.0 | Base case: dL/dL | 1.0 |
| f | −2.0 | 4.0 | Multiply node hands over the other input, d | 3.9999999999995595 |
| d | 4.0 | −2.0 | Multiply node hands over the other input, f | −1.9999999999953388 |
| c | 10.0 | −2.0 | Plus node routes: 1.0 × d.grad | −1.9999999999953388 |
| e | −6.0 | −2.0 | Plus node routes: 1.0 × d.grad | −2.0 |
| a | 2.0 | 6.0 | b.data × e.grad = −3.0 × −2.0 | 6.000000000021544 |
| b | −3.0 | −4.0 | a.data × e.grad = 2.0 × −2.0 | −3.9999999999995595 |
The numerical column is what he actually prints in the video, and the trailing digits are not noise to be embarrassed about — they are the fingerprint of finite differences. The estimate is biased by O(h) and corrupted by floating-point cancellation in the numerator, so agreement to ten digits is as good as it gets. That mismatch pattern is itself diagnostic: an error in the eleventh decimal is float behaviour, an error in the first is a wrong derivative.
that's what backpropagation is — just a recursive application of chain rule backwards through the computation graph.Karpathy, 50:37
One optimization step, previewed
At 51:10 Karpathy stops to show what the gradients are for, and it takes four lines. Each leaf moves a small distance in the direction of its own gradient, x.data += 0.01 * x.grad, and then the forward pass is re-run. Note the sign: he is deliberately making L go up here, because L is not yet a loss — it is just an output. In P7, once L is a loss to be minimized, the same line gets a minus sign in front of the step and nothing else changes.
The arithmetic: a goes 2.0 → 2.06, b goes −3.0 → −3.04, c goes 10.0 → 9.98, f goes −2.0 → −1.96. Forward again: e = −6.2624, d = 3.7176, L = −7.286496. He guesses "maybe negative six or so" before running it and lands on −7.29, which is worth noticing as an honest miss — the first-order prediction is that L rises by 0.01 × (6² + 4² + 2² + 4²) = 0.72, and it rises by 0.713. The gradient tells you the direction reliably and the magnitude only for small steps. That gap is the entire subject of learning rates in P7.
The neuron: the same procedure, on a shape that matters
The second worked example is chosen so the machinery lands on something recognizable. A neuron in the standard cartoon takes inputs x, multiplies each by a synaptic weight w, sums them, adds a bias b that Karpathy glosses as the neuron's innate "trigger-happiness", and passes the total through a squashing activation function. Written out with every intermediate given its own labelled Value — which is the point, since you need pointers to those intermediates to inspect their gradients — it becomes n = x1·w1 + x2·w2 + b and o = tanh(n).
tanh forces a decision. It cannot be built from the + and * that Value currently supports; it needs exponentiation. Karpathy notes he could implement exp and compose — and in P5 he goes back and does exactly that, as a demonstration — but here he implements tanh directly as one node, and makes the general point explicit.
we can create functions at arbitrary points of abstraction … the only thing that matters is that we know how to differentiate through any one function.Karpathy, 58:15
This is the load-bearing design principle of every autodiff library. A node can be as coarse or as fine as you like — a single addition, a whole tanh, an entire fused attention kernel — provided you can state its local derivative. Granularity is a performance and convenience decision, not a correctness one.
Then the bias. At 61:19 Karpathy sets b = 6.8813735870195432 and says plainly that he picked it so the numbers come out nice. They do, and it is worth seeing why. With x1 = 2.0, w1 = −3.0, x2 = 0.0, w2 = 1.0, the weighted sum is −6.0, so n = 0.8813735870195432 — and that is precisely artanh(1/√2). Therefore o = tanh(n) = 0.7071067811865476 = 1/√2, and the tanh derivative 1 − o² is exactly 1 − ½ = 0.5. Every gradient in the backward pass is then a small multiple of one half. It is a staged demo, and he says so; the honesty is the point, because it means you can check his arithmetic mentally instead of trusting the printout.
The gradients through the neuron
Same procedure, five steps. Start at o.grad = 1.0. Through the tanh: the derivative of tanh(x) is 1 − tanh²(x), which Karpathy looks up on Wikipedia rather than deriving, and since tanh(n) is already sitting in o.data the local derivative is 1 − o.data² = 0.5. So n.grad = 0.5. Then two plus nodes in a row, each routing that 0.5 unchanged to both children: the bias b, the partial sum, and then x1·w1 and x2·w2 all get 0.5. Finally two multiply nodes, each handing over the other input scaled by the 0.5 that arrived.
| Node | data | grad | Rule applied |
|---|---|---|---|
| o | 0.7071067811865476 | 1.0 | base case |
| n | 0.8813735870195432 | 0.5 | (1 − o.data²) × o.grad |
| b | 6.8813735870195432 | 0.5 | plus routes |
| x1*w1 + x2*w2 | −6.0 | 0.5 | plus routes |
| x1*w1 | −6.0 | 0.5 | plus routes |
| x2*w2 | 0.0 | 0.5 | plus routes |
| x1 | 2.0 | −1.5 | w1.data × 0.5 = −3.0 × 0.5 |
| w1 | −3.0 | 1.0 | x1.data × 0.5 = 2.0 × 0.5 |
| x2 | 0.0 | 0.5 | w2.data × 0.5 = 1.0 × 0.5 |
| w2 | 1.0 | 0.0 | x2.data × 0.5 = 0.0 × 0.5 |
The interesting entry is the zero. w2.grad = 0 because x2 = 0, and Karpathy pauses at 67:01 to make sure the intuition lands: wiggling w2 cannot change the output, because whatever you set it to gets multiplied by zero. A gradient of zero is not a failure of the method, it is a true statement about this network at this input — that weight is invisible to this example, and gradient descent will leave it exactly where it is. The mirror-image fact is w1.grad = 1.0, the largest gradient in the graph: raise w1 and the neuron fires harder, at a rate of one to one.
Note also the practical asymmetry. In training you want the gradients on w1 and w2, because weights are what you change. The gradients on x1 and x2 fall out of the same pass for free, and in a single neuron they are useless — data is fixed. In a deeper network they are not useless at all: the x of one layer is the o of the layer below, so those input gradients are exactly the signal that keeps propagating. That is why backprop is worth automating rather than special-casing.
The code at the end of this part
The gradient-check harness, verbatim from the notebook. It bumps b; in the video the bumped line moves around as he checks each variable in turn.
def lol():
h = 0.001
a = Value(2.0, label='a')
b = Value(-3.0, label='b')
c = Value(10.0, label='c')
e = a*b; e.label = 'e'
d = e + c; d.label = 'd'
f = Value(-2.0, label='f')
L = d * f; L.label = 'L'
L1 = L.data
a = Value(2.0, label='a')
b = Value(-3.0, label='b')
b.data += h
c = Value(10.0, label='c')
e = a*b; e.label = 'e'
d = e + c; d.label = 'd'
f = Value(-2.0, label='f')
L = d * f; L.label = 'L'
L2 = L.data
print((L2 - L1)/h)
lol()
Prints -3.9999999999995595, matching the hand-derived b.grad = −4. The whole expression is rebuilt from scratch in both halves because Value nodes cache their data at construction time — there is no forward-pass replay yet, so mutating b.data after the fact would not update e, d or L.
One optimization step, previewed:
a.data += 0.01 * a.grad
b.data += 0.01 * b.grad
c.data += 0.01 * c.grad
f.data += 0.01 * f.grad
e = a * b
d = e + c
L = d * f
print(L.data)
Four update lines and a forward pass. This is the entire skeleton of the training loop that arrives in P7; the only differences there are that the parameters come from n.parameters() instead of being named by hand, and the step is negated because the objective becomes a loss.
The neuron, verbatim, with every intermediate labelled so it can be read off the graph:
# inputs x1,x2
x1 = Value(2.0, label='x1')
x2 = Value(0.0, label='x2')
# weights w1,w2
w1 = Value(-3.0, label='w1')
w2 = Value(1.0, label='w2')
# bias of the neuron
b = Value(6.8813735870195432, label='b')
# x1*w1 + x2*w2 + b
x1w1 = x1*w1; x1w1.label = 'x1*w1'
x2w2 = x2*w2; x2w2.label = 'x2*w2'
x1w1x2w2 = x1w1 + x2w2; x1w1x2w2.label = 'x1*w1 + x2*w2'
n = x1w1x2w2 + b; n.label = 'n'
o = n.tanh(); o.label = 'o'
And the tanh method as it stands at the end of this part — the exponential form of tanh, one child in the tuple, 'tanh' as the op label so draw_dot prints it:
def tanh(self):
x = self.data
t = (math.exp(2*x) - 1)/(math.exp(2*x) + 1)
out = Value(t, (self, ), 'tanh')
return out
Two notes on that cell. First, (self, ) needs the trailing comma — it is a one-element tuple, not a parenthesised expression, and dropping the comma is a real bug that silently makes _prev iterate over the Value instead of containing it. Second, in the saved notebook this method has a _backward closure attached; that is added in P4, not here, which is why the version above stops at return out. The finished library takes a different route again — engine.py ships relu, not tanh.
Finally, the by-hand backward pass through the neuron. In the notebook these are separate cells typed bottom-up as he worked backwards; reordered here into the order the gradients actually flow:
o.grad = 1.0
n.grad = 0.5 # 1 - o.data**2, looked up on Wikipedia
x1w1x2w2.grad = 0.5 # plus routes
b.grad = 0.5 # plus routes
x1w1.grad = 0.5 # plus routes
x2w2.grad = 0.5 # plus routes
x2.grad = w2.data * x2w2.grad # 1.0 * 0.5 -> 0.5
w2.grad = x2.data * x2w2.grad # 0.0 * 0.5 -> 0.0
x1.grad = w1.data * x1w1.grad # -3.0 * 0.5 -> -1.5
w1.grad = x1.data * x1w1.grad # 2.0 * 0.5 -> 1.0
Ten assignments for a neuron with two inputs. This is the "obviously ridiculous" that P4 opens by putting an end to: each of those right-hand sides is a fixed function of the node's operation and its children, so each belongs on the node itself. Compare the last two lines with the closure inside __mul__ in engine.py — self.grad += other.data * out.grad — and the correspondence is exact except for the +=, which is the subject of the accumulation bug in P4.
Where people get stuck
- "Why is the derivative of a + b with respect to a equal to 1, and why does that mean the gradient is copied?" — Bump a by h and the sum goes up by exactly h; rise over run is 1. Then the chain rule multiplies the upstream gradient by that 1, which leaves it unchanged. The two facts are separate and both matter: the local derivative is 1 because addition is unit-slope in each argument, and the routing behaviour follows from multiplying by 1. If you skip the chain-rule step you will be surprised later when a plus node does change the gradient — it never does, but a plus node whose output feeds two places contributes twice, which is a different phenomenon (P4's accumulation bug).
- "Why is w2.grad zero? Did something break?" — No. w2 is multiplied by x2 = 0, so no change to w2 can change the output. The gradient is a statement about this network at this input, and zero is the honest answer. Practical consequence: gradient descent will not move that weight on this example. Feed a different example where x2 ≠ 0 and it moves. This is also the mechanism behind dead ReLU units and behind why data that is identically zero in a feature teaches the model nothing about the corresponding weight.
- "Why 1 − o² and not 1 − n²?" — The derivative of tanh at a point n is 1 − tanh²(n), and tanh(n) is the node's own output, which is already computed and stored in o.data. Writing it in terms of the output is not a different formula, it is the same formula with a free lunch: the forward pass already paid for the expensive part. Autodiff libraries do this constantly — PyTorch's own tanh backward is the same expression on the saved output, which the lecture goes and reads in P8.
- "He writes = here but the library writes +=. Which is right?" — += is right, and this part gets away with = only because in both examples every node feeds exactly one consumer. The moment a value is used twice — b = a + a is the minimal case — the second assignment overwrites the first and you lose half the gradient. Karpathy hits this deliberately at 82:28 in P4. If you are typing along by hand, use = here as he does, but expect the change and understand that it is the multivariable chain rule: contributions along distinct paths add.
Go deeper, verified
- engine.py, __add__ (L13–L22) — Andrej Karpathy · the plus node's routing behaviour, mechanized: self.grad += out.grad for both children, exactly the by-hand step at 45:29.
- engine.py, __mul__ (L24–L33) — Andrej Karpathy · the swap-the-other-input rule, mechanized. Read it next to the last two lines of the by-hand neuron backward pass above.
- Chain rule — Wikipedia · the page Karpathy pulls up at 41:21, including the car/bicycle/walking-man intuition he quotes. The section on the multivariable case is the one that explains P4's +=.
- Hyperbolic functions — Wikipedia · where d/dx tanh(x) = 1 − tanh²(x) comes from, and the exponential definition used in the tanh method.
- CS231n: Backpropagation, Intuitions — Andrej Karpathy et al. (2015–) · field map extra · the written version of this same material, with the "add gate is a gradient distributor, multiply gate is a gradient switcher" framing and staged circuit diagrams. The closest thing to a textbook for this part.
- Learning representations by back-propagating errors — Rumelhart, Hinton & Williams (1986) · field map extra · the paper that put this procedure in front of the field. Worth reading once to see how much of the modern practice is already there and how differently it is written.
- Gradcheck mechanics — PyTorch docs · field map extra · the industrial version of the lol() function: why the central difference beats the forward difference, how to choose h, and why the tolerances are what they are.
Exercises
- Redo backprop #1 with one constant changedcode — Set f = 3.0 instead of −2.0 and, before running anything, write down all six gradients on paper. Then rebuild the expression, fill in .grad by hand exactly as in the video, and confirm each one with a lol()-style bump. A good answer explains why a.grad flips sign from 6 to −9 in terms of the two multiply nodes it passes through, and notes that c.grad and e.grad stay equal to each other no matter what you change, because the plus node in between routes.
- Break the numerical check on purposecode — Take dL/da = 6 and estimate it with h = 1, 1e-2, 1e-4, 1e-6, 1e-8, 1e-10, 1e-14. Plot |estimate − 6| against h on log axes. You should see error fall as h shrinks, bottom out somewhere around 1e-6 to 1e-8, and then climb again as catastrophic cancellation in L2 − L1 takes over. Then repeat with the central difference (f(x+h) − f(x−h))/(2h) and note that the floor drops by several orders of magnitude. This is section 1 of the official exercise Colab, whose third part asks for exactly this symmetric-difference improvement — the backprop half of that notebook belongs to P5, but the gradient-checking half is this part's material.
- The neuron with x2 switched on — Re-run the manual neuron backward pass with x2 = 1.0 instead of 0.0, pencil first. Note that n is no longer 0.8813735870195432, so 1 − o² is no longer a clean 0.5 and every downstream number gets messy — which is the whole reason Karpathy staged the bias. A good answer states what w2.grad becomes and why it is no longer zero, and checks the largest and smallest gradients against a numerical bump.