What micrograd is, and what a derivative is
Transcript: this part, with timestamps
This part is the setup, and it does two jobs that look unrelated but are not. First it shows you the finished toy — an expression graph, a .backward() call, gradients appearing on inputs — so you know what you are building toward and can judge how small the finished thing is. Then it deliberately backs all the way off to high-school calculus and re-derives what a derivative is, numerically, with no symbols. Those two halves meet in the third section of the lecture: the numbers that .backward() fills in are exactly the numbers you would get by bumping each input by hand, and the reason to build the engine at all is that bumping by hand does not scale.
Outline, with timestamps
- 00:00 — Intro: a blank notebook now, a trained neural net by the end, with nothing hidden.
- 00:25 — Micrograd overview: build an expression out of Value objects, read g.data for the forward pass, call g.backward() and read a.grad and b.grad.
- 04:34 — The demo expression is nonsense on purpose; what we actually care about is neural nets, which are also just expressions.
- 05:34 — Scalar-valued on purpose: tensors are packaging for parallelism, and none of the math changes.
- 06:50 — Two files, roughly 150 lines total: engine.py is the autograd, nn.py is the whole neural-net library.
- 08:08 — Derivative of a function of one input: define f(x) = 3x² − 4x + 5, evaluate it, plot the parabola.
- 10:06 — Why we will not do symbolic differentiation, and what the limit definition actually says.
- 11:13 — The bump experiment at x = 3 with h = 0.001: rise over run gives 14.003, and the analytic answer is 14.
- 13:15 — The same experiment at x = −3 (negative slope, −22) and near x = 2/3 (slope zero).
- 14:12 — Several inputs: d = a*b + c with a=2, b=−3, c=10, and three separate derivatives to find.
- 16:18 — Predict the sign before running it: bumping a makes d go down, slope −3; then b gives 2 and c gives 1.
- 18:48 — Neural nets are enormous expressions, so we need a data structure that remembers them. Enter Value, in P2.
What an autograd engine buys you
Strip the marketing off and micrograd does one thing: it lets you write ordinary arithmetic in Python and then, afterwards, ask a question about that arithmetic that ordinary arithmetic cannot answer. The question is "how sensitive was the result to each of the numbers I put in?"
The demo Karpathy scrolls through is the README example from the repo. Two inputs, a = Value(-4.0) and b = Value(2.0), get pushed through a deliberately silly chain — adds, multiplies, a power, a negation, a couple of ReLUs, a division — producing an output g. The forward pass is unremarkable: g.data is 24.7041, which is just the number the arithmetic produces. The interesting call is g.backward(). After it returns, a.grad is 138.8338 and b.grad is 645.5773. (In the video he reads these off screen as "138" and "645"; the repo's own README prints them to four decimals.)
Read those two numbers as physical claims about the expression, not as abstract calculus. They say: if you increase a by a tiny amount, g increases about 138.8 times as fast; do the same to b and g climbs about 645.6 times as fast. Both inputs push g up, and b has roughly 4.6× the leverage. Nobody wrote a formula for dg/da. The engine walked backwards through the graph it had recorded during the forward pass, applying the chain rule node by node, and left the answer sitting on each input.
Two framings from this stretch are worth holding onto. The first is that the demo expression is meaningless on purpose:
This expression, by the way, is completely meaningless. I just made it up. I'm just flexing about the kinds of operations that are supported by micrograd.04:34
The point of the flex is that backpropagation does not know or care what a neural network is. It works on arbitrary differentiable expressions; neural networks just happen to be a particular family of them, and one we care about. Keeping those two ideas separate — the general algorithm, the specific application — is what makes the rest of the lecture legible.
The second framing is the scale one. Micrograd works on individual scalars: a neuron is chopped into its component adds and multiplies, and every one of them is a node. That is not a pedagogical shortcut around some harder "real" math. It is the other way around. Real libraries pack thousands of scalars into tensors so that a GPU can do the adds in parallel, and none of the calculus changes when they do. Karpathy is explicit that this is purely efficiency. He is also, quietly, making a claim about how much of deep learning is actually conceptual:
My claim is that micrograd is what you need to train neural networks, and everything else is just efficiency.06:50
He backs it with a file count. The repo has two source files. engine.py is the autograd engine and knows nothing about neural nets; nn.py is the entire neural-network library on top of it. He calls them "100 lines" and "a joke", respectively; if you count the current master today it is 94 lines and 60 lines. That is the whole target of the lecture, and it is worth opening both files now, failing to understand them, and coming back at the end.
The derivative, without any symbols
At 08:08 the lecture restarts from nothing. New function, chosen to be boring: f(x) = 3x² − 4x + 5. Call it at 3.0 and you get 20.0. Plot it over np.arange(-5, 5, 0.25) and you get a parabola with its minimum a bit right of zero.
Now the question: what is the derivative of this at a given x? Everyone in the audience knows the symbolic route — apply the power rule, get 6x − 4, plug in. Karpathy refuses it, and the refusal is the load-bearing part of the whole video:
No one in neural networks actually writes out the expression for the neural net. It would be a massive expression… no one actually derives the derivative.10:06
So instead he goes back to the limit definition — the derivative at x is the limit as h → 0 of (f(x+h) − f(x)) / h — and reads it as an experiment rather than a formula. Nudge the input up by a small h. Measure how much the output moved. Divide by how much you moved the input, because you want a rate, not a raw change. That is rise over run, and the resulting number is a slope: sign tells you which way the function goes, magnitude tells you how hard.
He takes h = 0.001 — the limit says take it to zero, but you cannot type zero — and evaluates at x = 3. Before running it he asks you to predict: is f(3.001) a bit above 20 or a bit below? The parabola is climbing there, so above. The measured slope comes out 14.003, and the analytic answer, 6(3) − 4, is 14. Close enough to trust the method, off by exactly the amount the method should be off by.
Then he moves the probe. At x = −3 the parabola is falling, so the slope must be negative; 6(−3) − 4 = −22, and the numerical version agrees at about −21.997. And there is a point where the slope is zero — for this function at x = 2/3 — where nudging the input barely moves the output at all. That flat point is the whole reason gradients are useful later: training is a search for somewhere the loss has stopped responding to the weights.
| Probe point | f(x) | Numerical slope, h = 0.001 | Analytic 6x − 4 | Reading |
|---|---|---|---|---|
| x = 3.0 | 20.0 | 14.003 | 14 | climbing, steeply |
| x = −3.0 | 44.0 | −21.997 | −22 | falling, steeply |
| x = 2/3 | 3.6667 | ≈ 3.0e−06 | 0 | flat: the minimum |
He also flags, without dwelling on it, that shrinking h does not improve the answer forever — floating point has finite precision and at some point you "get into trouble". That hedge is worth unpacking, because it bites people the first time they write a gradient check; see Where people get stuck below for the numbers.
More than one input, and the shape of the answer changes
At 14:12 the function grows a second and third input: d = a*b + c, with a = 2.0, b = −3.0, c = 10.0, so d = 4.0. Now "the derivative" is not one number. There are three, one per input, each answering the same bump question about a different knob while the other two are held still. This is the moment the object of interest quietly becomes a vector of sensitivities — which is what a gradient is, and what a neural network's training step consumes.
The procedure is the same three lines each time, and it is worth typing rather than reading: fix the inputs, compute d1, bump exactly one input by h, recompute as d2, and print (d2 − d1) / h.
The bit of teaching craft here is that he predicts the sign out loud before running the cell, and gets one of them counter-intuitively right. Bumping a upward makes d go down, because a is multiplied by a negative b: more a means more of a negative contribution. The printout shows d2 = 3.999699999999999 against d1 = 4.0 — down, as predicted — and the slope lands at −3, which is exactly b. (On screen the caption reads this as 3.9996; the actual float is 3.9997 with a trailing rounding artifact.)
| Bump (h = 0.0001) | d1 | d2 | Measured slope | Analytic | Why |
|---|---|---|---|---|---|
| a += h | 4.0 | 3.9997 | −3.0000000000108 | ∂d/∂a = b = −3 | more a, more of a negative product |
| b += h | 4.0 | 4.0002 | 2.0000000000042 | ∂d/∂b = a = 2 | a is positive, so d rises |
| c += h | 4.0 | 4.0001 | 0.9999999999977 | ∂d/∂c = 1 | c is added straight on, one for one |
Look at the analytic column and notice something that will matter enormously in P3 and P4: ∂d/∂a is not a constant, it is b — the value of the other input at that point. Local derivatives at a multiply node are the sibling's data. At an add node they are 1, for every input, always. Those two facts are the entire content of __mul__._backward and __add__._backward when the engine gets written, and you can already see them here in a cell with no classes in it.
Why the hand method has to be replaced
Karpathy closes the section by saying neural nets will be "pretty massive expressions", so we need a data structure to hold them — and cuts to Value. That is true but it undersells the actual problem, and it is worth stating the sharper version because it explains why backpropagation is a named algorithm rather than an obvious trick. (Field-map extra: the cost argument below is not made explicitly in this part of the video; he gets to it implicitly in P3–P4.)
The bump method needs one full forward pass per input. Three inputs, three extra evaluations — fine. A small MLP has forty-one parameters, so forty-two evaluations — still fine. A network with a hundred million parameters needs a hundred million forward passes for one gradient, and you need a fresh gradient at every training step. Backpropagation gets all of them in a single backward pass whose cost is about the same as one forward pass, regardless of how many inputs there are. That asymmetry is why the rest of the lecture exists.
The bump method does not become useless, though. It becomes the test. Every time the engine produces a gradient you do not believe, you can re-derive it by nudging and dividing, exactly as in these cells — which is what PyTorch's own gradcheck does to this day, and what Karpathy does by hand at several points later in the video.
The code at the end of this part
import math
import numpy as np
import matplotlib.pyplot as plt
%matplotlib inline
def f(x):
return 3*x**2 - 4*x + 5
f(3.0)
xs = np.arange(-5, 5, 0.25)
ys = f(xs)
plt.plot(xs, ys)
h = 0.000001
x = 2/3
(f(x + h) - f(x))/h
# les get more complex
a = 2.0
b = -3.0
c = 10.0
d = a*b + c
print(d)
h = 0.0001
# inputs
a = 2.0
b = -3.0
c = 10.0
d1 = a*b + c
c += h
d2 = a*b + c
print('d1', d1)
print('d2', d2)
print('slope', (d2 - d1)/h)
Line by line, with the two places the saved notebook differs from the video:
- import math / numpy / matplotlib — the boilerplate he pastes into every notebook. math is not used yet; it shows up in P3 for math.exp inside tanh. Note what is absent: no torch, no autograd, nothing that could be doing the work for us.
- def f(x): return 3*x**2 - 4*x + 5 — deliberately a scalar function of a scalar, and deliberately a quadratic so its slope changes sign across the plotted range. Because it is written with numpy-compatible operators, the same f applies elementwise to the whole xs array a line later.
- xs = np.arange(-5, 5, 0.25) — 40 points, right endpoint excluded, purely to draw the parabola so you have a picture to point at when you talk about slope.
- h = 0.000001; x = 2/3; (f(x + h) - f(x))/h — the numerical-derivative cell, saved in the state he left it in after exploring. In the video this cell starts as h = 0.001 and x = 3.0, printing 14.003; he then walks it to x = -3.0 for −22 and finally to x = 2/3 for the flat point. If you are following along with the saved notebook, retype it at 3.0 first — the whole argument of the section is in the sequence, not the final cell.
- d = a*b + c — the three-input expression. Same numbers reappear in P2 as Value(2.0, label='a') and friends, and again in P3 as the front half of the L = (a*b + c) * f graph he backpropagates by hand. It is worth noticing that the arithmetic here is plain Python floats: nothing is being recorded, which is precisely the gap Value fills.
- d1 = a*b + c; c += h; d2 = a*b + c — the bump. The saved cell bumps c, which is the last of the three he tries; the video runs a += h first (slope −3), then b (slope 2), then c (slope 1). Also note the mutation style: he modifies the input in place and rebuilds d2, so re-running the cell twice without re-fixing the inputs silently bumps c twice. That is the kind of stateful-notebook trap that will bite much harder in P7, where forgetting to reset something (the gradients) actually breaks training.
- Nothing in this part exists in the finished library — engine.py line 5 is where Value.__init__ begins, and P2 is where the lecture reaches it. These cells are scaffolding for your intuition, not code that survives.
Where people get stuck
- "Why is the answer 14.003 and not 14? Is the method broken?" — No, and the error is predictable. The one-sided difference (f(x+h) − f(x))/h has error proportional to h, so with h = 0.001 you expect to be wrong in the third decimal, and you are: 14.00300000000243. Halve h and the error roughly halves. If you want a better estimate for free, use the central difference (f(x+h) − f(x−h))/(2h), whose error scales with h² — it returns 14 to machine precision here. Karpathy stays with the one-sided version because it matches the textbook limit definition he is trying to make concrete.
- "So make h tiny and it gets exact, right?" — This is the trap he hedges about at 12:14 without showing. It gets better, then worse. On this function at x = 3: h = 1e−3 gives 14.00300, h = 1e−6 gives 14.000003, h = 1e−10 gives 14.0000012 — and h = 1e−14 gives 14.2109, which is worse than where you started. The cause is catastrophic cancellation: f(x+h) and f(x) agree in the first fifteen digits, so subtracting them throws away almost all the significant bits, and then you divide the noise by a tiny number and amplify it. The sweet spot for a one-sided difference in float64 is around h ≈ 1e−8, i.e. roughly the square root of machine epsilon.
- "Why is the derivative of a*b + c with respect to a equal to b — isn't a derivative supposed to be a function of the variable you differentiate?" — It is a function of the whole input point; it just happens not to depend on a here. Holding b and c fixed, d is a straight line in a with slope b, so the sensitivity to a is whatever b currently is. Change b to 100 and ∂d/∂a becomes 100. This is exactly why the engine has to record the forward values before it can run backwards: the local derivative at a multiply node is the other operand's data at that point.
- "Scalar-valued sounds like a simplified version of the real thing." — It is a simplified version of the implementation, not of the mathematics. A tensor library computes the same chain rule; it just applies it to arrays so the hardware can vectorize. If you understand backprop over scalars you understand backprop, and the delta to a real framework is memory layout, kernel fusion and hardware, not calculus. Treat the "everything else is efficiency" line as a real claim you will be able to audit by P8, not as a slogan.
Go deeper, verified
- karpathy/micrograd — Andrej Karpathy (2020) · The finished library the lecture reconstructs. Read the README example first; it is the exact expression he demos at 00:25, with g.data = 24.7041, a.grad = 138.8338, b.grad = 645.5773 printed in the source.
- nn-zero-to-hero / lectures / micrograd — Andrej Karpathy (2022) · The two notebooks he actually types in, saved cell by cell. micrograd_lecture_first_half_roughly.ipynb holds every cell in this part; open it beside the video rather than pausing to transcribe.
- engine.py, Value.__init__ (line 5) — Andrej Karpathy (2020) · Where the wrapper object you have not met yet begins: data, grad, _backward, _prev, _op. Five attributes, 94 lines in the file. Skim it now to see the size of the destination.
- The official exercise Colab — Andrej Karpathy (2022) · Its opening section belongs to this part: write the analytic gradient of a three-input function by hand, then estimate the same gradient numerically and compare. Do it before P2, not after the whole lecture.
- CS231n: Backpropagation, Intuitions — Andrej Karpathy / Stanford CS231n · Field map extra. The written version of the same argument, with the "gates" framing and the local-gradient-times-upstream-gradient picture that P3 arrives at. Its section on numerical gradient checking covers the h-too-small failure in more detail than the video does.
- torch.autograd.gradcheck — PyTorch docs · Field map extra. Production code that does exactly the bump-and-divide from these cells (central differences, tolerances) to verify analytic gradients. Proof that the "hacky" method in this part is the standard correctness test, not a teaching-only device.
- Learning representations by back-propagating errors — Rumelhart, Hinton & Williams (1986) · Field map extra. The canonical backprop paper. Two pages, and the algorithm being automated in P4 is recognisably the one described here; useful for seeing how little the idea has changed.
Exercises
- Sweep h until it breaks.code — Write slope(x, h) = (f(x+h) - f(x))/h for f(x) = 3x² − 4x + 5, and evaluate it at x = 3 for h in [1e-1, 1e-2, … 1e-16]. Print the absolute error against the exact answer, 14. A good answer: the error falls linearly with h down to about 1e-8, then rises, and is catastrophic by 1e-15. Then add a central-difference version and sweep it too — its error should fall faster (quadratically) and bottom out lower. Say in one sentence why the two curves have a U-shape at all.
- Gradient by hand for three inputs.code — Take d(a, b, c) = a*b + c at a = 2, b = −3, c = 10. (1) Write the three partial derivatives symbolically. (2) Write a loop that bumps each input in turn by h = 1e-6 and prints the numerical estimate, so you get a three-element gradient vector. (3) Check that they agree to five decimals. (4) Now change b to 100 and rerun without touching your symbolic answers — predict which entries move before you look. This is the exercise Colab's opening section in miniature; do that section for the real version, which uses a nastier function.
- Predict the sign, then run. — Before touching a keyboard, write down the sign of ∂g/∂a and ∂g/∂b for the README expression at a = −4, b = 2, and which of the two you expect to have the larger magnitude. Then pip install micrograd, run the README snippet, and compare with 138.8338 and 645.5773. A good answer explains which operations in the chain amplify b's influence — the b**3 term and the squaring of e are the places to look — rather than just noting that you were right or wrong.