MICROGRAD // FIELD MAP
← field map
PART 01 · DERIVATIVES00:00–19:09 · 19 min

What micrograd is, and what a derivative is

Andrej Karpathy · building micrograd (2022) · part 01 of 8

Transcript: this part, with timestamps

TL;DR — Micrograd is an autograd engine: you build a mathematical expression out of a wrapper object called Value, call .backward() on the output, and every input in the expression is left holding the derivative of that output with respect to itself. That derivative is not a symbolic formula you derived on paper — it is a slope, the answer to "if I nudge this input by a hair, how much and in which direction does the output move?" Karpathy spends these nineteen minutes making sure you feel that definition in your fingers, by literally bumping numbers by h = 0.001 and dividing. The one thing to carry forward: a gradient is a per-input sensitivity number, and the entire rest of the lecture is machinery for getting all of them cheaply.

This part is the setup, and it does two jobs that look unrelated but are not. First it shows you the finished toy — an expression graph, a .backward() call, gradients appearing on inputs — so you know what you are building toward and can judge how small the finished thing is. Then it deliberately backs all the way off to high-school calculus and re-derives what a derivative is, numerically, with no symbols. Those two halves meet in the third section of the lecture: the numbers that .backward() fills in are exactly the numbers you would get by bumping each input by hand, and the reason to build the engine at all is that bumping by hand does not scale.

Outline, with timestamps

What an autograd engine buys you

Strip the marketing off and micrograd does one thing: it lets you write ordinary arithmetic in Python and then, afterwards, ask a question about that arithmetic that ordinary arithmetic cannot answer. The question is "how sensitive was the result to each of the numbers I put in?"

The demo Karpathy scrolls through is the README example from the repo. Two inputs, a = Value(-4.0) and b = Value(2.0), get pushed through a deliberately silly chain — adds, multiplies, a power, a negation, a couple of ReLUs, a division — producing an output g. The forward pass is unremarkable: g.data is 24.7041, which is just the number the arithmetic produces. The interesting call is g.backward(). After it returns, a.grad is 138.8338 and b.grad is 645.5773. (In the video he reads these off screen as "138" and "645"; the repo's own README prints them to four decimals.)

Read those two numbers as physical claims about the expression, not as abstract calculus. They say: if you increase a by a tiny amount, g increases about 138.8 times as fast; do the same to b and g climbs about 645.6 times as fast. Both inputs push g up, and b has roughly 4.6× the leverage. Nobody wrote a formula for dg/da. The engine walked backwards through the graph it had recorded during the forward pass, applying the chain rule node by node, and left the answer sitting on each input.

Two framings from this stretch are worth holding onto. The first is that the demo expression is meaningless on purpose:

This expression, by the way, is completely meaningless. I just made it up. I'm just flexing about the kinds of operations that are supported by micrograd.04:34

The point of the flex is that backpropagation does not know or care what a neural network is. It works on arbitrary differentiable expressions; neural networks just happen to be a particular family of them, and one we care about. Keeping those two ideas separate — the general algorithm, the specific application — is what makes the rest of the lecture legible.

The second framing is the scale one. Micrograd works on individual scalars: a neuron is chopped into its component adds and multiplies, and every one of them is a node. That is not a pedagogical shortcut around some harder "real" math. It is the other way around. Real libraries pack thousands of scalars into tensors so that a GPU can do the adds in parallel, and none of the calculus changes when they do. Karpathy is explicit that this is purely efficiency. He is also, quietly, making a claim about how much of deep learning is actually conceptual:

My claim is that micrograd is what you need to train neural networks, and everything else is just efficiency.06:50

He backs it with a file count. The repo has two source files. engine.py is the autograd engine and knows nothing about neural nets; nn.py is the entire neural-network library on top of it. He calls them "100 lines" and "a joke", respectively; if you count the current master today it is 94 lines and 60 lines. That is the whole target of the lecture, and it is worth opening both files now, failing to understand them, and coming back at the end.

The derivative, without any symbols

At 08:08 the lecture restarts from nothing. New function, chosen to be boring: f(x) = 3x² − 4x + 5. Call it at 3.0 and you get 20.0. Plot it over np.arange(-5, 5, 0.25) and you get a parabola with its minimum a bit right of zero.

Now the question: what is the derivative of this at a given x? Everyone in the audience knows the symbolic route — apply the power rule, get 6x − 4, plug in. Karpathy refuses it, and the refusal is the load-bearing part of the whole video:

No one in neural networks actually writes out the expression for the neural net. It would be a massive expression… no one actually derives the derivative.10:06

So instead he goes back to the limit definition — the derivative at x is the limit as h → 0 of (f(x+h) − f(x)) / h — and reads it as an experiment rather than a formula. Nudge the input up by a small h. Measure how much the output moved. Divide by how much you moved the input, because you want a rate, not a raw change. That is rise over run, and the resulting number is a slope: sign tells you which way the function goes, magnitude tells you how hard.

He takes h = 0.001 — the limit says take it to zero, but you cannot type zero — and evaluates at x = 3. Before running it he asks you to predict: is f(3.001) a bit above 20 or a bit below? The parabola is climbing there, so above. The measured slope comes out 14.003, and the analytic answer, 6(3) − 4, is 14. Close enough to trust the method, off by exactly the amount the method should be off by.

Then he moves the probe. At x = −3 the parabola is falling, so the slope must be negative; 6(−3) − 4 = −22, and the numerical version agrees at about −21.997. And there is a point where the slope is zero — for this function at x = 2/3 — where nudging the input barely moves the output at all. That flat point is the whole reason gradients are useful later: training is a search for somewhere the loss has stopped responding to the weights.

Probe pointf(x)Numerical slope, h = 0.001Analytic 6x − 4Reading
x = 3.020.014.00314climbing, steeply
x = −3.044.0−21.997−22falling, steeply
x = 2/33.6667≈ 3.0e−060flat: the minimum

He also flags, without dwelling on it, that shrinking h does not improve the answer forever — floating point has finite precision and at some point you "get into trouble". That hedge is worth unpacking, because it bites people the first time they write a gradient check; see Where people get stuck below for the numbers.

More than one input, and the shape of the answer changes

At 14:12 the function grows a second and third input: d = a*b + c, with a = 2.0, b = −3.0, c = 10.0, so d = 4.0. Now "the derivative" is not one number. There are three, one per input, each answering the same bump question about a different knob while the other two are held still. This is the moment the object of interest quietly becomes a vector of sensitivities — which is what a gradient is, and what a neural network's training step consumes.

The procedure is the same three lines each time, and it is worth typing rather than reading: fix the inputs, compute d1, bump exactly one input by h, recompute as d2, and print (d2 − d1) / h.

The bit of teaching craft here is that he predicts the sign out loud before running the cell, and gets one of them counter-intuitively right. Bumping a upward makes d go down, because a is multiplied by a negative b: more a means more of a negative contribution. The printout shows d2 = 3.999699999999999 against d1 = 4.0 — down, as predicted — and the slope lands at −3, which is exactly b. (On screen the caption reads this as 3.9996; the actual float is 3.9997 with a trailing rounding artifact.)

Bump (h = 0.0001)d1d2Measured slopeAnalyticWhy
a += h4.03.9997−3.0000000000108∂d/∂a = b = −3more a, more of a negative product
b += h4.04.00022.0000000000042∂d/∂b = a = 2a is positive, so d rises
c += h4.04.00010.9999999999977∂d/∂c = 1c is added straight on, one for one

Look at the analytic column and notice something that will matter enormously in P3 and P4: ∂d/∂a is not a constant, it is b — the value of the other input at that point. Local derivatives at a multiply node are the sibling's data. At an add node they are 1, for every input, always. Those two facts are the entire content of __mul__._backward and __add__._backward when the engine gets written, and you can already see them here in a cell with no classes in it.

Carry this away: a gradient is a measurement, not a formula. a.grad = −3 means "push a up by ε and d moves down by 3ε." Every later abstraction — _backward closures, topological order, p.data += -0.01 * p.grad — exists only to produce that same measurement for millions of inputs without running millions of forward passes. If a gradient ever confuses you later in this lecture, come back and bump the thing by h.

Why the hand method has to be replaced

Karpathy closes the section by saying neural nets will be "pretty massive expressions", so we need a data structure to hold them — and cuts to Value. That is true but it undersells the actual problem, and it is worth stating the sharper version because it explains why backpropagation is a named algorithm rather than an obvious trick. (Field-map extra: the cost argument below is not made explicitly in this part of the video; he gets to it implicitly in P3–P4.)

The bump method needs one full forward pass per input. Three inputs, three extra evaluations — fine. A small MLP has forty-one parameters, so forty-two evaluations — still fine. A network with a hundred million parameters needs a hundred million forward passes for one gradient, and you need a fresh gradient at every training step. Backpropagation gets all of them in a single backward pass whose cost is about the same as one forward pass, regardless of how many inputs there are. That asymmetry is why the rest of the lecture exists.

The bump method does not become useless, though. It becomes the test. Every time the engine produces a gradient you do not believe, you can re-derive it by nudging and dividing, exactly as in these cells — which is what PyTorch's own gradcheck does to this day, and what Karpathy does by hand at several points later in the video.

The code at the end of this part

import math
import numpy as np
import matplotlib.pyplot as plt
%matplotlib inline

def f(x):
  return 3*x**2 - 4*x + 5

f(3.0)

xs = np.arange(-5, 5, 0.25)
ys = f(xs)
plt.plot(xs, ys)

h = 0.000001
x = 2/3
(f(x + h) - f(x))/h

# les get more complex
a = 2.0
b = -3.0
c = 10.0
d = a*b + c
print(d)

h = 0.0001

# inputs
a = 2.0
b = -3.0
c = 10.0

d1 = a*b + c
c += h
d2 = a*b + c

print('d1', d1)
print('d2', d2)
print('slope', (d2 - d1)/h)

Line by line, with the two places the saved notebook differs from the video:

Where people get stuck

Go deeper, verified

Exercises

  1. Sweep h until it breaks.code — Write slope(x, h) = (f(x+h) - f(x))/h for f(x) = 3x² − 4x + 5, and evaluate it at x = 3 for h in [1e-1, 1e-2, … 1e-16]. Print the absolute error against the exact answer, 14. A good answer: the error falls linearly with h down to about 1e-8, then rises, and is catastrophic by 1e-15. Then add a central-difference version and sweep it too — its error should fall faster (quadratically) and bottom out lower. Say in one sentence why the two curves have a U-shape at all.
  2. Gradient by hand for three inputs.code — Take d(a, b, c) = a*b + c at a = 2, b = −3, c = 10. (1) Write the three partial derivatives symbolically. (2) Write a loop that bumps each input in turn by h = 1e-6 and prints the numerical estimate, so you get a three-element gradient vector. (3) Check that they agree to five decimals. (4) Now change b to 100 and rerun without touching your symbolic answers — predict which entries move before you look. This is the exercise Colab's opening section in miniature; do that section for the real version, which uses a nastier function.
  3. Predict the sign, then run. — Before touching a keyboard, write down the sign of ∂g/∂a and ∂g/∂b for the README expression at a = −4, b = 2, and which of the two you expect to have the larger magnitude. Then pip install micrograd, run the README snippet, and compare with 138.8338 and 645.5773. A good answer explains which operations in the chain amplify b's influence — the b**3 term and the squaring of e are the places to look — rather than just noting that you were right or wrong.
Next: P02 The Value object and the expression graph · Back to the map.