The real code: micrograd on GitHub, PyTorch's tanh backward, what comes next
Transcript: this part, with timestamps
Every part before this one built something. This part audits. The lecture has spent two and a quarter hours deriving a scalar autograd engine and a tiny neural net library from nothing, and there is a fair question hanging over it: was that the real thing, or a teaching model of the real thing? Karpathy answers twice. First by showing that the library on PyPI is, line for line, what you just wrote — so the answer is "the real thing, at least at this scale". Then by digging into PyTorch to show that the production version of the one derivative you understand best is also literally what you wrote — so the answer is "the real thing at scale, too". What separates them is not insight; it is dtypes, devices, and a dispatch table.
Outline, with timestamps
- 136:46 — Walkthrough begins: the promise from the intro, and the caveat that the repo will keep moving after the recording.
- 137:05 — engine.py: data, grad, _backward, _prev, _op — all recognisable; add, multiply, scalar power, and ReLU where the lecture used tanh.
- 138:07 — nn.py: identical Neuron/Layer/MLP, plus one new thing — a Module base class that mirrors PyTorch's nn.Module and owns zero_grad.
- 138:38 — The test suite: the same expression written twice, once in micrograd and once in PyTorch, with both forward values and input gradients asserted equal.
- 139:09 — demo.ipynb: a real binary classifier on a two-moons dataset, with batching, a max-margin loss, L2 regularisation and a learning-rate schedule — the parts the lecture skipped.
- 141:10 — "Real stuff": what we are looking for in PyTorch is something shaped like (1 - t**2) * out.grad.
- 141:42 — Fifteen minutes of searching, 2,800 hits across 406 files, and a candid aside about what large codebases are optimised for.
- 142:12 — The CPU kernel, found via a GitHub issue; why it is long (complex types, bfloat16, vectorised paths) and where the actual arithmetic hides.
- 143:14 — The CUDA kernel: the same rule, one line, and Karpathy's puzzlement that any of this lives under "binary ops".
- 143:45 — Extending PyTorch: subclass torch.autograd.Function, supply forward and backward, and your operation is a first-class citizen of autograd.
- 144:39 — Conclusion: links in the description, questions in the comments, a follow-up video promised.
- 145:20 — Outtakes: thirty seconds of the takes that did not survive, including "microcrab".
engine.py: the same object, with the teaching scaffolding removed
Open engine.py next to the notebook you have been building and the first useful observation is how little there is to reconcile. The constructor stores data and grad, sets _backward to a no-op lambda, freezes the children into a set, and records the operation string for the graph drawing. Every one of those five attributes was introduced somewhere in P2 through P4 and none of them is new here.
The differences are worth cataloguing precisely, because "it's basically the same" is the kind of claim that hides the one place where it is not.
| Lecture notebook | Library engine.py | Why |
|---|---|---|
| label='' on every Value | gone | Labels existed only so draw_dot could print readable node names while you were learning. Nothing computes with them. |
| tanh(), and exp() | relu() only | The one substantive change. ReLU is cheaper, is what modern nets actually use, and its derivative is a step function rather than a smooth curve. Karpathy taught with tanh precisely because its local gradient is more interesting to work out. |
| self.grad += 1.0 * out.grad in __add__ | self.grad += out.grad | Identical arithmetic. The explicit 1.0 was there to make the local derivative of addition visible; once you believe it, it is noise. |
| grad initialised to 0.0 | initialised to 0 | Cosmetic. Python promotes on the first += with a float. |
| __radd__, __rmul__, __neg__, __sub__, __truediv__ | those plus __rsub__ and __rtruediv__ | The reflected forms you did not happen to need. All are one-liners on top of __add__, __mul__ and __pow__ — the library adds no new backward rules to support them. |
That last row is the structurally interesting one. There are exactly four primitive operations in the finished engine — add, multiply, raise to a scalar power, ReLU — and therefore exactly four _backward closures to be right about. Subtraction is a + (-b), negation is a * -1, division is a * b**-1. Every one of those compiles down to nodes whose gradients you already derived. Getting the derivative surface this small is a design decision, and it is why the whole file fits in 94 lines.
def relu(self):
out = Value(0 if self.data < 0 else self.data, (self,), 'ReLU')
def _backward():
self.grad += (out.data > 0) * out.grad
out._backward = _backward
return out
The one line to read carefully is (out.data > 0) * out.grad. That comparison produces a Python bool, which multiplies as 1 or 0 — so the gradient passes straight through when the unit was active and is killed when it was not. Note that the test is on out.data, the value after the ReLU, not on self.data; for this function they agree everywhere except exactly at zero, where the derivative does not exist and the code silently picks zero. That is the same pragmatic choice every framework makes, and it is invisible unless you go looking.
nn.py: the one genuinely new idea is Module
nn.py is 60 lines and you have written all but eleven of them. Neuron holds a list of weight Values and a bias; Layer holds neurons; MLP holds layers; parameters() at each level is a flatten of the level below. What is new is that all three now inherit from a base class:
class Module:
def zero_grad(self):
for p in self.parameters():
p.grad = 0
def parameters(self):
return []
This is a refactor with a motive. In P7, forgetting to reset gradients before backward() silently turned a training loop into an accumulator and produced a giant effective step size. The fix there was a loop over n.parameters() written by hand at the top of every iteration. Hoisting it into a base class means the fix is available on every object in the library — model.zero_grad() works on a single neuron, a layer, or the whole network — and, more to the point, it is now the same call you would make in PyTorch, where nn.Module also owns zero_grad and parameters. Karpathy is explicit that the API match is deliberate: the code you take away should transfer.
Two small divergences from the lecture's version are worth noticing because they change behaviour rather than style. The library's Neuron.__init__ takes a nonlin=True flag and initialises the bias to exactly Value(0) instead of a random draw. And MLP.__init__ passes nonlin=i != len(nouts) - 1, which makes the final layer linear. Both matter for real training: a zero bias is a standard, less arbitrary starting point than a uniform sample, and a squashed output layer cannot emit the unbounded scores that a margin loss wants. The four-example toy problem in P6 was forgiving enough that neither choice showed up.
The test that makes "we agree with PyTorch" checkable
Two functions in test/test_engine.py, both built the same way: write an expression in micrograd, write the byte-identical expression in PyTorch, call backward() on both, assert the forward values and the input gradients match. It is the single most valuable file in the repo for someone who has just finished this lecture, because it converts a claim into a command you can run.
The details reward a look. test_sanity_check asserts with plain == — exact floating-point equality between a hand-rolled Python engine and a C++ framework. That works because the expression is short and every intermediate is exactly representable in the doubles both sides use, so both compute the identical sequence of IEEE operations. test_more_ops chains far more operations, including division and squaring, so the two implementations' orderings diverge in the last bits; that test drops to a 1e-6 tolerance. The progression from exact to approximate is the honest thing to do and is a good habit to copy.
Note also what the tests use as inputs: a single negative starting value, -4.0, run through ReLUs. That is deliberate — it exercises both sides of the ReLU kink, which is where a naive implementation gets the gradient wrong. The README's worked example uses the same expression and publishes the numbers, so you can check yourself without running anything: g.data is 24.7041, a.grad is 138.8338, b.grad is 645.5773.
demo.ipynb: the three things the lecture left out
The last file in the repo is the one that shows what the toy loop was missing. demo.ipynb trains an MLP(2, [16, 16, 1]) — 337 parameters, against the lecture net's 41 — to separate two interleaved crescents of points, and it produces an actual decision surface at the end. Three additions carry it from demonstration to something recognisably like practice, and Karpathy names all three while explicitly declining to teach them here.
| Addition | What it is | Why the lecture could skip it |
|---|---|---|
| Batching | The loss takes an optional batch_size and, when given one, forwards a random subset of rows instead of the whole dataset. | With four examples, "the batch" and "the dataset" are the same thing. With a million rows a full forward pass per step is unaffordable, and the random subset is a cheap unbiased estimate of the gradient. |
| A different loss + L2 regularisation | An SVM max-margin hinge, (1 + -yi*scorei).relu() averaged over the batch, plus alpha * sum(p*p) with alpha = 1e-4. | Mean squared error was chosen for being the simplest thing that is obviously a Value. The hinge and the penalty term are about generalisation, which never came up because the toy net was asked to memorise four points. |
| A learning-rate schedule | learning_rate = 1.0 - 0.9*k/100 over 100 steps — 1.0 at the start, 0.109 at the end. | P7 showed by hand that a big step converges fast and then overshoots. Decay is the systematic version of the thing Karpathy was doing manually by editing the number between cells. |
Everything else in that notebook is the loop you already know: forward, zero_grad, backward, nudge every parameter against its gradient. The regularisation term is a nice illustration of why it was worth making the loss a Value rather than a float — sum(p*p for p in model.parameters()) is an expression over the parameters themselves, so backward() pushes a gradient into every weight through a path that never touched the data at all.
Finding 1 − tanh² inside PyTorch
Then the part of this section that is not a repo tour. Karpathy sets a concrete target: in micrograd the backward for tanh is (1 - t**2) * out.grad, where t is the output of the tanh and the multiplication is the chain rule. Somewhere in PyTorch there must be a line that does that. He goes to find it, on camera, and mostly fails.
these libraries unfortunately they grow in size and entropy … they're meant to be used, not really inspected141:42
He reports about fifteen minutes of searching, 2,800 matches for "tanh" across 406 files, and eventual success only by stumbling across a GitHub issue where someone else had already located the kernels. That is a real experience of a large C++ codebase and it is good that it is in the video unedited. But it is also avoidable, and the disciplined route is worth knowing, because it generalises to every operation in the framework rather than just this one. PyTorch keeps a single declarative table of derivatives, and it is grep-able:
| Layer | Where | What it says |
|---|---|---|
| The derivative table | tools/autograd/derivatives.yaml | tanh's gradient with respect to its input is tanh_backward(grad, result) — note result, the saved output, exactly as in micrograd. |
| The dispatch stub | aten/src/ATen/native/BinaryOps.h | tanh_backward_stub — one symbol, filled in per device at registration time. |
| CPU kernel | cpu/BinaryOpsKernel.cpp | Three branches — complex, reduced-precision float, everything else — each computing a * (1 - b*b) scalar-wise and again vectorised. |
| CUDA kernel | cuda/BinaryMiscBackwardOpsKernels.cu | The same rule, in a device lambda, in one line. |
Codebase archaeology has a half-life, and this one has decayed since 2022. The CPU kernel is still in the file Karpathy landed on, but the CUDA kernel he showed has since been split out of BinaryMiscOpsKernels.cu into a dedicated backward-ops file. Line numbers here are from main as of September 2026 and will drift again; if an anchor lands in the wrong place, search the file for tanh_backward_kernel and you will be within a few lines of it. The function names have been stable for years even as the files have moved, which is the general lesson: navigate by symbol, cite by line only with a date attached.
# tools/autograd/derivatives.yaml
- name: tanh(Tensor self) -> Tensor
self: tanh_backward(grad, result)
result: auto_element_wise
// aten/src/ATen/native/cpu/BinaryOpsKernel.cpp — the plain floating-point branch
AT_DISPATCH_FLOATING_TYPES(iter.dtype(), "tanh_backward_cpu", [&]() {
auto one_vec = Vectorized<scalar_t>(scalar_t{1});
cpu_kernel_vec(
iter,
[=](scalar_t a, scalar_t b) -> scalar_t {
return a * (scalar_t{1} - b * b); // <-- (1 - t**2) * out.grad
},
[=](Vectorized<scalar_t> a, Vectorized<scalar_t> b) {
return a * (one_vec - b * b); // ...and again, eight lanes at a time
});
});
// aten/src/ATen/native/cuda/BinaryMiscBackwardOpsKernels.cu — the GPU version
AT_DISPATCH_FLOATING_TYPES_AND2(at::ScalarType::Half, at::ScalarType::BFloat16, dtype, "tanh_backward_cuda", [&]() {
gpu_kernel(iter, [] GPU_LAMBDA(scalar_t a, scalar_t b) -> scalar_t {
return a * (scalar_t{1.} - b * b);
});
});
Read a as out.grad and b as t and the correspondence with your notebook is exact, character for character in the arithmetic. Karpathy notes on camera that the CPU kernel is long and wonders aloud why it lives in a file about binary operations when tanh takes one argument. Both puzzles have the same answer, and it is a good one to hold onto. The length is dtype coverage: a branch for complex numbers, a branch for bfloat16 and half that upcasts to float before doing the arithmetic and casts back, a branch for ordinary floats and doubles — and inside each, both a scalar lambda and a SIMD lambda so the loop can be vectorised. The filename is about arity of the backward function, not the forward one: tanh_backward(grad_output, output) consumes two tensors elementwise, so as far as the kernel machinery is concerned it is a binary elementwise op, and it gets filed with the other binary elementwise ops. Nothing is misplaced; the taxonomy is just one level down from where you were reading.
Registering your own Lego block
The last technical minute of the lecture closes a loop that was opened in P4. Micrograd's design rule was: to add an operation, supply its forward value and a closure that pushes gradient into its inputs. PyTorch's rule is the same rule with a class instead of a closure. Karpathy pulls up the official example, which implements the third Legendre polynomial, ½(5x³ − 3x), as a custom autograd node:
class LegendrePolynomial3(torch.autograd.Function):
@staticmethod
def forward(ctx, input):
ctx.save_for_backward(input) # micrograd: capturing self in the closure
return 0.5 * (5 * input ** 3 - 3 * input)
@staticmethod
def backward(ctx, grad_output):
input, = ctx.saved_tensors # micrograd: reading self.data
return grad_output * 1.5 * (5 * input ** 2 - 1) # local derivative x upstream grad
you can use this as a lego block in a larger lego castle of all the different lego blocks that PyTorch already has144:15
Line up the three pieces against the engine you built. ctx.save_for_backward is doing what Python's closure did for free when _backward captured self and out; PyTorch needs it explicit because it must be able to free the graph and because the saved tensors participate in its own correctness checks. grad_output is out.grad. The return is the local derivative multiplied by the upstream gradient — the chain rule, written once, by you, per operation. That really is the entire contract: PyTorch does not need to know anything about your function beyond how to evaluate it and how to differentiate it locally, and in exchange it will route gradients through it from anywhere in an arbitrarily large graph. Which is precisely the deal micrograd offers, at a different scale.
The code at the end of this part
Not a notebook cell this time — the finished library, in full. This is the file the lecture's opening promise was about, and after two and a quarter hours there should be nothing in it you cannot account for.
class Value:
""" stores a single scalar value and its gradient """
def __init__(self, data, _children=(), _op=''):
self.data = data
self.grad = 0
# internal variables used for autograd graph construction
self._backward = lambda: None
self._prev = set(_children)
self._op = _op # the op that produced this node, for graphviz / debugging / etc
def __add__(self, other):
other = other if isinstance(other, Value) else Value(other)
out = Value(self.data + other.data, (self, other), '+')
def _backward():
self.grad += out.grad
other.grad += out.grad
out._backward = _backward
return out
def __mul__(self, other):
other = other if isinstance(other, Value) else Value(other)
out = Value(self.data * other.data, (self, other), '*')
def _backward():
self.grad += other.data * out.grad
other.grad += self.data * out.grad
out._backward = _backward
return out
def __pow__(self, other):
assert isinstance(other, (int, float)), "only supporting int/float powers for now"
out = Value(self.data**other, (self,), f'**{other}')
def _backward():
self.grad += (other * self.data**(other-1)) * out.grad
out._backward = _backward
return out
def relu(self):
out = Value(0 if self.data < 0 else self.data, (self,), 'ReLU')
def _backward():
self.grad += (out.data > 0) * out.grad
out._backward = _backward
return out
def backward(self):
# topological order all of the children in the graph
topo = []
visited = set()
def build_topo(v):
if v not in visited:
visited.add(v)
for child in v._prev:
build_topo(child)
topo.append(v)
build_topo(self)
# go one variable at a time and apply the chain rule to get its gradient
self.grad = 1
for v in reversed(topo):
v._backward()
def __neg__(self): # -self
return self * -1
def __radd__(self, other): # other + self
return self + other
def __sub__(self, other): # self - other
return self + (-other)
def __rsub__(self, other): # other - self
return other + (-self)
def __rmul__(self, other): # other * self
return self * other
def __truediv__(self, other): # self / other
return self * other**-1
def __rtruediv__(self, other): # other / self
return other * self**-1
def __repr__(self):
return f"Value(data={self.data}, grad={self.grad})"
- Lines 5–11, the constructor — five fields. Two are the mathematics (data, grad); two are the graph (_backward, _prev); one is for drawing pictures (_op). The default _backward is a lambda that does nothing, which is exactly right for a leaf — a weight has no children to push gradient into.
- Lines 17–19, the add closure — +=, never =. If a value feeds two consumers, both of their closures write to it and the contributions must sum; assignment would let the last writer win. This is the bug from P4, fixed permanently by a two-character choice.
- Lines 28–30, the multiply closure — each input's local derivative is the other input's data. Note it reads other.data, not other.grad; local derivatives come from the forward pass, which is why the forward has to run first and why its intermediate values must survive.
- Lines 35–43, __pow__ — the assert restricts the exponent to a plain number, which keeps the rule to the schoolbook n·x^(n−1) and avoids needing a gradient with respect to the exponent. It is also what makes division work: a / b is a * b**-1, so one rule buys two operations.
- Lines 54–70, backward — build a topological order by depth-first post-order with a visited set, seed self.grad = 1 because dL/dL is 1, then walk the list in reverse. The ordering is the whole correctness argument: a node's _backward may only run once every consumer downstream of it has already deposited its contribution, and reverse topological order is precisely the guarantee that this has happened.
- Lines 72–91, the operator tail — seven dunder methods, zero new gradient rules. This is where the "only four primitives" design pays off, and it is also what lets 2 * x and 10.0 / f work with a raw Python number on the left.
- The neural net library on top is nn.py: Module at lines 4–11, Neuron at 13–28, Layer at 30–43, MLP at 45–60.
Where people get stuck
- "I installed micrograd and .tanh() doesn't exist." Correct — it never shipped. Karpathy says on camera at 137:05 that tanh is not present "as of right now" but that he intends to add it later; as of this writing, years on, engine.py still exposes only relu. So every notebook you wrote in P3–P7 uses an operation the published library does not have. This is not a version mismatch to debug, it is a genuine divergence between the lecture and the repo, and adding it back is a ten-line exercise (below).
- "Why is tanh's backward in a file called BinaryOpsKernel? tanh takes one argument." Karpathy flags this as odd at 143:14 and moves on. The resolution: the kernel is not for tanh, it is for tanh_backward, whose signature is (grad_output, output) — two tensors in, one out, elementwise. That is a binary elementwise op by the only definition the kernel layer cares about, so it is filed with the others. The forward tanh kernel is where you would expect it — in UnaryOpsKernel.cpp, in the same directory.
- "The library's MLP makes the last layer linear — is that a bug?" No, and it is the change most likely to bite if you copy the lecture's version into a real problem. nonlin=i != len(nouts) - 1 turns off the nonlinearity on the output layer only. A ReLU'd output can never be negative, and a tanh'd output can never leave (−1, 1); either one silently caps what your loss can ask for. The moons demo uses a max-margin loss that wants unbounded signed scores, so the head must be linear. The lecture's toy targets happened to be ±1, which is exactly the range tanh can hit, which is why the issue never surfaced.
- "Your line numbers into PyTorch don't match what I'm seeing." Expect that. The tanh backward code has already moved once since the video — the CUDA kernel Karpathy showed in BinaryMiscOpsKernels.cu now lives in BinaryMiscBackwardOpsKernels.cu — and the CPU file has grown by hundreds of lines. Anchors on this page point at main as of September 2026. The durable handles are the symbol names: tanh_backward in derivatives.yaml, tanh_backward_stub for the dispatch, tanh_backward_kernel / tanh_backward_kernel_cuda for the implementations. Search for those, not for line 974.
Go deeper, verified
- micrograd/engine.py — Andrej Karpathy (2020–) · The whole autograd engine in 94 lines; the file the lecture's opening promise was about. Read lines 54–70 last — the topological sort is the only non-obvious thing in it.
- micrograd/nn.py — Andrej Karpathy · Sixty lines of neural net library, of which the only idea not in the lecture is the Module base class at lines 4–11, deliberately shaped like PyTorch's nn.Module.
- test/test_engine.py — Andrej Karpathy · The claim "we agree with PyTorch" as sixty-seven runnable lines. Worth reading for the exact-equality/tolerance split between the two tests.
- PyTorch: derivatives.yaml, the tanh entry — PyTorch contributors · The front door Karpathy did not use. One declarative table mapping every differentiable operation to its gradient formula; if you ever want to know what a framework really computes for an op, start here rather than in the kernels.
- PyTorch: tanh_backward CPU kernel — PyTorch contributors · field map extra: the line Karpathy hunted for, at its current address, alongside the CUDA kernel, which has moved files since the video was recorded.
- PyTorch: Defining New autograd Functions — PyTorch tutorials · The LegendrePolynomial3 example from 143:45, in full and runnable. Pair it with the Extending PyTorch notes for the rules on saving tensors and the gradient-checking utility.
- Learning representations by back-propagating errors — Rumelhart, Hinton & Williams (1986) · field map extra: four pages, and the algorithm you just implemented. Useful as a check on how little has changed; for a modern treatment of the same material with worked circuit diagrams, the CS231n backprop notes are the standard companion.
Exercises
- Put tanh back into the real enginecode — clone micrograd and add the operation the lecture used but the library never shipped. (1) Write tanh in engine.py next to relu, in the library's style — compute the value, capture it, and have _backward do self.grad += (1 - t*t) * out.grad. (2) Decide deliberately whether to compute t via math.tanh or the exponential form from P5, and say which is more numerically stable for large |x| and why. (3) Add a third test to test/test_engine.py that runs an expression containing your tanh through both micrograd and torch.tanh, asserting forward and gradient agreement to 1e-6. (4) Give Neuron a nonlin option that selects tanh, and check the moons demo still converges. A good answer notices that the backward reads the saved output, never recomputes the tanh, and works whether t is captured as a local or read from out.data.
- Re-run the treasure hunt, todaycode — Karpathy's search took fifteen minutes and ended at a GitHub issue. Do it in five, from the top. (1) Find the tanh entry in derivatives.yaml and note that its gradient is expressed in terms of result, the saved output. (2) Grep for tanh_backward_stub to find the dispatch declaration and both registrations. (3) Open the CPU and CUDA kernels and write down the current file and line for the plain-float branch of each. (4) Compare against the anchors on this page and record what has moved. (5) Then do the same for one operation this lecture never covered — sigmoid is a good second, since its backward sits directly above tanh's in both kernel files. A good answer ends with a repeatable four-step recipe you would trust for any op, and an explanation of why grad_output and the saved output are the only two tensors the kernel needs.
- The official exercise notebook — Karpathy sets exactly one exercise for the whole lecture, in the video description: this Colab. Its three sections map back across the map — deriving a gradient analytically and then numerically belongs to P1, extending Value with the operations a softmax and a negative-log-likelihood loss need belongs to P5, and the training questions belong to P7. The section that belongs here is the last one: verifying your gradients against PyTorch's. Do it the way test_engine.py does — write the expression twice, once in your Value and once in tensors with requires_grad=True, and assert on both the forward number and every input gradient rather than eyeballing them. If a gradient disagrees, the fastest bisect is to shrink the expression until it doesn't.
What comes next. The next video in Neural Networks: Zero to Hero is makemore part 1, which drops the scalar engine for PyTorch tensors and trains a bigram character-level language model — the first time the loss becomes a negative log likelihood and the first time the data is something you did not type by hand. Everything about the training loop will be familiar; what changes is that gradients now flow through arrays rather than one number at a time. The series climbs from there to a transformer, and the seventh video, Let's build GPT, is mapped the same way this one is at Let's build GPT, mapped. If you would rather see how far the engine on this page can be pushed instead, Karpathy's microgpt trains a GPT-2-class transformer in dependency-free Python on a descendant of this very Value class — a direct answer to "could you actually build something real out of this?"
And then the outtakes at 145:20, in which the man who just derived backpropagation from first principles cannot reliably say the word "micrograd".