MICROGRAD // FIELD MAP
← field map
PART 08 · THE REAL CODE136:46–145:52 · 9 min

The real code: micrograd on GitHub, PyTorch's tanh backward, what comes next

Andrej Karpathy · building micrograd (2022) · part 08 of 8

Transcript: this part, with timestamps

TL;DR — Karpathy opened the lecture with a promise: by the end you will understand every line of micrograd. This nine-minute coda is him collecting on it. He scrolls the repo top to bottom — engine.py, nn.py, the test that pins micrograd's gradients to PyTorch's, the moons demo — and almost nothing in it is new, which is the point. Then he goes hunting inside PyTorch for the C++ that computes the derivative of tanh, spends fifteen minutes failing, and eventually finds a * (1 - b * b) buried in a kernel file: the same one line you wrote in a notebook, wrapped in dispatch machinery. The thing to remember is the shape of that comparison — real frameworks are not doing something smarter than you did, they are doing the same thing across many dtypes and two devices, and the price of that generality is that the idea becomes hard to find.

Every part before this one built something. This part audits. The lecture has spent two and a quarter hours deriving a scalar autograd engine and a tiny neural net library from nothing, and there is a fair question hanging over it: was that the real thing, or a teaching model of the real thing? Karpathy answers twice. First by showing that the library on PyPI is, line for line, what you just wrote — so the answer is "the real thing, at least at this scale". Then by digging into PyTorch to show that the production version of the one derivative you understand best is also literally what you wrote — so the answer is "the real thing at scale, too". What separates them is not insight; it is dtypes, devices, and a dispatch table.

Outline, with timestamps

engine.py: the same object, with the teaching scaffolding removed

Open engine.py next to the notebook you have been building and the first useful observation is how little there is to reconcile. The constructor stores data and grad, sets _backward to a no-op lambda, freezes the children into a set, and records the operation string for the graph drawing. Every one of those five attributes was introduced somewhere in P2 through P4 and none of them is new here.

The differences are worth cataloguing precisely, because "it's basically the same" is the kind of claim that hides the one place where it is not.

Lecture notebookLibrary engine.pyWhy
label='' on every ValuegoneLabels existed only so draw_dot could print readable node names while you were learning. Nothing computes with them.
tanh(), and exp()relu() onlyThe one substantive change. ReLU is cheaper, is what modern nets actually use, and its derivative is a step function rather than a smooth curve. Karpathy taught with tanh precisely because its local gradient is more interesting to work out.
self.grad += 1.0 * out.grad in __add__self.grad += out.gradIdentical arithmetic. The explicit 1.0 was there to make the local derivative of addition visible; once you believe it, it is noise.
grad initialised to 0.0initialised to 0Cosmetic. Python promotes on the first += with a float.
__radd__, __rmul__, __neg__, __sub__, __truediv__those plus __rsub__ and __rtruediv__The reflected forms you did not happen to need. All are one-liners on top of __add__, __mul__ and __pow__ — the library adds no new backward rules to support them.

That last row is the structurally interesting one. There are exactly four primitive operations in the finished engine — add, multiply, raise to a scalar power, ReLU — and therefore exactly four _backward closures to be right about. Subtraction is a + (-b), negation is a * -1, division is a * b**-1. Every one of those compiles down to nodes whose gradients you already derived. Getting the derivative surface this small is a design decision, and it is why the whole file fits in 94 lines.

def relu(self):
    out = Value(0 if self.data < 0 else self.data, (self,), 'ReLU')

    def _backward():
        self.grad += (out.data > 0) * out.grad
    out._backward = _backward

    return out

The one line to read carefully is (out.data > 0) * out.grad. That comparison produces a Python bool, which multiplies as 1 or 0 — so the gradient passes straight through when the unit was active and is killed when it was not. Note that the test is on out.data, the value after the ReLU, not on self.data; for this function they agree everywhere except exactly at zero, where the derivative does not exist and the code silently picks zero. That is the same pragmatic choice every framework makes, and it is invisible unless you go looking.

nn.py: the one genuinely new idea is Module

nn.py is 60 lines and you have written all but eleven of them. Neuron holds a list of weight Values and a bias; Layer holds neurons; MLP holds layers; parameters() at each level is a flatten of the level below. What is new is that all three now inherit from a base class:

class Module:

    def zero_grad(self):
        for p in self.parameters():
            p.grad = 0

    def parameters(self):
        return []

This is a refactor with a motive. In P7, forgetting to reset gradients before backward() silently turned a training loop into an accumulator and produced a giant effective step size. The fix there was a loop over n.parameters() written by hand at the top of every iteration. Hoisting it into a base class means the fix is available on every object in the library — model.zero_grad() works on a single neuron, a layer, or the whole network — and, more to the point, it is now the same call you would make in PyTorch, where nn.Module also owns zero_grad and parameters. Karpathy is explicit that the API match is deliberate: the code you take away should transfer.

Two small divergences from the lecture's version are worth noticing because they change behaviour rather than style. The library's Neuron.__init__ takes a nonlin=True flag and initialises the bias to exactly Value(0) instead of a random draw. And MLP.__init__ passes nonlin=i != len(nouts) - 1, which makes the final layer linear. Both matter for real training: a zero bias is a standard, less arbitrary starting point than a uniform sample, and a squashed output layer cannot emit the unbounded scores that a margin loss wants. The four-example toy problem in P6 was forgiving enough that neither choice showed up.

The test that makes "we agree with PyTorch" checkable

Two functions in test/test_engine.py, both built the same way: write an expression in micrograd, write the byte-identical expression in PyTorch, call backward() on both, assert the forward values and the input gradients match. It is the single most valuable file in the repo for someone who has just finished this lecture, because it converts a claim into a command you can run.

The details reward a look. test_sanity_check asserts with plain == — exact floating-point equality between a hand-rolled Python engine and a C++ framework. That works because the expression is short and every intermediate is exactly representable in the doubles both sides use, so both compute the identical sequence of IEEE operations. test_more_ops chains far more operations, including division and squaring, so the two implementations' orderings diverge in the last bits; that test drops to a 1e-6 tolerance. The progression from exact to approximate is the honest thing to do and is a good habit to copy.

Note also what the tests use as inputs: a single negative starting value, -4.0, run through ReLUs. That is deliberate — it exercises both sides of the ReLU kink, which is where a naive implementation gets the gradient wrong. The README's worked example uses the same expression and publishes the numbers, so you can check yourself without running anything: g.data is 24.7041, a.grad is 138.8338, b.grad is 645.5773.

demo.ipynb: the three things the lecture left out

The last file in the repo is the one that shows what the toy loop was missing. demo.ipynb trains an MLP(2, [16, 16, 1]) — 337 parameters, against the lecture net's 41 — to separate two interleaved crescents of points, and it produces an actual decision surface at the end. Three additions carry it from demonstration to something recognisably like practice, and Karpathy names all three while explicitly declining to teach them here.

AdditionWhat it isWhy the lecture could skip it
BatchingThe loss takes an optional batch_size and, when given one, forwards a random subset of rows instead of the whole dataset.With four examples, "the batch" and "the dataset" are the same thing. With a million rows a full forward pass per step is unaffordable, and the random subset is a cheap unbiased estimate of the gradient.
A different loss + L2 regularisationAn SVM max-margin hinge, (1 + -yi*scorei).relu() averaged over the batch, plus alpha * sum(p*p) with alpha = 1e-4.Mean squared error was chosen for being the simplest thing that is obviously a Value. The hinge and the penalty term are about generalisation, which never came up because the toy net was asked to memorise four points.
A learning-rate schedulelearning_rate = 1.0 - 0.9*k/100 over 100 steps — 1.0 at the start, 0.109 at the end.P7 showed by hand that a big step converges fast and then overshoots. Decay is the systematic version of the thing Karpathy was doing manually by editing the number between cells.

Everything else in that notebook is the loop you already know: forward, zero_grad, backward, nudge every parameter against its gradient. The regularisation term is a nice illustration of why it was worth making the loss a Value rather than a float — sum(p*p for p in model.parameters()) is an expression over the parameters themselves, so backward() pushes a gradient into every weight through a path that never touched the data at all.

Finding 1 − tanh² inside PyTorch

Then the part of this section that is not a repo tour. Karpathy sets a concrete target: in micrograd the backward for tanh is (1 - t**2) * out.grad, where t is the output of the tanh and the multiplication is the chain rule. Somewhere in PyTorch there must be a line that does that. He goes to find it, on camera, and mostly fails.

these libraries unfortunately they grow in size and entropy … they're meant to be used, not really inspected141:42

He reports about fifteen minutes of searching, 2,800 matches for "tanh" across 406 files, and eventual success only by stumbling across a GitHub issue where someone else had already located the kernels. That is a real experience of a large C++ codebase and it is good that it is in the video unedited. But it is also avoidable, and the disciplined route is worth knowing, because it generalises to every operation in the framework rather than just this one. PyTorch keeps a single declarative table of derivatives, and it is grep-able:

LayerWhereWhat it says
The derivative tabletools/autograd/derivatives.yamltanh's gradient with respect to its input is tanh_backward(grad, result) — note result, the saved output, exactly as in micrograd.
The dispatch stubaten/src/ATen/native/BinaryOps.htanh_backward_stub — one symbol, filled in per device at registration time.
CPU kernelcpu/BinaryOpsKernel.cppThree branches — complex, reduced-precision float, everything else — each computing a * (1 - b*b) scalar-wise and again vectorised.
CUDA kernelcuda/BinaryMiscBackwardOpsKernels.cuThe same rule, in a device lambda, in one line.

Codebase archaeology has a half-life, and this one has decayed since 2022. The CPU kernel is still in the file Karpathy landed on, but the CUDA kernel he showed has since been split out of BinaryMiscOpsKernels.cu into a dedicated backward-ops file. Line numbers here are from main as of September 2026 and will drift again; if an anchor lands in the wrong place, search the file for tanh_backward_kernel and you will be within a few lines of it. The function names have been stable for years even as the files have moved, which is the general lesson: navigate by symbol, cite by line only with a date attached.

# tools/autograd/derivatives.yaml
- name: tanh(Tensor self) -> Tensor
  self: tanh_backward(grad, result)
  result: auto_element_wise

// aten/src/ATen/native/cpu/BinaryOpsKernel.cpp — the plain floating-point branch
AT_DISPATCH_FLOATING_TYPES(iter.dtype(), "tanh_backward_cpu", [&]() {
  auto one_vec = Vectorized<scalar_t>(scalar_t{1});
  cpu_kernel_vec(
      iter,
      [=](scalar_t a, scalar_t b) -> scalar_t {
        return a * (scalar_t{1} - b * b);      // <-- (1 - t**2) * out.grad
      },
      [=](Vectorized<scalar_t> a, Vectorized<scalar_t> b) {
        return a * (one_vec - b * b);          // ...and again, eight lanes at a time
      });
});

// aten/src/ATen/native/cuda/BinaryMiscBackwardOpsKernels.cu — the GPU version
AT_DISPATCH_FLOATING_TYPES_AND2(at::ScalarType::Half, at::ScalarType::BFloat16, dtype, "tanh_backward_cuda", [&]() {
  gpu_kernel(iter, [] GPU_LAMBDA(scalar_t a, scalar_t b) -> scalar_t {
    return a * (scalar_t{1.} - b * b);
  });
});

Read a as out.grad and b as t and the correspondence with your notebook is exact, character for character in the arithmetic. Karpathy notes on camera that the CPU kernel is long and wonders aloud why it lives in a file about binary operations when tanh takes one argument. Both puzzles have the same answer, and it is a good one to hold onto. The length is dtype coverage: a branch for complex numbers, a branch for bfloat16 and half that upcasts to float before doing the arithmetic and casts back, a branch for ordinary floats and doubles — and inside each, both a scalar lambda and a SIMD lambda so the loop can be vectorised. The filename is about arity of the backward function, not the forward one: tanh_backward(grad_output, output) consumes two tensors elementwise, so as far as the kernel machinery is concerned it is a binary elementwise op, and it gets filed with the other binary elementwise ops. Nothing is misplaced; the taxonomy is just one level down from where you were reading.

Registering your own Lego block

The last technical minute of the lecture closes a loop that was opened in P4. Micrograd's design rule was: to add an operation, supply its forward value and a closure that pushes gradient into its inputs. PyTorch's rule is the same rule with a class instead of a closure. Karpathy pulls up the official example, which implements the third Legendre polynomial, ½(5x³ − 3x), as a custom autograd node:

class LegendrePolynomial3(torch.autograd.Function):

    @staticmethod
    def forward(ctx, input):
        ctx.save_for_backward(input)                 # micrograd: capturing self in the closure
        return 0.5 * (5 * input ** 3 - 3 * input)

    @staticmethod
    def backward(ctx, grad_output):
        input, = ctx.saved_tensors                   # micrograd: reading self.data
        return grad_output * 1.5 * (5 * input ** 2 - 1)   # local derivative x upstream grad
you can use this as a lego block in a larger lego castle of all the different lego blocks that PyTorch already has144:15

Line up the three pieces against the engine you built. ctx.save_for_backward is doing what Python's closure did for free when _backward captured self and out; PyTorch needs it explicit because it must be able to free the graph and because the saved tensors participate in its own correctness checks. grad_output is out.grad. The return is the local derivative multiplied by the upstream gradient — the chain rule, written once, by you, per operation. That really is the entire contract: PyTorch does not need to know anything about your function beyond how to evaluate it and how to differentiate it locally, and in exchange it will route gradients through it from anywhere in an arbitrarily large graph. Which is precisely the deal micrograd offers, at a different scale.

Carry this away: the difference between a 94-line autograd engine and a production framework is not conceptual, it is coverage. Same chain rule, same saved output, same local-derivative-times-upstream multiply — plus complex and reduced-precision dtypes, two device backends, vectorised inner loops, a dispatch table, and a declarative derivative registry so the whole thing can be generated rather than hand-written. When you next need to know what a framework really does with some operation, look for its derivative table first; the kernels are an implementation detail of the entry you find there.

The code at the end of this part

Not a notebook cell this time — the finished library, in full. This is the file the lecture's opening promise was about, and after two and a quarter hours there should be nothing in it you cannot account for.

class Value:
    """ stores a single scalar value and its gradient """

    def __init__(self, data, _children=(), _op=''):
        self.data = data
        self.grad = 0
        # internal variables used for autograd graph construction
        self._backward = lambda: None
        self._prev = set(_children)
        self._op = _op # the op that produced this node, for graphviz / debugging / etc

    def __add__(self, other):
        other = other if isinstance(other, Value) else Value(other)
        out = Value(self.data + other.data, (self, other), '+')

        def _backward():
            self.grad += out.grad
            other.grad += out.grad
        out._backward = _backward

        return out

    def __mul__(self, other):
        other = other if isinstance(other, Value) else Value(other)
        out = Value(self.data * other.data, (self, other), '*')

        def _backward():
            self.grad += other.data * out.grad
            other.grad += self.data * out.grad
        out._backward = _backward

        return out

    def __pow__(self, other):
        assert isinstance(other, (int, float)), "only supporting int/float powers for now"
        out = Value(self.data**other, (self,), f'**{other}')

        def _backward():
            self.grad += (other * self.data**(other-1)) * out.grad
        out._backward = _backward

        return out

    def relu(self):
        out = Value(0 if self.data < 0 else self.data, (self,), 'ReLU')

        def _backward():
            self.grad += (out.data > 0) * out.grad
        out._backward = _backward

        return out

    def backward(self):

        # topological order all of the children in the graph
        topo = []
        visited = set()
        def build_topo(v):
            if v not in visited:
                visited.add(v)
                for child in v._prev:
                    build_topo(child)
                topo.append(v)
        build_topo(self)

        # go one variable at a time and apply the chain rule to get its gradient
        self.grad = 1
        for v in reversed(topo):
            v._backward()

    def __neg__(self): # -self
        return self * -1

    def __radd__(self, other): # other + self
        return self + other

    def __sub__(self, other): # self - other
        return self + (-other)

    def __rsub__(self, other): # other - self
        return other + (-self)

    def __rmul__(self, other): # other * self
        return self * other

    def __truediv__(self, other): # self / other
        return self * other**-1

    def __rtruediv__(self, other): # other / self
        return other * self**-1

    def __repr__(self):
        return f"Value(data={self.data}, grad={self.grad})"

Where people get stuck

Go deeper, verified

Exercises

  1. Put tanh back into the real enginecode — clone micrograd and add the operation the lecture used but the library never shipped. (1) Write tanh in engine.py next to relu, in the library's style — compute the value, capture it, and have _backward do self.grad += (1 - t*t) * out.grad. (2) Decide deliberately whether to compute t via math.tanh or the exponential form from P5, and say which is more numerically stable for large |x| and why. (3) Add a third test to test/test_engine.py that runs an expression containing your tanh through both micrograd and torch.tanh, asserting forward and gradient agreement to 1e-6. (4) Give Neuron a nonlin option that selects tanh, and check the moons demo still converges. A good answer notices that the backward reads the saved output, never recomputes the tanh, and works whether t is captured as a local or read from out.data.
  2. Re-run the treasure hunt, todaycode — Karpathy's search took fifteen minutes and ended at a GitHub issue. Do it in five, from the top. (1) Find the tanh entry in derivatives.yaml and note that its gradient is expressed in terms of result, the saved output. (2) Grep for tanh_backward_stub to find the dispatch declaration and both registrations. (3) Open the CPU and CUDA kernels and write down the current file and line for the plain-float branch of each. (4) Compare against the anchors on this page and record what has moved. (5) Then do the same for one operation this lecture never covered — sigmoid is a good second, since its backward sits directly above tanh's in both kernel files. A good answer ends with a repeatable four-step recipe you would trust for any op, and an explanation of why grad_output and the saved output are the only two tensors the kernel needs.
  3. The official exercise notebook — Karpathy sets exactly one exercise for the whole lecture, in the video description: this Colab. Its three sections map back across the map — deriving a gradient analytically and then numerically belongs to P1, extending Value with the operations a softmax and a negative-log-likelihood loss need belongs to P5, and the training questions belong to P7. The section that belongs here is the last one: verifying your gradients against PyTorch's. Do it the way test_engine.py does — write the expression twice, once in your Value and once in tensors with requires_grad=True, and assert on both the forward number and every input gradient rather than eyeballing them. If a gradient disagrees, the fastest bisect is to shrink the expression until it doesn't.

What comes next. The next video in Neural Networks: Zero to Hero is makemore part 1, which drops the scalar engine for PyTorch tensors and trains a bigram character-level language model — the first time the loss becomes a negative log likelihood and the first time the data is something you did not type by hand. Everything about the training loop will be familiar; what changes is that gradients now flow through arrays rather than one number at a time. The series climbs from there to a transformer, and the seventh video, Let's build GPT, is mapped the same way this one is at Let's build GPT, mapped. If you would rather see how far the engine on this page can be pushed instead, Karpathy's microgpt trains a GPT-2-class transformer in dependency-free Python on a descendant of this very Value class — a direct answer to "could you actually build something real out of this?"

And then the outtakes at 145:20, in which the man who just derived backpropagation from first principles cannot reliably say the word "micrograd".

Next: the exercise for the whole lecture · Back to the map.