MICROGRAD // FIELD MAP
← field map
PART 05 · BACKPROPAGATION, BY HAND AND THEN AUTOMATED87:05–103:55 · 17 min

Breaking up tanh, more operations, and the same thing in PyTorch

Andrej Karpathy · building micrograd (2022) · part 05 of 8

Transcript: this part, with timestamps

TL;DR — The engine already works, so this stretch is a stress test rather than a new idea. Karpathy replaces the single tanh node with the expression it actually is — (e**(2n) − 1) / (e**(2n) + 1) — which forces him to teach Value five more things: how to accept a plain 2, how to survive 2 * a, how to exponentiate, how to raise to a constant power (division is a special case), and how to subtract. The graph gets much longer and every leaf gradient comes out identical: 0.5, 0, −1.5, 1.0. The point to carry away is that the granularity of an "operation" is a free design choice — anything you can evaluate forward and differentiate locally can be a node. Then he types the same neuron in PyTorch and gets the same four numbers, which is the moment micrograd stops being a toy and starts being a small, honest version of the real thing.

By 87:05 the machine is finished: backward() sorts the graph topologically, each node knows how to push its gradient into its children, and gradients accumulate with += so a value used twice is handled correctly. Nothing in this part is required to make that work. What it does instead is answer the two questions a skeptical viewer has at exactly this moment — did the tanh node quietly do the hard part for me? and is any of this what a real library does? Karpathy answers the first by dissolving tanh into arithmetic and showing the gradients don't move, and the second by rewriting the neuron in PyTorch. Both answers are demonstrations, not new machinery, which is why this is the shortest and least tense stretch of the middle hour.

Outline, with timestamps

Why take tanh apart at all

In P04 tanh got a _backward of one line — self.grad += (1 - t**2) * out.grad — because Karpathy happened to know that identity from calculus. That's legitimate, but it leaves a suspicion: the interesting part of a neuron is the squashing function, and the squashing function was handled by a closed-form derivative someone else derived. If you replace that one node with a pile of +, *, exp and / nodes, each with a derivative a first-year student could produce, and the answer at the leaves is bit-for-bit the same, the suspicion dies.

The identity he uses is the exponential definition of the hyperbolic tangent, tanh(x) = (e^{2x} − 1) / (e^{2x} + 1). Note what it demands: exponentiation, subtraction, addition of a bare 1, and division. Value at this point supports exactly two operations on two Values — + and * — plus tanh. So the detour is really a shopping list, and Karpathy is candid that this is half the motivation: it forces four or five more backward passes to be written, which is the actual practice.

Making Value tolerate a plain number

The first thing that breaks isn't calculus, it's Python. Value(2.0) + 1 raises AttributeError: 'int' object has no attribute 'data', because __add__ immediately does other.data and 1 is an int. The fix at 88:08 is one line at the top of both __add__ and __mul__:

other = other if isinstance(other, Value) else Value(other)

Leave it alone if it's already a Value; otherwise wrap it. The wrapper is a genuine graph node — it appears in _prev, it has a grad field, and gradient flows into it and is then ignored, because nothing holds a reference to it. That's harmless and it keeps the rest of the class uniform: from out = Value(...) onward, other is always a Value.

The second break is subtler and is the one people remember. a * 2 now works; 2 * a does not. Python evaluates a * 2 by calling a.__mul__(2), and evaluates 2 * a by first calling (2).__mul__(a) — the int's method, which knows nothing about Value and returns NotImplemented. Only then does Python try the reflected method on the right-hand operand, a.__rmul__(2). So the fix is to define it, and the definition is a one-liner that flips the order back:

def __rmul__(self, other): # other * self
    return self * other

def __radd__(self, other): # other + self
    return self + other

This is worth pausing on because it is not a micrograd idea at all — it's the Python data model, and every numeric class in the ecosystem (NumPy arrays, Decimal, PyTorch tensors) implements the same pair of hooks for the same reason. Karpathy needs __rmul__ specifically because the expression he is about to write is (2*n).exp(), with the literal on the left.

exp, pow, and division: three ops, three local derivatives

exp() is structurally a clone of tanh(): one child, pop the Python float out with self.data, compute math.exp(x), wrap it back up. The only new thinking is the local derivative, and it's the friendliest one in calculus — d/dx e^x = e^x — which means the derivative is a number the forward pass has already computed and stored. So the closure reads self.grad += out.data * out.grad: local derivative out.data, times the gradient flowing in from above. Karpathy calls this out at 91:13 as looking confusing and being nothing more than the chain rule.

Division he refuses to implement directly. At 92:14 he rewrites a / b as a * b**-1 and implements the more general thing instead: raising a Value to a constant power. The constant matters. __pow__ opens with

assert isinstance(other, (int, float)), "only supporting int/float powers for now"

because the exponent is deliberately not a node in the graph — it appears in the op label as f'**{other}' and nowhere in _children. That restriction is what lets the backward pass be the plain power rule, d/dx xⁿ = n·xⁿ⁻¹, which he goes and looks up on a derivative-rules table rather than assert from memory. Written out with the chain rule attached:

def __pow__(self, other):
    assert isinstance(other, (int, float)), "only supporting int/float powers for now"
    out = Value(self.data**other, (self,), f'**{other}')

    def _backward():
        self.grad += other * (self.data ** (other - 1)) * out.grad
    out._backward = _backward

    return out

Subtraction gets the same treatment at 95:46: -a is a * -1, and a - b is a + (-b). No new _backward is written, because both compose out of __mul__ and __add__, which already know what to do. The nodes really do appear in the drawn graph — a * node with a wrapped −1 hanging off it — which is slightly wasteful and completely correct.

Op addedForwardLocal derivative w.r.t. selfNew _backward?
exp()math.exp(x)out.data (the output itself)yes
__pow__(k)x**k, k an int/floatk * x**(k-1)yes
__truediv__self * other**-1—no, composes
__neg__self * -1—no, composes
__sub__self + (-other)—no, composes
__rmul__, __radd__swap the operands—no, dispatch only

The same neuron, the long way round

The test case is the neuron from P03, unchanged: x1 = 2.0, w1 = −3.0, x2 = 0.0, w2 = 1.0, and the deliberately ugly bias b = 6.8813735870195432 that was chosen so the output lands on a round number. The pre-activation is n = 0.8813735870195432. The only edited line is the last one — o = n.tanh() becomes:

e = (2*n).exp()
o = (e - 1) / (e + 1)

He binds e to a name explicitly because it is used twice, and a value used twice is exactly the case that P04's += fix exists for. If += had still been =, this expression would silently produce a wrong gradient — the accumulation bug and the tanh decomposition are wired together more tightly than the ordering of the video suggests.

Forward, o comes out at 0.7071067811865477, matching the single-node version to the last digit or two of float noise. Backward, the gradient now walks a chain of five nodes to get from o to n where before it took one hop. Here is that walk, with the numbers:

NodeValueHow its grad arrivesgrad
o = num * inv0.70710678seeded by backward()1.0
num = e - 14.82842712inv.data * o.grad0.14644661
inv = (e+1)**-10.14644661num.data * o.grad4.82842712
den = e + 16.82842712-1 * den**-2 * inv.grad−0.10355339
e5.828427121·num.grad + 1·den.grad, accumulated0.04289322
2*n1.76274717e.data * e.grad0.25
n0.881373592 * 0.250.5

That final 0.5 is the whole argument in one number. In P04 it came from 1 − tanh²(n) = 1 − 0.7071² = 0.5, one multiplication. Here it is the product of seven local derivatives, two of which are the + node's boring 1.0s and one of which is the negative −0.1035… coming back through the reciprocal — and they compose to the same number. Everything downstream of n is untouched, so the leaves read out as before: w2.grad = 0, x2.grad = 0.5, w1.grad = 1.0, x1.grad = −1.5. Karpathy checks them on screen at 97:49 and warns that they appear in a different order in the drawing, which is the only place this comparison can trip you.

All that matters is we have some kind of inputs and some kind of an output, and this output is a function of the inputs in some way, and as long as you can do forward pass and the backward pass of that little operation it doesn't matter what that operation is and how composite it is. Karpathy, 98:49
The thing to carry away: an operation is not a mathematical category, it's an implementation decision. A node earns its place in an autograd engine if you can (1) compute its output from its inputs and (2) write down the derivative of its output with respect to each input. Fuse as much as you like into one node — that's what a CUDA kernel for a fused attention block is — or split as far down as you like. The gradients are the same either way; only speed, numerical stability and code volume change. This is exactly why real libraries ship both torch.tanh and the arithmetic to build it yourself.

The same thing in PyTorch

I would like to show you how you can do the exact same thing by using a modern deep neural network library, like for example PyTorch, which I've roughly modeled micrograd by. Karpathy, 99:19

The PyTorch cell is deliberately awkward, and the awkwardness is the lesson. PyTorch's atom is the tensor, an n-dimensional array; micrograd's atom is a single scalar. To make the two comparable, every leaf has to be built as a tensor holding exactly one element, which is why the cell is full of torch.Tensor([2.0]) rather than torch.Tensor(2.0). He points out around 100:21 that this is not how anyone actually writes PyTorch — a normal tensor is a 2×3 block of numbers with a .shape, and the whole point of the library is doing thousands of these operations in parallel.

Two more lines per leaf are pure impedance matching. .double() casts from PyTorch's default float32 to float64, because Python floats are double precision and he wants the digits to line up exactly with micrograd's. And requires_grad = True has to be set explicitly on every leaf, because PyTorch defaults leaves to not tracking gradients — in real training the inputs to the network are data, nobody wants ∂loss/∂pixel, and computing it would be wasted work. Micrograd has no such switch: every Value always has a grad, because at this scale nobody cares.

import torch

x1 = torch.Tensor([2.0]).double()                ; x1.requires_grad = True
x2 = torch.Tensor([0.0]).double()                ; x2.requires_grad = True
w1 = torch.Tensor([-3.0]).double()               ; w1.requires_grad = True
w2 = torch.Tensor([1.0]).double()                ; w2.requires_grad = True
b  = torch.Tensor([6.8813735870195432]).double() ; b.requires_grad = True
n = x1*w1 + x2*w2 + b
o = torch.tanh(n)

print(o.data.item())
o.backward()

print('---')
print('x2', x2.grad.item())
print('w2', w2.grad.item())
print('x1', x1.grad.item())
print('w1', w1.grad.item())

From n = x1*w1 + x2*w2 + b onward the two programs are character-for-character the same. Tensors have .data and .grad just as Value does; o.backward() is spelled identically and does the identical thing. The only extra ceremony is .item(), which strips a one-element tensor down to the bare Python number so it prints as 0.7071067811865476 and not as tensor([0.7071], dtype=torch.float64). Run it and the forward pass agrees, and the four gradients print as 0.5, 0, −1.5, 1.0 — the numbers this lecture has now derived three separate ways.

Karpathy's closing framing at 103:28 is the honest one: what micrograd does is what PyTorch does in the special case where every tensor holds one element. The difference isn't conceptual, it's throughput — plus the several hundred operations, the device backends, the memory management and the kernel fusion that make up the other 99% of a production framework. It's worth noticing what he does not claim: micrograd has no broadcasting, no no_grad context, no graph freeing, no second derivatives, and its backward() recurses in Python, which will hit the recursion limit on a deep enough graph. The API agrees; the engineering does not.

The code at the end of this part

The full Value class as it stands when the PyTorch comparison finishes — the version that carries into P06, where it gets wrapped in Neuron, Layer and MLP:

class Value:

  def __init__(self, data, _children=(), _op='', label=''):
    self.data = data
    self.grad = 0.0
    self._backward = lambda: None
    self._prev = set(_children)
    self._op = _op
    self.label = label

  def __repr__(self):
    return f"Value(data={self.data})"

  def __add__(self, other):
    other = other if isinstance(other, Value) else Value(other)
    out = Value(self.data + other.data, (self, other), '+')

    def _backward():
      self.grad += 1.0 * out.grad
      other.grad += 1.0 * out.grad
    out._backward = _backward

    return out

  def __mul__(self, other):
    other = other if isinstance(other, Value) else Value(other)
    out = Value(self.data * other.data, (self, other), '*')

    def _backward():
      self.grad += other.data * out.grad
      other.grad += self.data * out.grad
    out._backward = _backward

    return out

  def __pow__(self, other):
    assert isinstance(other, (int, float)), "only supporting int/float powers for now"
    out = Value(self.data**other, (self,), f'**{other}')

    def _backward():
        self.grad += other * (self.data ** (other - 1)) * out.grad
    out._backward = _backward

    return out

  def __rmul__(self, other): # other * self
    return self * other

  def __truediv__(self, other): # self / other
    return self * other**-1

  def __neg__(self): # -self
    return self * -1

  def __sub__(self, other): # self - other
    return self + (-other)

  def __radd__(self, other): # other + self
    return self + other

  def tanh(self):
    x = self.data
    t = (math.exp(2*x) - 1)/(math.exp(2*x) + 1)
    out = Value(t, (self, ), 'tanh')

    def _backward():
      self.grad += (1 - t**2) * out.grad
    out._backward = _backward

    return out

  def exp(self):
    x = self.data
    out = Value(math.exp(x), (self, ), 'exp')

    def _backward():
      self.grad += out.data * out.grad # NOTE: in the video I incorrectly used = instead of +=. Fixed here.
    out._backward = _backward

    return out

  def backward(self):

    topo = []
    visited = set()
    def build_topo(v):
      if v not in visited:
        visited.add(v)
        for child in v._prev:
          build_topo(child)
        topo.append(v)
    build_topo(self)

    self.grad = 1.0
    for node in reversed(topo):
      node._backward()

Notes, line by line, against the finished library:

Where people get stuck

Go deeper, verified

Exercises

  1. Rebuild tanh a third waycode — the sigmoid route. Add o = 1 / (1 + (-2*n).exp()) as an intermediate and express tanh(n) = 2·o − 1. Steps: (1) confirm it needs __rsub__ or a rewrite, and pick one; (2) run o.backward() on the P03 neuron; (3) check n.grad is 0.5 and the four leaves are 0, 0.5, 1.0, −1.5; (4) count the nodes in each of the three graphs. A good answer notes that the sigmoid route still costs one exp but trades the division-by-a-sum for a shift and a scale, and that all three forms agree to within about 1e−16.
  2. Add log and build a softmaxcode — the operation this part stops one short of. Implement def log(self) with math.log(self.data) and local derivative 1/self.data, then write softmax(logits) as a list comprehension over exp() divided by the sum. Steps: (1) build a 4-element logit list of Values; (2) take −softmax(logits)[3].log() as a negative-log-likelihood loss; (3) call backward(); (4) rebuild the identical expression in PyTorch with requires_grad=True and compare each gradient to within 1e−9. This is the official exercise Colab's softmax section — it is deliberately reachable the moment you finish this part, and it needs nothing from P06–P08.
  3. Break __pow__ on purpose — remove the assert and pass a Value as the exponent. Steps: (1) predict what goes wrong before running it; (2) run it and see that the forward pass works and the exponent's grad stays 0; (3) derive both partial derivatives of x^y and say which one the current code implements; (4) write the version that adds other to _children and pushes out.data * math.log(self.data) * out.grad into it, and name the input domain where that blows up. A good answer identifies x ≤ 0 as the failure case and explains why the constant-exponent restriction is the sane default for a teaching engine.
Next: P06 An MLP in micrograd: Neuron, Layer, MLP, a tiny dataset, the loss, the parameters · Back to the map.