Breaking up tanh, more operations, and the same thing in PyTorch
Transcript: this part, with timestamps
tanh node with the expression it actually is — (e**(2n) − 1) / (e**(2n) + 1) — which forces him to teach Value five more things: how to accept a plain 2, how to survive 2 * a, how to exponentiate, how to raise to a constant power (division is a special case), and how to subtract. The graph gets much longer and every leaf gradient comes out identical: 0.5, 0, −1.5, 1.0. The point to carry away is that the granularity of an "operation" is a free design choice — anything you can evaluate forward and differentiate locally can be a node. Then he types the same neuron in PyTorch and gets the same four numbers, which is the moment micrograd stops being a toy and starts being a small, honest version of the real thing.By 87:05 the machine is finished: backward() sorts the graph topologically, each node knows how to push its gradient into its children, and gradients accumulate with += so a value used twice is handled correctly. Nothing in this part is required to make that work. What it does instead is answer the two questions a skeptical viewer has at exactly this moment — did the tanh node quietly do the hard part for me? and is any of this what a real library does? Karpathy answers the first by dissolving tanh into arithmetic and showing the gradients don't move, and the second by rewriting the neuron in PyTorch. Both answers are demonstrations, not new machinery, which is why this is the shortest and least tense stretch of the middle hour.
Outline, with timestamps
- 87:05 — Why take tanh apart: he already knows its derivative, so the exercise is really an excuse to implement exp, division and subtraction.
- 87:38 —
a + 1crashes:__add__reaches forother.dataand aninthas none. Wrap non-Values on the way in. - 89:10 —
a * 2works but2 * astill doesn't.__rmul__is Python's fallback, and it just swaps the operands. - 90:41 —
exp(), a second single-input node shaped exactly liketanh; its local derivative is the output itself. - 91:43 — Division, rearranged:
a / bisa * b**−1, so implementing powers gets division for free. - 93:15 —
__pow__with a constant exponent, anassertthat keeps it constant, and the power rule as the local derivative. - 95:46 — Subtraction, built rather than implemented: negate by multiplying by
−1, then add. - 96:17 — The same two-input neuron, with
o = n.tanh()swapped fore = (2*n).exp(); o = (e - 1) / (e + 1). - 97:49 — Reading the long graph: same forward value, and the leaves still carry
0,0.5,1,−1.5. - 98:19 — The punchline: how coarse or fine an operation is, is entirely up to whoever writes the library.
- 99:31 — The neuron in PyTorch: one-element tensors,
.double(),requires_grad = True,o.backward(),.item(). - 103:28 — Where micrograd sits: PyTorch restricted to tensors with exactly one element, minus the parallelism.
Why take tanh apart at all
In P04 tanh got a _backward of one line — self.grad += (1 - t**2) * out.grad — because Karpathy happened to know that identity from calculus. That's legitimate, but it leaves a suspicion: the interesting part of a neuron is the squashing function, and the squashing function was handled by a closed-form derivative someone else derived. If you replace that one node with a pile of +, *, exp and / nodes, each with a derivative a first-year student could produce, and the answer at the leaves is bit-for-bit the same, the suspicion dies.
The identity he uses is the exponential definition of the hyperbolic tangent, tanh(x) = (e^{2x} − 1) / (e^{2x} + 1). Note what it demands: exponentiation, subtraction, addition of a bare 1, and division. Value at this point supports exactly two operations on two Values — + and * — plus tanh. So the detour is really a shopping list, and Karpathy is candid that this is half the motivation: it forces four or five more backward passes to be written, which is the actual practice.
Making Value tolerate a plain number
The first thing that breaks isn't calculus, it's Python. Value(2.0) + 1 raises AttributeError: 'int' object has no attribute 'data', because __add__ immediately does other.data and 1 is an int. The fix at 88:08 is one line at the top of both __add__ and __mul__:
other = other if isinstance(other, Value) else Value(other)
Leave it alone if it's already a Value; otherwise wrap it. The wrapper is a genuine graph node — it appears in _prev, it has a grad field, and gradient flows into it and is then ignored, because nothing holds a reference to it. That's harmless and it keeps the rest of the class uniform: from out = Value(...) onward, other is always a Value.
The second break is subtler and is the one people remember. a * 2 now works; 2 * a does not. Python evaluates a * 2 by calling a.__mul__(2), and evaluates 2 * a by first calling (2).__mul__(a) — the int's method, which knows nothing about Value and returns NotImplemented. Only then does Python try the reflected method on the right-hand operand, a.__rmul__(2). So the fix is to define it, and the definition is a one-liner that flips the order back:
def __rmul__(self, other): # other * self
return self * other
def __radd__(self, other): # other + self
return self + other
This is worth pausing on because it is not a micrograd idea at all — it's the Python data model, and every numeric class in the ecosystem (NumPy arrays, Decimal, PyTorch tensors) implements the same pair of hooks for the same reason. Karpathy needs __rmul__ specifically because the expression he is about to write is (2*n).exp(), with the literal on the left.
exp, pow, and division: three ops, three local derivatives
exp() is structurally a clone of tanh(): one child, pop the Python float out with self.data, compute math.exp(x), wrap it back up. The only new thinking is the local derivative, and it's the friendliest one in calculus — d/dx e^x = e^x — which means the derivative is a number the forward pass has already computed and stored. So the closure reads self.grad += out.data * out.grad: local derivative out.data, times the gradient flowing in from above. Karpathy calls this out at 91:13 as looking confusing and being nothing more than the chain rule.
Division he refuses to implement directly. At 92:14 he rewrites a / b as a * b**-1 and implements the more general thing instead: raising a Value to a constant power. The constant matters. __pow__ opens with
assert isinstance(other, (int, float)), "only supporting int/float powers for now"
because the exponent is deliberately not a node in the graph — it appears in the op label as f'**{other}' and nowhere in _children. That restriction is what lets the backward pass be the plain power rule, d/dx xⁿ = n·xⁿ⁻¹, which he goes and looks up on a derivative-rules table rather than assert from memory. Written out with the chain rule attached:
def __pow__(self, other):
assert isinstance(other, (int, float)), "only supporting int/float powers for now"
out = Value(self.data**other, (self,), f'**{other}')
def _backward():
self.grad += other * (self.data ** (other - 1)) * out.grad
out._backward = _backward
return out
Subtraction gets the same treatment at 95:46: -a is a * -1, and a - b is a + (-b). No new _backward is written, because both compose out of __mul__ and __add__, which already know what to do. The nodes really do appear in the drawn graph — a * node with a wrapped −1 hanging off it — which is slightly wasteful and completely correct.
| Op added | Forward | Local derivative w.r.t. self | New _backward? |
|---|---|---|---|
exp() | math.exp(x) | out.data (the output itself) | yes |
__pow__(k) | x**k, k an int/float | k * x**(k-1) | yes |
__truediv__ | self * other**-1 | — | no, composes |
__neg__ | self * -1 | — | no, composes |
__sub__ | self + (-other) | — | no, composes |
__rmul__, __radd__ | swap the operands | — | no, dispatch only |
The same neuron, the long way round
The test case is the neuron from P03, unchanged: x1 = 2.0, w1 = −3.0, x2 = 0.0, w2 = 1.0, and the deliberately ugly bias b = 6.8813735870195432 that was chosen so the output lands on a round number. The pre-activation is n = 0.8813735870195432. The only edited line is the last one — o = n.tanh() becomes:
e = (2*n).exp()
o = (e - 1) / (e + 1)
He binds e to a name explicitly because it is used twice, and a value used twice is exactly the case that P04's += fix exists for. If += had still been =, this expression would silently produce a wrong gradient — the accumulation bug and the tanh decomposition are wired together more tightly than the ordering of the video suggests.
Forward, o comes out at 0.7071067811865477, matching the single-node version to the last digit or two of float noise. Backward, the gradient now walks a chain of five nodes to get from o to n where before it took one hop. Here is that walk, with the numbers:
| Node | Value | How its grad arrives | grad |
|---|---|---|---|
o = num * inv | 0.70710678 | seeded by backward() | 1.0 |
num = e - 1 | 4.82842712 | inv.data * o.grad | 0.14644661 |
inv = (e+1)**-1 | 0.14644661 | num.data * o.grad | 4.82842712 |
den = e + 1 | 6.82842712 | -1 * den**-2 * inv.grad | −0.10355339 |
e | 5.82842712 | 1·num.grad + 1·den.grad, accumulated | 0.04289322 |
2*n | 1.76274717 | e.data * e.grad | 0.25 |
n | 0.88137359 | 2 * 0.25 | 0.5 |
That final 0.5 is the whole argument in one number. In P04 it came from 1 − tanh²(n) = 1 − 0.7071² = 0.5, one multiplication. Here it is the product of seven local derivatives, two of which are the + node's boring 1.0s and one of which is the negative −0.1035… coming back through the reciprocal — and they compose to the same number. Everything downstream of n is untouched, so the leaves read out as before: w2.grad = 0, x2.grad = 0.5, w1.grad = 1.0, x1.grad = −1.5. Karpathy checks them on screen at 97:49 and warns that they appear in a different order in the drawing, which is the only place this comparison can trip you.
All that matters is we have some kind of inputs and some kind of an output, and this output is a function of the inputs in some way, and as long as you can do forward pass and the backward pass of that little operation it doesn't matter what that operation is and how composite it is. Karpathy, 98:49
torch.tanh and the arithmetic to build it yourself.The same thing in PyTorch
I would like to show you how you can do the exact same thing by using a modern deep neural network library, like for example PyTorch, which I've roughly modeled micrograd by. Karpathy, 99:19
The PyTorch cell is deliberately awkward, and the awkwardness is the lesson. PyTorch's atom is the tensor, an n-dimensional array; micrograd's atom is a single scalar. To make the two comparable, every leaf has to be built as a tensor holding exactly one element, which is why the cell is full of torch.Tensor([2.0]) rather than torch.Tensor(2.0). He points out around 100:21 that this is not how anyone actually writes PyTorch — a normal tensor is a 2×3 block of numbers with a .shape, and the whole point of the library is doing thousands of these operations in parallel.
Two more lines per leaf are pure impedance matching. .double() casts from PyTorch's default float32 to float64, because Python floats are double precision and he wants the digits to line up exactly with micrograd's. And requires_grad = True has to be set explicitly on every leaf, because PyTorch defaults leaves to not tracking gradients — in real training the inputs to the network are data, nobody wants ∂loss/∂pixel, and computing it would be wasted work. Micrograd has no such switch: every Value always has a grad, because at this scale nobody cares.
import torch
x1 = torch.Tensor([2.0]).double() ; x1.requires_grad = True
x2 = torch.Tensor([0.0]).double() ; x2.requires_grad = True
w1 = torch.Tensor([-3.0]).double() ; w1.requires_grad = True
w2 = torch.Tensor([1.0]).double() ; w2.requires_grad = True
b = torch.Tensor([6.8813735870195432]).double() ; b.requires_grad = True
n = x1*w1 + x2*w2 + b
o = torch.tanh(n)
print(o.data.item())
o.backward()
print('---')
print('x2', x2.grad.item())
print('w2', w2.grad.item())
print('x1', x1.grad.item())
print('w1', w1.grad.item())
From n = x1*w1 + x2*w2 + b onward the two programs are character-for-character the same. Tensors have .data and .grad just as Value does; o.backward() is spelled identically and does the identical thing. The only extra ceremony is .item(), which strips a one-element tensor down to the bare Python number so it prints as 0.7071067811865476 and not as tensor([0.7071], dtype=torch.float64). Run it and the forward pass agrees, and the four gradients print as 0.5, 0, −1.5, 1.0 — the numbers this lecture has now derived three separate ways.
Karpathy's closing framing at 103:28 is the honest one: what micrograd does is what PyTorch does in the special case where every tensor holds one element. The difference isn't conceptual, it's throughput — plus the several hundred operations, the device backends, the memory management and the kernel fusion that make up the other 99% of a production framework. It's worth noticing what he does not claim: micrograd has no broadcasting, no no_grad context, no graph freeing, no second derivatives, and its backward() recurses in Python, which will hit the recursion limit on a deep enough graph. The API agrees; the engineering does not.
The code at the end of this part
The full Value class as it stands when the PyTorch comparison finishes — the version that carries into P06, where it gets wrapped in Neuron, Layer and MLP:
class Value:
def __init__(self, data, _children=(), _op='', label=''):
self.data = data
self.grad = 0.0
self._backward = lambda: None
self._prev = set(_children)
self._op = _op
self.label = label
def __repr__(self):
return f"Value(data={self.data})"
def __add__(self, other):
other = other if isinstance(other, Value) else Value(other)
out = Value(self.data + other.data, (self, other), '+')
def _backward():
self.grad += 1.0 * out.grad
other.grad += 1.0 * out.grad
out._backward = _backward
return out
def __mul__(self, other):
other = other if isinstance(other, Value) else Value(other)
out = Value(self.data * other.data, (self, other), '*')
def _backward():
self.grad += other.data * out.grad
other.grad += self.data * out.grad
out._backward = _backward
return out
def __pow__(self, other):
assert isinstance(other, (int, float)), "only supporting int/float powers for now"
out = Value(self.data**other, (self,), f'**{other}')
def _backward():
self.grad += other * (self.data ** (other - 1)) * out.grad
out._backward = _backward
return out
def __rmul__(self, other): # other * self
return self * other
def __truediv__(self, other): # self / other
return self * other**-1
def __neg__(self): # -self
return self * -1
def __sub__(self, other): # self - other
return self + (-other)
def __radd__(self, other): # other + self
return self + other
def tanh(self):
x = self.data
t = (math.exp(2*x) - 1)/(math.exp(2*x) + 1)
out = Value(t, (self, ), 'tanh')
def _backward():
self.grad += (1 - t**2) * out.grad
out._backward = _backward
return out
def exp(self):
x = self.data
out = Value(math.exp(x), (self, ), 'exp')
def _backward():
self.grad += out.data * out.grad # NOTE: in the video I incorrectly used = instead of +=. Fixed here.
out._backward = _backward
return out
def backward(self):
topo = []
visited = set()
def build_topo(v):
if v not in visited:
visited.add(v)
for child in v._prev:
build_topo(child)
topo.append(v)
build_topo(self)
self.grad = 1.0
for node in reversed(topo):
node._backward()
Notes, line by line, against the finished library:
- The wrapping line in
__add__and__mul__survives verbatim into the shipped engine — engine.py#L14 and #L25. The1.0 *in__add__'s backward is pedagogical padding to show the local derivative explicitly; the library drops it (#L17–L19). __pow__is the same function with the same assert in the library at #L35–L43, down to the parenthesisation of(other * self.data**(other-1)) * out.grad. Because the exponent is not a child, no gradient flows to it and none can —_childrenis the one-tuple(self,).- The composed operators —
__neg__,__radd__,__sub__,__rmul__,__truediv__— sit together at #L72–L91. The library adds two the lecture never needs:__rsub__and__rtruediv__, for1 - aand1 / a. expandtanhdo not exist in the shipped library at all.engine.pyhas exactly one nonlinearity,relu(#L45–L52), whose backward is the delightfully bluntself.grad += (out.data > 0) * out.grad— a Python bool multiplying as 1 or 0. That the two nonlinearities are interchangeable without touching anything else is the same abstraction-is-a-choice point, made structurally.- The
expcomment is a real erratum. In the video Karpathy typesself.grad = out.data * out.grad; the notebook fixes it to+=and says so in the comment. It doesn't change any number in the lecture because2*nis only used once — but in a graph where the input to anexpfed two places, it would be P04's bug all over again.
Where people get stuck
- "Why does
a * 2work but2 * anot?" — because Python asks the left operand first.2 * abegins as(2).__mul__(a);intdoesn't know what aValueis, returnsNotImplemented, and Python then tries the reflected methoda.__rmul__(2). If you never define it, you getTypeError: unsupported operand type(s). This has nothing to do with autograd and everything to do with the numeric data model; the same fix appears in every numeric class in Python. - "Why is
exp's backwardout.data * out.gradand notmath.exp(self.data) * out.grad?" — they are the same number.out.dataise^x, computed in the forward pass and cached on the node. Reaching for it is the general pattern, not a trick: an operation's backward pass is allowed to use anything the forward pass already produced, which is exactly why PyTorch keeps activations alive untilbackward()runs and why training uses so much more memory than inference. - "Why does
__pow__refuse aValueexponent?" — because the derivative would be a different formula. With a constantkthe answer is the power rule,k·xᵏ⁻¹. With aValuein the exponent you havex^y, whose derivative with respect toyisx^y·ln x— a second edge out of the node and alogthe engine doesn't have yet. Theassertis a load-bearing scope limit, not defensive noise. - "The long graph and the short graph give the same gradients — is that a coincidence?" — no, it's the chain rule being associative. The gradient at any node is the product of local derivatives along every path to the output, summed over paths. Grouping seven of those factors into one closed-form
1 − tanh²multiplies out to the same number as evaluating them one at a time. What does differ is floating-point rounding (the twoovalues here differ in the last digit) and speed — the fused version does onemath.expwhere the split version builds nine nodes.
Go deeper, verified
- micrograd/engine.py,
__pow__and the composed operators — Andrej Karpathy (2020) · the shipped versions of everything added in this part, in 20 lines; useful for spotting the two reflected operators the lecture skips. - micrograd_lecture_second_half_roughly.ipynb — Andrej Karpathy (2022) · the notebook this part builds; cell 11 is the broken-up tanh neuron and cell 13 is the PyTorch comparison, both runnable as-is.
- Python data model — emulating numeric types — Python Software Foundation · the authoritative statement of why
__rmul__and__radd__exist and exactly when the interpreter reaches for them. Field map extra. - PyTorch — Autograd mechanics — PyTorch docs · what
requires_gradpropagation actually does, why leaves default toFalse, and the graph-freeing behaviour micrograd has no equivalent of. Field map extra. - torch.tanh — PyTorch docs · the fused node micrograd is being compared against; P08 goes and reads its C++ backward implementation. Field map extra.
- CS231n — Backpropagation, intuitions — Andrej Karpathy et al. (2015–) · the written predecessor of this lecture, with the "staged computation" and sigmoid-gate examples that make the same fuse-or-split point at more length. Field map extra.
- Learning representations by back-propagating errors — Rumelhart, Hinton & Williams (1986) · the original; the local-derivative-times-upstream rule being exercised here is their δ recursion, stated for arbitrary differentiable units. Field map extra.
- Derivative rules table — Maths Is Fun · the sort of lookup table Karpathy pulls up on screen at 94:15 to confirm the power rule before typing it. Field map extra.
Exercises
- Rebuild tanh a third waycode — the sigmoid route. Add
o = 1 / (1 + (-2*n).exp())as an intermediate and expresstanh(n) = 2·o − 1. Steps: (1) confirm it needs__rsub__or a rewrite, and pick one; (2) runo.backward()on the P03 neuron; (3) checkn.gradis0.5and the four leaves are0, 0.5, 1.0, −1.5; (4) count the nodes in each of the three graphs. A good answer notes that the sigmoid route still costs oneexpbut trades the division-by-a-sum for a shift and a scale, and that all three forms agree to within about 1e−16. - Add
logand build a softmaxcode — the operation this part stops one short of. Implementdef log(self)withmath.log(self.data)and local derivative1/self.data, then writesoftmax(logits)as a list comprehension overexp()divided by the sum. Steps: (1) build a 4-element logit list ofValues; (2) take−softmax(logits)[3].log()as a negative-log-likelihood loss; (3) callbackward(); (4) rebuild the identical expression in PyTorch withrequires_grad=Trueand compare each gradient to within 1e−9. This is the official exercise Colab's softmax section — it is deliberately reachable the moment you finish this part, and it needs nothing from P06–P08. - Break
__pow__on purpose — remove theassertand pass aValueas the exponent. Steps: (1) predict what goes wrong before running it; (2) run it and see that the forward pass works and the exponent'sgradstays0; (3) derive both partial derivatives ofx^yand say which one the current code implements; (4) write the version that addsotherto_childrenand pushesout.data * math.log(self.data) * out.gradinto it, and name the input domain where that blows up. A good answer identifiesx ≤ 0as the failure case and explains why the constant-exponent restriction is the sane default for a teaching engine.