An MLP in micrograd: Neuron, Layer, MLP, a tiny dataset, the loss, the parameters
Transcript: this part, with timestamps
tanh — three lines of Value arithmetic. A layer is a list of neurons; an MLP is a list of layers called in sequence. Four hand-written examples and a sum of squared errors turn the network's four predictions into one Value, and calling backward() on that one number pushes a gradient into all 41 weights and biases at once. The thing to remember: the loss is not a float that got computed, it is a node at the top of a graph containing four full forward passes, and that is the only reason the weights learn anything.By 103:55 the engine is finished and verified against PyTorch. What is missing is a reason to care: so far every graph has been an expression somebody typed by hand. This stretch of the video builds the smallest thing that deserves the name "neural network" — a two-hidden-layer perceptron — entirely out of Value, and then does the one move that makes training possible: collapsing the network's several outputs into a single scalar so that backward() has somewhere to start. Nothing here is new mathematics. It is the moment the machinery gets pointed at a target.
Outline, with timestamps
- 103:55 — Neural nets are a class of expression: the plan is a two-layer MLP, built to mirror PyTorch's module API.
- 104:28 —
class Neuron:ninrandom weights and a bias, both uniform in [−1, 1]. - 105:28 —
__call__so thatn(x)works; a stub returning 0.0 first, then the real forward pass withzip. - 107:02 —
act.tanh(), and the tidy version:sum(...)started atself.binstead of at 0. - 107:34 —
class Layer: independent neurons, all fed the same input, evaluated in a list comprehension. - 109:06 —
class MLP:sz = [nin] + nouts, consecutive pairs become layers, and the call chains them. - 110:06 — Unwrapping a one-element output list, then
draw_dot(n(x))on the whole network. - 111:04 — Four inputs, four desired targets: a binary classifier's worth of data, typed by hand.
- 112:38 — Mean squared error: pair predictions with targets, subtract, square, sum. Loss ≈ 7.
- 114:39 —
loss.backward(), and the first weight deep inside the net that now has a gradient. - 116:10 —
draw_dot(loss): four forward passes in one picture, ending at 7.12. - 117:56 —
parameters()on all three classes, and the count: 41.
A neuron is four lines of Value arithmetic
The design constraint Karpathy sets himself at 103:55 is the same one that shaped the engine: mirror PyTorch. The engine already matched torch.Tensor on the autograd side — .data, .grad, .backward(). Now the network classes will match torch.nn.Module: a constructor that allocates parameters, a call that runs a forward pass, and a parameters() method that hands you everything trainable. That API choice is the reason this code reads as familiar to anyone who has written PyTorch, and it is worth noticing that it is a choice, not a consequence.
A neuron takes nin inputs. Its constructor allocates one weight per input plus a single bias, all drawn uniformly from −1 to 1. Every one of those is a Value, which means every one of them is a leaf that a gradient can eventually be deposited on. The weights say how much each input matters; the bias is a free offset that shifts the neuron's activation before the squashing function — Karpathy calls it the neuron's
a bias that controls the overall trigger happiness of this neuron104:58
The forward pass is w · x + b put through tanh. He builds it in two moves. First __call__ is stubbed to return 0.0 just to demonstrate that defining __call__ is what makes the expression n(x) legal Python at 105:28 — the object becomes callable, which is the whole trick behind PyTorch's model(x) too. Then the body gets filled in: zip(self.w, x) pairs each weight with its input, a generator multiplies each pair, and sum adds them up.
class Neuron:
def __init__(self, nin):
self.w = [Value(random.uniform(-1,1)) for _ in range(nin)]
self.b = Value(random.uniform(-1,1))
def __call__(self, x):
# w * x + b
act = sum((wi*xi for wi, xi in zip(self.w, x)), self.b)
out = act.tanh()
return out
The detail that trips people up is the second argument to sum. Python's sum takes a start value, defaulting to the integer 0, and the bias is exactly the thing you want the products to accumulate onto — so passing self.b as the start folds the + b into the same call and saves a line (107:34). Two consequences follow. One is syntactic: because sum now has two arguments, the generator expression must be wrapped in its own parentheses. The other is a quiet vindication of the previous part — if you had left the start at 0, the very first addition would be int + Value, which only works because __radd__ was added while breaking up tanh. The dunder chores from P05 are load-bearing here; they are the reason ordinary Python idioms compose with Value at all.
Each call also produces a fresh graph. The weights are persistent objects, but wi*xi, the running sums, the pre-activation and the tanh output are all new Value nodes allocated on every forward pass, each holding references to its children. Nothing is cached, nothing is reused, and the graph is rebuilt from scratch every time you call the network — which is exactly the define-by-run model PyTorch uses, and exactly why the memory cost of a forward pass is proportional to the number of operations, not the number of parameters.
Layer and MLP: containers, nothing more
A layer is a set of neurons that all see the same input and never talk to each other (108:05). That is the entire content of the class: hold a list of Neuron(nin), and on call, evaluate each of them on x and return the list of outputs. "Fully connected" is not implemented anywhere — it is simply what you get when every neuron in the list is constructed with nin weights and handed the whole input vector.
An MLP is the same trick one level up. You give it the input width and a list of layer widths; it forms sz = [nin] + nouts and creates a Layer(sz[i], sz[i+1]) for each consecutive pair, so the widths chain automatically. The forward pass is a loop that reassigns x to each layer's output in turn (109:06). MLP(3, [4, 4, 1]) is therefore three inputs, two hidden layers of four, one output — the picture he has been drawing on the whiteboard all lecture.
One piece of ergonomics matters more than it looks. Layer.__call__ always returns a list, so a network ending in a single neuron would return a one-element list instead of a number. At 110:06 he adds return outs[0] if len(outs) == 1 else outs, and this is not cosmetic: the loss written a minute later does yout - ygt, which needs yout to be a Value, not a list containing one. Small unwrap, large downstream consequence.
Then draw_dot(n(x)) on the whole network, which produces a graph too wide to read. This is deliberate — it is the visual argument that you would never do this by hand, followed immediately by the claim that the machinery does not care how big it gets.
Four examples, one number
The dataset at 111:04 is four three-dimensional inputs and four targets, [1.0, -1.0, -1.0, 1.0], typed as literals. It is a binary classification problem with four training points and forty-one parameters — hopelessly over-parameterised, which is the point: it will fit perfectly and quickly, and nothing about generalisation is being taught here.
Run the untrained network on all four and the predictions are all positive and all clustered high — around 0.91, 0.88, 0.8, 0.8 in his run (111:36). Exact digits depend on the random initialisation, so yours will differ; what will not differ is that two of the four have the wrong sign entirely. The network has no idea what it is doing yet.
| Input | Target | Untrained prediction | Signed error |
|---|---|---|---|
| [2.0, 3.0, −1.0] | +1.0 | ≈ 0.91 | small, right sign |
| [3.0, −1.0, 0.5] | −1.0 | ≈ 0.88 | large, wrong sign |
| [0.5, 1.0, 1.0] | −1.0 | ≈ 0.8 | large, wrong sign |
| [1.0, 1.0, −1.0] | +1.0 | ≈ 0.8 | small, right sign |
Now the move the whole lecture has been building toward. You cannot backpropagate from four numbers; backward() seeds a single output with grad = 1.0 and needs exactly one root. So you define a single number that measures how badly the network is doing overall — the loss — and differentiate that. Mean squared error is the choice here: pair each prediction with its target, subtract, square, and sum the four.
xs = [
[2.0, 3.0, -1.0],
[3.0, -1.0, 0.5],
[0.5, 1.0, 1.0],
[1.0, 1.0, -1.0],
]
ys = [1.0, -1.0, -1.0, 1.0] # desired targets
ypred = [n(x) for x in xs]
loss = sum((yout - ygt)**2 for ygt, yout in zip(ys, ypred))
Three things about that one line. The squaring is there to discard the sign — he says explicitly at 113:38 that absolute value would serve the same purpose, and he is right that for this purpose it would; he does not get into why squaring is preferred in practice (it is differentiable everywhere and penalises large errors quadratically). The loss is zero exactly when every prediction equals its target and grows as you get further away, which is the only property being asked of it. And **2 and - on Value objects exist only because of the operator work in P05 — __pow__ and __sub__. Without them this line raises a TypeError and the lecture stops.
Because the term for the well-fit first example is near zero and the badly-fit ones are near 4 each, the total lands around 7 — he reads off 7.12 at 116:41. And critically, loss is not a Python float. It is a Value sitting at the top of a graph that contains four complete forward passes of a 41-parameter network, wired together by the subtraction, squaring and summation. Four separate runs of the same network, sharing the same weight objects, joined into one expression.
loss.backward() can tell you how every parameter affects overall performance across the whole dataset — and why the += accumulation from P04 is not an optimisation but a correctness requirement.backward() reaches the weights
Then he calls loss.backward(), notes that something magical just happened, and proves it by reaching into the network: n.layers[0].neurons[0].w[0] — the first weight of the first neuron of the first layer, as deep in the structure as it is possible to be from the loss — now has a nonzero .grad (115:09). In his run it is negative, which reads directly: nudging that particular weight up would push the total loss down. There is one such number per parameter, and each is an instruction about which direction to move.
The gradient reached that weight by the same rule as everywhere else in the lecture — local derivative times upstream gradient, accumulated with +=, applied node by node in reverse topological order (backward, engine.py L54). The only difference from P03's hand-worked neuron is scale. Concretely, that first weight receives four contributions, one from each example's forward pass, and they sum: the reported gradient is the derivative of the total loss, not of any single example's error.
Drawing draw_dot(loss) at 116:10 shows the whole thing, and he calls it excessive — which it is, and which is the honest lesson: nothing about this graph is efficient. It is four forward passes of scalar arithmetic with a Python object per intermediate. A real framework batches this into tensor operations. The graph is here for comprehension, not throughput.
He also notices something worth pausing on at 117:12: the input data has gradients too. The 2.0 and 3.0 in the first example each got a derivative, because they are leaves of the graph exactly like the weights are. Those gradients are perfectly real and completely useless for training — the data is given, not adjustable. They are not useless in general, though; the same numbers are what adversarial-example generation and input-optimisation methods use. Here, they are just noise the engine happened to compute. It is a good reminder that the engine has no concept of "parameter": every leaf gets a gradient, and it is the training loop that decides which ones to act on.
parameters(): forty-one numbers in one list
To update the weights you first need to enumerate them, and they are buried three containers deep. Hence parameters() on each class, named after PyTorch's method of the same name (117:56). A neuron returns its weights plus its bias. A layer returns the concatenation of its neurons' parameters. An MLP returns the concatenation of its layers'. He writes the layer version first as an explicit loop with params.extend(...), then compresses it to a nested list comprehension at 119:14 — read the two for clauses left to right, outer loop first, exactly as if they were nested statements.
Then a very ordinary Jupyter trap, which he hits live at 120:45: the network object n was constructed from the old class definitions, so it has no parameters method no matter how many times you re-run the cell that defines the classes. You have to re-instantiate the MLP, which re-randomises every weight and changes all the numbers he has been showing. Worth internalising if you are following along in a notebook.
len(n.parameters()) comes out to 41. The arithmetic:
| Layer | Shape | Neurons | Params each (w + b) | Total |
|---|---|---|---|---|
| 1 (hidden) | 3 → 4 | 4 | 3 + 1 = 4 | 16 |
| 2 (hidden) | 4 → 4 | 4 | 4 + 1 = 5 | 20 |
| 3 (output) | 4 → 1 | 1 | 4 + 1 = 5 | 5 |
| MLP(3, [4, 4, 1]) | 3 → 1 | 9 | — | 41 |
Forty-one Value objects, each with a .data you can change and a .grad telling you which way to change it. That is a complete, trainable neural network, and the part ends one line short of actually training it.
The code at the end of this part
class Neuron:
def __init__(self, nin):
self.w = [Value(random.uniform(-1,1)) for _ in range(nin)]
self.b = Value(random.uniform(-1,1))
def __call__(self, x):
# w * x + b
act = sum((wi*xi for wi, xi in zip(self.w, x)), self.b)
out = act.tanh()
return out
def parameters(self):
return self.w + [self.b]
class Layer:
def __init__(self, nin, nout):
self.neurons = [Neuron(nin) for _ in range(nout)]
def __call__(self, x):
outs = [n(x) for n in self.neurons]
return outs[0] if len(outs) == 1 else outs
def parameters(self):
return [p for neuron in self.neurons for p in neuron.parameters()]
class MLP:
def __init__(self, nin, nouts):
sz = [nin] + nouts
self.layers = [Layer(sz[i], sz[i+1]) for i in range(len(nouts))]
def __call__(self, x):
for layer in self.layers:
x = layer(x)
return x
def parameters(self):
return [p for layer in self.layers for p in layer.parameters()]
x = [2.0, 3.0, -1.0]
n = MLP(3, [4, 4, 1])
n(x)
xs = [
[2.0, 3.0, -1.0],
[3.0, -1.0, 0.5],
[0.5, 1.0, 1.0],
[1.0, 1.0, -1.0],
]
ys = [1.0, -1.0, -1.0, 1.0] # desired targets
ypred = [n(x) for x in xs]
loss = sum((yout - ygt)**2 for ygt, yout in zip(ys, ypred))
loss.backward()
len(n.parameters()) # 41
Line by line, against the finished library:
self.w = [...]; self.b = Value(...)— the same allocation as nn.py L16–L17, with one change in the released version: the bias there is initialised toValue(0)rather than a random number. A zero bias is the standard modern default — the symmetry that random initialisation is there to break lives in the weights, so randomising the bias buys nothing.act = sum((wi*xi ...), self.b)— identical to nn.py L21. Everywi*xiis a new graph node whose_backwardroutesout.gradtowiscaled byxi.dataand vice versa: for a product, the local derivative with respect to one factor is the other factor. Each+in the sum routes its gradient to both children unchanged, because the derivative ofa + bwith respect to either input is 1 — addition is a gradient distributor.out = act.tanh()— the lecture keepstanhas a single node with local derivative1 - t**2. nn.py L20–L22 usesrelu()instead and adds anonlinflag so the output layer can be linear. P05 already showed why either is fine: the node boundary is a choice about where you are willing to write a derivative by hand, not a mathematical constraint.return self.w + [self.b]— nn.py L24–L25, unchanged. Note it returns the objects, not copies: mutatingp.datathrough this list mutates the network. That aliasing is what makes the next part's update loop work.outs[0] if len(outs) == 1 else outs— nn.py L35–L37, unchanged. Convenience, but load-bearing for the loss expression.sz = [nin] + nouts— nn.py L47–L49. The library addsnonlin=i!=len(nouts)-1to that same comprehension, which is the only structural difference between the lecture's MLP and the shipped one.loss = sum((yout - ygt)**2 ...)— the outersumstarts at the integer 0, so its first addition is0 + Valueand goes through__radd__.yout - ygtgoes through__sub__, which is implemented asself + (-other)and therefore as a multiply by −1 plus an add — no new backward function needed.- What is not here yet: a
zero_grad. The library has one on theModulebase class (nn.py L4–L11). The lecture does not, and the resulting bug is the centrepiece of P07.
Where people get stuck
- "Why does
sumtake a second argument, and why the extra parentheses?" — The second argument is Python'sstartvalue for the accumulation, and usingself.bfor it means the running total begins at the bias instead of at 0, folding+ binto the same expression. Oncesumhas two arguments, the generator can no longer be bare; it needs its own parentheses or Python cannot tell where the first argument ends. Leaving the start at 0 also works, but only becauseValue.__radd__exists to handleint + Value. - "My loss line raises a TypeError." — Almost always one of two things. Either your
Valueclass is missing__pow__,__sub__,__neg__or__radd__(they are added in P05, and the loss uses all four), or your network is returning a one-element list because you skipped theouts[0] if len(outs) == 1unwrap, and you are trying to subtract a float from a list. - "Why does the deepest weight get a gradient at all — the loss doesn't mention it?" — Because the loss is not a number that was computed and discarded; it is the root of a graph that still holds references, through
_prev, all the way down to that weight.backward()topologically sorts that graph and walks it in reverse, so every node between the loss and the weight passes its gradient down by the chain rule. The distance is irrelevant; only connectivity matters. And because the same weight object participates in all four forward passes, its four contributions accumulate with+=— remove the+=from P04 and this network silently trains on one example instead of four. - "I added
parameters()and got AttributeError." — Re-running a cell that defines a class does not retrofit the method onto objects already built from the previous definition. You must construct a freshMLP, which re-randomises all 41 parameters and changes every number on screen. Karpathy hits this live; it is not you.
Go deeper, verified
- micrograd/nn.py — Andrej Karpathy (2020) · The whole neural-net library, 60 lines:
Module(L4),Neuron(L13),Layer(L30),MLP(L45). Compare against the cells above to see exactly what the polished version adds —zero_grad, ReLU, a linear output layer,__repr__. - micrograd/engine.py — Andrej Karpathy (2020) · The operators this part leans on:
__pow__at L35,__sub__at L78,__radd__at L75, andbackwardat L54. - torch.nn.Module — PyTorch docs · The API being imitated. Read
parameters(),zero_grad()and__call__/forwardand the resemblance is exact; the difference is that PyTorch discovers parameters by attribute registration rather than by a hand-written list comprehension. - torch.nn.Linear — PyTorch docs · Field map extra. What a whole
Layerbecomes when you batch it: one matrix multiply plus a bias vector, instead ofnoutPython objects each running a scalar loop. Worth reading now to see precisely which line of ourLayerthe GPU version replaces. - torch.nn.MSELoss — PyTorch docs · Field map extra. The same loss with a
reductionargument. Note the default is'mean', not'sum'as in the lecture — which effectively divides every gradient by the number of examples and so changes what learning rate you need. - CS231n: Neural Networks Part 1 — Karpathy, Johnson, Li (Stanford, 2016) · Field map extra. The same neuron-and-layer construction written out with the biological analogy, activation-function comparison, and a parameter count for real architectures. The direct ancestor of this lecture's framing.
- Learning representations by back-propagating errors — Rumelhart, Hinton & Williams (1986) · Field map extra. The original paper, and its loss is exactly the sum of squared differences between output and target that you just typed. Reading section 1 after this part is a short, satisfying trip.
- TensorFlow Playground — Smilkov & Carter (Google, 2016) · Field map extra. Build the same shape of network in a browser and watch the decision boundary move. Useful for intuition about what four points and 41 parameters actually get you.
Exercises
- Predict the parameter count, then check it code — Without running anything, work out
len(MLP(2, [16, 16, 1]).parameters()). Then write a one-line formula forMLP(nin, nouts)in general, implement it as a function, and assert it againstlen(n.parameters())for at least four different shapes including a single-layer net. A good answer names the two things the formula must handle: the bias adds one per neuron, and each layer's input width is the previous layer's output width. Then say how many parametersMLP(3, [4, 4, 1])would have if the biases were dropped, and whether the network could still fit the four-point dataset. - Break the unwrap, then break the accumulation code — (a) Delete the
outs[0] if len(outs) == 1 else outsline fromLayer.__call__and re-run the loss cell; read the error and explain in one sentence which object the interpreter was actually complaining about. Restore it. (b) Now changeself.grad +=toself.grad =in__mul__and__add__only, rebuild the network, and comparen.layers[0].neurons[0].w[0].gradbefore and after. A good answer explains why the broken version returns a gradient that is close to correct for a single example but wrong for the four-example loss, and connects it to theb = a + abug from P04. - Change the loss and predict the consequence code — Replace the sum-of-squares with (i) the mean of squares and (ii) a sum of absolute differences (you will need
abs— implement it as aValuemethod with the right_backward, or express it with the operations you already have). For each, callbackward()on a freshly initialised network and recordw[0].gradof the first neuron. A good answer states the ratio between the mean and sum gradients before running it and gets it right, and says what happens to the absolute-value derivative when a prediction exactly hits its target. Note that the official exercise Colab does this properly in its second section, building a softmax and negative-log-likelihood loss instead — its sections belong to P1, P5 and P7 rather than here, but this part is the setup that makes section 2 make sense.