MICROGRAD // FIELD MAP
← field map
PART 06 · A NEURAL NET, TRAINED BY HAND103:55–121:12 · 17 min

An MLP in micrograd: Neuron, Layer, MLP, a tiny dataset, the loss, the parameters

Andrej Karpathy · building micrograd (2022) · part 06 of 8

Transcript: this part, with timestamps

TL;DR — Everything hard is already done; this part is bookkeeping. A neuron is a dot product of weights with inputs, plus a bias, through tanh — three lines of Value arithmetic. A layer is a list of neurons; an MLP is a list of layers called in sequence. Four hand-written examples and a sum of squared errors turn the network's four predictions into one Value, and calling backward() on that one number pushes a gradient into all 41 weights and biases at once. The thing to remember: the loss is not a float that got computed, it is a node at the top of a graph containing four full forward passes, and that is the only reason the weights learn anything.

By 103:55 the engine is finished and verified against PyTorch. What is missing is a reason to care: so far every graph has been an expression somebody typed by hand. This stretch of the video builds the smallest thing that deserves the name "neural network" — a two-hidden-layer perceptron — entirely out of Value, and then does the one move that makes training possible: collapsing the network's several outputs into a single scalar so that backward() has somewhere to start. Nothing here is new mathematics. It is the moment the machinery gets pointed at a target.

Outline, with timestamps

A neuron is four lines of Value arithmetic

The design constraint Karpathy sets himself at 103:55 is the same one that shaped the engine: mirror PyTorch. The engine already matched torch.Tensor on the autograd side — .data, .grad, .backward(). Now the network classes will match torch.nn.Module: a constructor that allocates parameters, a call that runs a forward pass, and a parameters() method that hands you everything trainable. That API choice is the reason this code reads as familiar to anyone who has written PyTorch, and it is worth noticing that it is a choice, not a consequence.

A neuron takes nin inputs. Its constructor allocates one weight per input plus a single bias, all drawn uniformly from −1 to 1. Every one of those is a Value, which means every one of them is a leaf that a gradient can eventually be deposited on. The weights say how much each input matters; the bias is a free offset that shifts the neuron's activation before the squashing function — Karpathy calls it the neuron's

a bias that controls the overall trigger happiness of this neuron104:58

The forward pass is w · x + b put through tanh. He builds it in two moves. First __call__ is stubbed to return 0.0 just to demonstrate that defining __call__ is what makes the expression n(x) legal Python at 105:28 — the object becomes callable, which is the whole trick behind PyTorch's model(x) too. Then the body gets filled in: zip(self.w, x) pairs each weight with its input, a generator multiplies each pair, and sum adds them up.

class Neuron:

  def __init__(self, nin):
    self.w = [Value(random.uniform(-1,1)) for _ in range(nin)]
    self.b = Value(random.uniform(-1,1))

  def __call__(self, x):
    # w * x + b
    act = sum((wi*xi for wi, xi in zip(self.w, x)), self.b)
    out = act.tanh()
    return out

The detail that trips people up is the second argument to sum. Python's sum takes a start value, defaulting to the integer 0, and the bias is exactly the thing you want the products to accumulate onto — so passing self.b as the start folds the + b into the same call and saves a line (107:34). Two consequences follow. One is syntactic: because sum now has two arguments, the generator expression must be wrapped in its own parentheses. The other is a quiet vindication of the previous part — if you had left the start at 0, the very first addition would be int + Value, which only works because __radd__ was added while breaking up tanh. The dunder chores from P05 are load-bearing here; they are the reason ordinary Python idioms compose with Value at all.

Each call also produces a fresh graph. The weights are persistent objects, but wi*xi, the running sums, the pre-activation and the tanh output are all new Value nodes allocated on every forward pass, each holding references to its children. Nothing is cached, nothing is reused, and the graph is rebuilt from scratch every time you call the network — which is exactly the define-by-run model PyTorch uses, and exactly why the memory cost of a forward pass is proportional to the number of operations, not the number of parameters.

Layer and MLP: containers, nothing more

A layer is a set of neurons that all see the same input and never talk to each other (108:05). That is the entire content of the class: hold a list of Neuron(nin), and on call, evaluate each of them on x and return the list of outputs. "Fully connected" is not implemented anywhere — it is simply what you get when every neuron in the list is constructed with nin weights and handed the whole input vector.

An MLP is the same trick one level up. You give it the input width and a list of layer widths; it forms sz = [nin] + nouts and creates a Layer(sz[i], sz[i+1]) for each consecutive pair, so the widths chain automatically. The forward pass is a loop that reassigns x to each layer's output in turn (109:06). MLP(3, [4, 4, 1]) is therefore three inputs, two hidden layers of four, one output — the picture he has been drawing on the whiteboard all lecture.

One piece of ergonomics matters more than it looks. Layer.__call__ always returns a list, so a network ending in a single neuron would return a one-element list instead of a number. At 110:06 he adds return outs[0] if len(outs) == 1 else outs, and this is not cosmetic: the loss written a minute later does yout - ygt, which needs yout to be a Value, not a list containing one. Small unwrap, large downstream consequence.

Then draw_dot(n(x)) on the whole network, which produces a graph too wide to read. This is deliberate — it is the visual argument that you would never do this by hand, followed immediately by the claim that the machinery does not care how big it gets.

Four examples, one number

The dataset at 111:04 is four three-dimensional inputs and four targets, [1.0, -1.0, -1.0, 1.0], typed as literals. It is a binary classification problem with four training points and forty-one parameters — hopelessly over-parameterised, which is the point: it will fit perfectly and quickly, and nothing about generalisation is being taught here.

Run the untrained network on all four and the predictions are all positive and all clustered high — around 0.91, 0.88, 0.8, 0.8 in his run (111:36). Exact digits depend on the random initialisation, so yours will differ; what will not differ is that two of the four have the wrong sign entirely. The network has no idea what it is doing yet.

InputTargetUntrained predictionSigned error
[2.0, 3.0, −1.0]+1.0≈ 0.91small, right sign
[3.0, −1.0, 0.5]−1.0≈ 0.88large, wrong sign
[0.5, 1.0, 1.0]−1.0≈ 0.8large, wrong sign
[1.0, 1.0, −1.0]+1.0≈ 0.8small, right sign

Now the move the whole lecture has been building toward. You cannot backpropagate from four numbers; backward() seeds a single output with grad = 1.0 and needs exactly one root. So you define a single number that measures how badly the network is doing overall — the loss — and differentiate that. Mean squared error is the choice here: pair each prediction with its target, subtract, square, and sum the four.

xs = [
  [2.0, 3.0, -1.0],
  [3.0, -1.0, 0.5],
  [0.5, 1.0, 1.0],
  [1.0, 1.0, -1.0],
]
ys = [1.0, -1.0, -1.0, 1.0] # desired targets

ypred = [n(x) for x in xs]
loss = sum((yout - ygt)**2 for ygt, yout in zip(ys, ypred))

Three things about that one line. The squaring is there to discard the sign — he says explicitly at 113:38 that absolute value would serve the same purpose, and he is right that for this purpose it would; he does not get into why squaring is preferred in practice (it is differentiable everywhere and penalises large errors quadratically). The loss is zero exactly when every prediction equals its target and grows as you get further away, which is the only property being asked of it. And **2 and - on Value objects exist only because of the operator work in P05 — __pow__ and __sub__. Without them this line raises a TypeError and the lecture stops.

Because the term for the well-fit first example is near zero and the badly-fit ones are near 4 each, the total lands around 7 — he reads off 7.12 at 116:41. And critically, loss is not a Python float. It is a Value sitting at the top of a graph that contains four complete forward passes of a 41-parameter network, wired together by the subtraction, squaring and summation. Four separate runs of the same network, sharing the same weight objects, joined into one expression.

The single most important idea in this part: the loss is a node, not a number. Every operation you performed to get from the weights to the four predictions to the total error is still sitting there as graph structure, and the shared weight objects mean each weight appears in all four forward passes. That is why a single loss.backward() can tell you how every parameter affects overall performance across the whole dataset — and why the += accumulation from P04 is not an optimisation but a correctness requirement.

backward() reaches the weights

Then he calls loss.backward(), notes that something magical just happened, and proves it by reaching into the network: n.layers[0].neurons[0].w[0] — the first weight of the first neuron of the first layer, as deep in the structure as it is possible to be from the loss — now has a nonzero .grad (115:09). In his run it is negative, which reads directly: nudging that particular weight up would push the total loss down. There is one such number per parameter, and each is an instruction about which direction to move.

The gradient reached that weight by the same rule as everywhere else in the lecture — local derivative times upstream gradient, accumulated with +=, applied node by node in reverse topological order (backward, engine.py L54). The only difference from P03's hand-worked neuron is scale. Concretely, that first weight receives four contributions, one from each example's forward pass, and they sum: the reported gradient is the derivative of the total loss, not of any single example's error.

Drawing draw_dot(loss) at 116:10 shows the whole thing, and he calls it excessive — which it is, and which is the honest lesson: nothing about this graph is efficient. It is four forward passes of scalar arithmetic with a Python object per intermediate. A real framework batches this into tensor operations. The graph is here for comprehension, not throughput.

He also notices something worth pausing on at 117:12: the input data has gradients too. The 2.0 and 3.0 in the first example each got a derivative, because they are leaves of the graph exactly like the weights are. Those gradients are perfectly real and completely useless for training — the data is given, not adjustable. They are not useless in general, though; the same numbers are what adversarial-example generation and input-optimisation methods use. Here, they are just noise the engine happened to compute. It is a good reminder that the engine has no concept of "parameter": every leaf gets a gradient, and it is the training loop that decides which ones to act on.

parameters(): forty-one numbers in one list

To update the weights you first need to enumerate them, and they are buried three containers deep. Hence parameters() on each class, named after PyTorch's method of the same name (117:56). A neuron returns its weights plus its bias. A layer returns the concatenation of its neurons' parameters. An MLP returns the concatenation of its layers'. He writes the layer version first as an explicit loop with params.extend(...), then compresses it to a nested list comprehension at 119:14 — read the two for clauses left to right, outer loop first, exactly as if they were nested statements.

Then a very ordinary Jupyter trap, which he hits live at 120:45: the network object n was constructed from the old class definitions, so it has no parameters method no matter how many times you re-run the cell that defines the classes. You have to re-instantiate the MLP, which re-randomises every weight and changes all the numbers he has been showing. Worth internalising if you are following along in a notebook.

len(n.parameters()) comes out to 41. The arithmetic:

LayerShapeNeuronsParams each (w + b)Total
1 (hidden)3 → 443 + 1 = 416
2 (hidden)4 → 444 + 1 = 520
3 (output)4 → 114 + 1 = 55
MLP(3, [4, 4, 1])3 → 19—41

Forty-one Value objects, each with a .data you can change and a .grad telling you which way to change it. That is a complete, trainable neural network, and the part ends one line short of actually training it.

The code at the end of this part

class Neuron:

  def __init__(self, nin):
    self.w = [Value(random.uniform(-1,1)) for _ in range(nin)]
    self.b = Value(random.uniform(-1,1))

  def __call__(self, x):
    # w * x + b
    act = sum((wi*xi for wi, xi in zip(self.w, x)), self.b)
    out = act.tanh()
    return out

  def parameters(self):
    return self.w + [self.b]

class Layer:

  def __init__(self, nin, nout):
    self.neurons = [Neuron(nin) for _ in range(nout)]

  def __call__(self, x):
    outs = [n(x) for n in self.neurons]
    return outs[0] if len(outs) == 1 else outs

  def parameters(self):
    return [p for neuron in self.neurons for p in neuron.parameters()]

class MLP:

  def __init__(self, nin, nouts):
    sz = [nin] + nouts
    self.layers = [Layer(sz[i], sz[i+1]) for i in range(len(nouts))]

  def __call__(self, x):
    for layer in self.layers:
      x = layer(x)
    return x

  def parameters(self):
    return [p for layer in self.layers for p in layer.parameters()]

x = [2.0, 3.0, -1.0]
n = MLP(3, [4, 4, 1])
n(x)

xs = [
  [2.0, 3.0, -1.0],
  [3.0, -1.0, 0.5],
  [0.5, 1.0, 1.0],
  [1.0, 1.0, -1.0],
]
ys = [1.0, -1.0, -1.0, 1.0] # desired targets

ypred = [n(x) for x in xs]
loss = sum((yout - ygt)**2 for ygt, yout in zip(ys, ypred))
loss.backward()

len(n.parameters())   # 41

Line by line, against the finished library:

Where people get stuck

Go deeper, verified

Exercises

  1. Predict the parameter count, then check it code — Without running anything, work out len(MLP(2, [16, 16, 1]).parameters()). Then write a one-line formula for MLP(nin, nouts) in general, implement it as a function, and assert it against len(n.parameters()) for at least four different shapes including a single-layer net. A good answer names the two things the formula must handle: the bias adds one per neuron, and each layer's input width is the previous layer's output width. Then say how many parameters MLP(3, [4, 4, 1]) would have if the biases were dropped, and whether the network could still fit the four-point dataset.
  2. Break the unwrap, then break the accumulation code — (a) Delete the outs[0] if len(outs) == 1 else outs line from Layer.__call__ and re-run the loss cell; read the error and explain in one sentence which object the interpreter was actually complaining about. Restore it. (b) Now change self.grad += to self.grad = in __mul__ and __add__ only, rebuild the network, and compare n.layers[0].neurons[0].w[0].grad before and after. A good answer explains why the broken version returns a gradient that is close to correct for a single example but wrong for the four-example loss, and connects it to the b = a + a bug from P04.
  3. Change the loss and predict the consequence code — Replace the sum-of-squares with (i) the mean of squares and (ii) a sum of absolute differences (you will need abs — implement it as a Value method with the right _backward, or express it with the operations you already have). For each, call backward() on a freshly initialised network and record w[0].grad of the first neuron. A good answer states the ratio between the mean and sum gradients before running it and gets it right, and says what happens to the absolute-value derivative when a prediction exactly hits its target. Note that the official exercise Colab does this properly in its second section, building a softmax and negative-log-likelihood loss instead — its sections belong to P1, P5 and P7 rather than here, but this part is the setup that makes section 2 make sense.
Next: P07 Training by hand: gradient descent, the learning rate, the zero-grad mistake · Back to the map.