MICROGRAD // FIELD MAP
← field map
PART 07 · A NEURAL NET, TRAINED BY HAND121:12–136:46 · 16 min

Training by hand: gradient descent, the learning rate, the zero-grad mistake

Andrej Karpathy · building micrograd (2022) · part 07 of 8

Transcript: this part, with timestamps

TL;DR — Everything needed to train the network already exists at 121:12 — a forward pass that builds a graph, a backward() that fills in every .grad, and a parameters() list of all 41 weights and biases. What is missing is one line: p.data += -lr * p.grad. This part earns that minus sign from first principles, runs the loop by hand until the loss falls from 4.84 to essentially zero, and then Karpathy discovers on camera that he never reset the gradients between steps — so they had been accumulating, silently inflating his step size. The fix is three lines. Remember: the gradient tells you which way is uphill, you must walk downhill, and .grad is an accumulator that you are responsible for emptying.

The lecture has spent two hours building machinery and never once used it for its actual purpose. Every previous part was in service of one number: dloss/dp for each parameter p. This part spends that number. It is the shortest conceptual step in the whole video — subtract a small multiple of the gradient from each parameter, repeat — and it is also where the lecture stops being a calculus exercise and starts being machine learning. The three things worth slowing down for are the sign of the update, the size of the step, and the fact that gradients accumulate.

Outline, with timestamps

The minus sign is the whole idea

At 121:12 the screen shows a single parameter — one weight from the first layer — with a .data of about 0.85 and a .grad that is slightly negative. Both numbers came from the previous part: .data from random initialization, .grad from a single call to loss.backward(). The gradient's meaning is fixed and narrow, and it is worth restating precisely because everything downstream depends on it: if you increase this weight by a tiny amount h, the loss changes by approximately grad × h. That is the entire contract. It says nothing about what happens for a large change.

So the weight's gradient is negative. Increase the weight, and the loss goes down. Since the goal is a small loss, the correct move is to increase the weight — which means moving against the sign of the gradient. Written out: p.data += -step_size * p.grad. With a negative grad the two minus signs cancel and the weight goes up. Karpathy first types the update without the minus, reasons out loud that it would push this weight down and therefore push the loss up, and adds the sign at 123:19.

The second framing he gives is the one worth keeping, because it survives into every optimizer you will ever use. Collect all 41 partial derivatives into one vector. That vector points in the direction of steepest increase of the loss. You want a decrease. So you step along the negative of it. "Gradient descent" is not a metaphor — it is literally the instruction "go down the gradient," and the minus sign is where the word descent lives.

QuantityValue at 121:12Why
p.data≈ 0.854random init from random.uniform(-1,1), unchanged so far
p.gradslightly negative (≈ −0.3)filled in by one loss.backward() over the 4-example loss
step size0.01chosen, not derived — the first guess
p.data after≈ 0.8570.854 + (−0.01 × −0.3): up, because raising it lowers the loss

The change is three thousandths. That is the point: the linear approximation the gradient offers is only trustworthy in a small neighbourhood, so the step is deliberately small. Multiply that tiny move across all 41 parameters at once, though, and the loss moves by a lot more than you would expect from any single one of them — because every one of them moved in a helpful direction simultaneously.

One step, then a loop

He nudges all the parameters, re-runs the forward pass — the dataset definition is untouched, only the network changed — and reads the loss. It goes from 4.84 to 4.36 (124:51). Then he does it by hand a few more times, and the sequence he reads aloud is the first evidence in the entire lecture that any of this works:

StepLearning rateLossNote
start—4.84random 41-parameter MLP, four examples
10.014.36first manual update
20.013.90
30.013.66
40.013.47predictions visibly drifting toward ±1
…raised0.31"we may be able to afford to go a bit faster"
…raised0.04predictions now ≈ 1, −1, −1, 1
…too high7e-9after a visible overshoot that could have gone badly

Note what is not happening here. There is no optimizer object, no framework, no magic. Each row of that table is three manual actions: recompute ypred and loss, call loss.backward(), then walk n.parameters() and adjust each .data. Low loss means the four predictions match the four targets [1.0, -1.0, -1.0, 1.0], because that is exactly what the mean-squared-error sum from P6 measures. When the loss reaches 7e-9 the predictions are correct to seven decimal places.

The learning rate is a knife edge, and nobody derives it

At 126:22 he raises the step size to move faster, and then pauses to explain the risk, which is the most important idea in this stretch. The gradient is a local object. It describes the loss surface in an infinitesimal neighbourhood of the current parameters and makes no promise whatsoever about the terrain further out. Take a step small enough and the linear approximation holds and the loss falls. Take a step large enough and you land somewhere the approximation never modelled — possibly a region where the loss is far higher than where you started. Do that repeatedly and training diverges.

Then it happens live. He steps too aggressively, the loss visibly blows up, and — through luck rather than design — the parameters land in a region that then optimizes down to 7e-9. He is candid that this is not a technique:

this learning rate and the tuning of it is a subtle art Andrej Karpathy · 127:54

His practical summary is the one every practitioner still uses: too low and you converge too slowly to be useful; too high and training is unstable and may explode. When he writes the real loop at 129:27 he says 0.01 is a little too small and 0.1 is dangerously high, and picks something in between. (The notebook checked into nn-zero-to-hero ships the loop with -0.1 — the cells were tidied after recording, so treat the notebook's constant as a later edit rather than the number on screen.)

Nothing in backpropagation tells you how far to step. Backprop gives you a direction and a first-order estimate of how fast the loss changes along it; the distance is a hyperparameter you choose, test, and often schedule. The real micrograd demo notebook makes this explicit — its loop sets learning_rate = 1.0 - 0.9*k/100, decaying from 1.0 down to 0.1 across 100 steps, so early steps move fast and late steps refine. See demo.ipynb.

The bug: gradients are an accumulator, and nobody empties it for you

At 130:31 the lecture turns, and this is the most instructive four minutes in the video. Karpathy stops, says the loop has a terrible and very common bug, and — rather than reshoot — leaves it in:

i can't believe i've done it for the 20th time in my life especially on camera Andrej Karpathy · 130:31

The bug traces directly back to the design decision from P4. Every _backward closure in the Value class writes self.grad += …, never =. That += was mandatory: a node used in several places downstream receives a gradient contribution along each path, and those contributions must sum. But += is indiscriminate. It has no notion of "this is a new backward pass, forget the last one." So across training steps, the arithmetic is:

Stepp.grad before backward()p.grad used by the update
10 (set in __init__)g₁
2g₁ (never cleared)g₁ + g₂
3g₁ + g₂g₁ + g₂ + g₃
ksum of all previousg₁ + … + g_k

So the update at step k is not lr × g_k but lr × the running sum of every gradient computed so far. Two things follow. First, the effective learning rate grows without bound as training proceeds. Second, the direction is wrong too — it is a cumulative sum of gradients evaluated at parameter settings the network has long since left. (It superficially resembles momentum, but momentum uses a decaying average with a coefficient below 1, which keeps the running sum bounded. This has no decay at all.)

The fix is to reset the accumulator immediately before each backward pass:

for p in n.parameters():
  p.grad = 0.0
loss.backward()

Ordering matters and is easy to get wrong. Zero after backward() and you wipe the gradients before the update reads them, so nothing learns at all. Zero before, and each backward pass starts from a clean slate and ends holding exactly the derivatives of the current loss. In PyTorch the same operation is optimizer.zero_grad() or model.zero_grad(), and it exists for exactly this reason — PyTorch's .backward() accumulates too. Karpathy later refactors it into micrograd as a Module.zero_grad() base method precisely to mirror nn.Module (nn.py L6–L8).

The subtlest part is what happens after the fix: training gets slower. At 132:33 the corrected loop descends in a visibly more controlled, more gradual way and needs more steps to reach a low loss. The buggy version had, by accident, been taking enormous steps on a problem so easy that enormous steps happened to work. Four examples, 41 parameters, a target the network can fit essentially exactly — there was no ravine to fall into. On any real problem the same bug would have destabilized training or quietly capped how good the model could get, and the code would still have run without error.

This is the transferable lesson of the part, and it is about debugging, not calculus: a neural network that trains is not evidence that the training code is correct. Gradient bugs do not raise exceptions. They degrade results by an amount you cannot see unless you have a correct version to compare against. This is why gradient checking against finite differences, and unit tests against PyTorch (as in test_engine.py), earn their keep.

The summary chapter: what he claims a neural net is

From 134:03 Karpathy compresses the whole lecture into one paragraph, and it is a good paragraph to be able to reproduce from memory. A neural net is a mathematical expression — in the multi-layer-perceptron case a fairly simple one — that takes the data and the parameters as inputs and produces predictions. Attached to it is a loss function that scores those predictions against the targets, arranged so that low loss means the network is doing what you want. Backpropagation through that whole expression yields the gradient of the loss with respect to every parameter. Gradient descent then follows that gradient downhill, repeatedly. That is the entire algorithm.

we just have a blob of neural stuff and we can make it do arbitrary things Andrej Karpathy · 135:04

He then draws the line to the frontier explicitly, and the claim is stronger than it first sounds. This network has 41 parameters; a GPT-class model has hundreds of billions. The learning problem changes — instead of four hand-written examples with ±1 targets, you take a large corpus of internet text and train the network to predict the next token in a sequence — and Karpathy notes that at that scale the resulting networks display genuinely surprising emergent behaviour. But the machinery does not change. The Value abstraction is there (as tensors rather than scalars), the gradient is there, backpropagation is there, gradient descent is there.

He names two honest differences at 136:04. The update rule in practice is not this plain stochastic gradient descent — production training uses variants such as Adam, and things like learning-rate decay, which the micrograd demo notebook already hints at. And the loss for next-token prediction is not mean squared error but cross-entropy. Both are swaps of a component, not of the architecture of the idea. If you want to watch the same skeleton get filled in at the next size up, that is makemore part 1, and eventually let's build GPT — same forward/backward/update loop, bigger expression.

The code at the end of this part

xs = [
  [2.0, 3.0, -1.0],
  [3.0, -1.0, 0.5],
  [0.5, 1.0, 1.0],
  [1.0, 1.0, -1.0],
]
ys = [1.0, -1.0, -1.0, 1.0] # desired targets

for k in range(20):

  # forward pass
  ypred = [n(x) for x in xs]
  loss = sum((yout - ygt)**2 for ygt, yout in zip(ys, ypred))

  # backward pass
  for p in n.parameters():
    p.grad = 0.0
  loss.backward()

  # update
  for p in n.parameters():
    p.data += -0.1 * p.grad

  print(k, loss.data)

Twelve lines, and every one of them is load-bearing:

Where people get stuck

Go deeper, verified

Exercises

  1. Reproduce the bug, then measure itcode — Take the loop above and make two runs from the same seed: one with the zero-grad lines, one without. (a) Set random.seed(1337) and re-create the MLP identically for both. (b) Record loss at every step for 30 steps. (c) In the buggy run, also record the norm of p.grad for one fixed parameter each step. A good answer shows the buggy run's gradient magnitude growing roughly linearly while the correct run's stays bounded, notes that the buggy run reaches a lower loss sooner, and explains why that is an artifact of the problem being trivially fittable rather than a virtue.
  2. Find the learning rate cliffcode — Sweep lr over [0.001, 0.01, 0.05, 0.1, 0.5, 1.0, 2.0], 50 steps each, correct zero-grad, same initialization every time. Plot final loss against lr on a log x-axis. A good answer identifies three regimes — too slow to converge, a working band, and divergence — reports roughly where the boundary falls for this network, and says why that boundary is a property of this loss surface and does not transfer to a different model.
  3. Write zero_grad the way the library doescode — Add a Module base class with zero_grad() and a default parameters() returning [], and make Neuron, Layer and MLP inherit from it, so the loop becomes n.zero_grad(). Then check your version against nn.py L4–L11. A good answer notices that the base zero_grad works for all three classes without overriding, because each one already implements parameters().
  4. The official exercise notebook, section 3 — The exercise Colab from the video description ends by building a softmax and negative-log-likelihood loss on top of Value and checking the gradients against PyTorch. That verification step is this part's discipline applied to a different loss: swap mean squared error for cross-entropy and confirm the same training loop still drives it down.
Next: P08 The real code: micrograd on GitHub, PyTorch's tanh backward, what comes next · Back to the map.