Training by hand: gradient descent, the learning rate, the zero-grad mistake
Transcript: this part, with timestamps
The lecture has spent two hours building machinery and never once used it for its actual purpose. Every previous part was in service of one number: dloss/dp for each parameter p. This part spends that number. It is the shortest conceptual step in the whole video — subtract a small multiple of the gradient from each parameter, repeat — and it is also where the lecture stops being a calculus exercise and starts being machine learning. The three things worth slowing down for are the sign of the update, the size of the step, and the fact that gradients accumulate.
Outline, with timestamps
- 121:12 — The parameter on screen: one weight with .data ≈ 0.85 and a slightly negative .grad, and the question of what to do with it.
- 122:15 — Gradient descent stated: loop over all 41 parameters, nudge each .data by a small step in the direction of its gradient.
- 122:47 — The sign argument, worked out loud on the negative-gradient weight; the missing minus.
- 123:49 — The nudge lands: 0.854 → 0.857 on that weight, which is the direction the gradient asked for.
- 124:51 — Re-run the forward pass and read the new loss: 4.84 → 4.36.
- 125:22 — Doing it again, and again, by hand: 3.90, 3.66, 3.47, and the predictions visibly moving toward their targets.
- 126:22 — A bigger learning rate; the loss drops to 0.31 then 0.04; what "overstepping" means when you only know the local slope.
- 127:54 — After an overshoot the loss lands at 7e-9 anyway; the learning rate as a tuned quantity, not a derived one.
- 128:25 — Writing an actual training loop: forward, backward, update, print, for 20 steps.
- 130:31 — "A really terrible bug": the gradients were never zeroed between steps.
- 132:02 — Zeroing p.grad before backward(), and the corrected — noticeably slower — descent.
- 133:03 — Why the buggy version converged faster, and why that is not luck you can count on.
- 134:03 — The summary chapter: what a neural net is, and the straight line from these 41 parameters to GPT.
The minus sign is the whole idea
At 121:12 the screen shows a single parameter — one weight from the first layer — with a .data of about 0.85 and a .grad that is slightly negative. Both numbers came from the previous part: .data from random initialization, .grad from a single call to loss.backward(). The gradient's meaning is fixed and narrow, and it is worth restating precisely because everything downstream depends on it: if you increase this weight by a tiny amount h, the loss changes by approximately grad × h. That is the entire contract. It says nothing about what happens for a large change.
So the weight's gradient is negative. Increase the weight, and the loss goes down. Since the goal is a small loss, the correct move is to increase the weight — which means moving against the sign of the gradient. Written out: p.data += -step_size * p.grad. With a negative grad the two minus signs cancel and the weight goes up. Karpathy first types the update without the minus, reasons out loud that it would push this weight down and therefore push the loss up, and adds the sign at 123:19.
The second framing he gives is the one worth keeping, because it survives into every optimizer you will ever use. Collect all 41 partial derivatives into one vector. That vector points in the direction of steepest increase of the loss. You want a decrease. So you step along the negative of it. "Gradient descent" is not a metaphor — it is literally the instruction "go down the gradient," and the minus sign is where the word descent lives.
| Quantity | Value at 121:12 | Why |
|---|---|---|
| p.data | ≈ 0.854 | random init from random.uniform(-1,1), unchanged so far |
| p.grad | slightly negative (≈ −0.3) | filled in by one loss.backward() over the 4-example loss |
| step size | 0.01 | chosen, not derived — the first guess |
| p.data after | ≈ 0.857 | 0.854 + (−0.01 × −0.3): up, because raising it lowers the loss |
The change is three thousandths. That is the point: the linear approximation the gradient offers is only trustworthy in a small neighbourhood, so the step is deliberately small. Multiply that tiny move across all 41 parameters at once, though, and the loss moves by a lot more than you would expect from any single one of them — because every one of them moved in a helpful direction simultaneously.
One step, then a loop
He nudges all the parameters, re-runs the forward pass — the dataset definition is untouched, only the network changed — and reads the loss. It goes from 4.84 to 4.36 (124:51). Then he does it by hand a few more times, and the sequence he reads aloud is the first evidence in the entire lecture that any of this works:
| Step | Learning rate | Loss | Note |
|---|---|---|---|
| start | — | 4.84 | random 41-parameter MLP, four examples |
| 1 | 0.01 | 4.36 | first manual update |
| 2 | 0.01 | 3.90 | |
| 3 | 0.01 | 3.66 | |
| 4 | 0.01 | 3.47 | predictions visibly drifting toward ±1 |
| … | raised | 0.31 | "we may be able to afford to go a bit faster" |
| … | raised | 0.04 | predictions now ≈ 1, −1, −1, 1 |
| … | too high | 7e-9 | after a visible overshoot that could have gone badly |
Note what is not happening here. There is no optimizer object, no framework, no magic. Each row of that table is three manual actions: recompute ypred and loss, call loss.backward(), then walk n.parameters() and adjust each .data. Low loss means the four predictions match the four targets [1.0, -1.0, -1.0, 1.0], because that is exactly what the mean-squared-error sum from P6 measures. When the loss reaches 7e-9 the predictions are correct to seven decimal places.
The learning rate is a knife edge, and nobody derives it
At 126:22 he raises the step size to move faster, and then pauses to explain the risk, which is the most important idea in this stretch. The gradient is a local object. It describes the loss surface in an infinitesimal neighbourhood of the current parameters and makes no promise whatsoever about the terrain further out. Take a step small enough and the linear approximation holds and the loss falls. Take a step large enough and you land somewhere the approximation never modelled — possibly a region where the loss is far higher than where you started. Do that repeatedly and training diverges.
Then it happens live. He steps too aggressively, the loss visibly blows up, and — through luck rather than design — the parameters land in a region that then optimizes down to 7e-9. He is candid that this is not a technique:
this learning rate and the tuning of it is a subtle art Andrej Karpathy · 127:54
His practical summary is the one every practitioner still uses: too low and you converge too slowly to be useful; too high and training is unstable and may explode. When he writes the real loop at 129:27 he says 0.01 is a little too small and 0.1 is dangerously high, and picks something in between. (The notebook checked into nn-zero-to-hero ships the loop with -0.1 — the cells were tidied after recording, so treat the notebook's constant as a later edit rather than the number on screen.)
The bug: gradients are an accumulator, and nobody empties it for you
At 130:31 the lecture turns, and this is the most instructive four minutes in the video. Karpathy stops, says the loop has a terrible and very common bug, and — rather than reshoot — leaves it in:
i can't believe i've done it for the 20th time in my life especially on camera Andrej Karpathy · 130:31
The bug traces directly back to the design decision from P4. Every _backward closure in the Value class writes self.grad += …, never =. That += was mandatory: a node used in several places downstream receives a gradient contribution along each path, and those contributions must sum. But += is indiscriminate. It has no notion of "this is a new backward pass, forget the last one." So across training steps, the arithmetic is:
| Step | p.grad before backward() | p.grad used by the update |
|---|---|---|
| 1 | 0 (set in __init__) | g₁ |
| 2 | g₁ (never cleared) | g₁ + g₂ |
| 3 | g₁ + g₂ | g₁ + g₂ + g₃ |
| k | sum of all previous | g₁ + … + g_k |
So the update at step k is not lr × g_k but lr × the running sum of every gradient computed so far. Two things follow. First, the effective learning rate grows without bound as training proceeds. Second, the direction is wrong too — it is a cumulative sum of gradients evaluated at parameter settings the network has long since left. (It superficially resembles momentum, but momentum uses a decaying average with a coefficient below 1, which keeps the running sum bounded. This has no decay at all.)
The fix is to reset the accumulator immediately before each backward pass:
for p in n.parameters():
p.grad = 0.0
loss.backward()
Ordering matters and is easy to get wrong. Zero after backward() and you wipe the gradients before the update reads them, so nothing learns at all. Zero before, and each backward pass starts from a clean slate and ends holding exactly the derivatives of the current loss. In PyTorch the same operation is optimizer.zero_grad() or model.zero_grad(), and it exists for exactly this reason — PyTorch's .backward() accumulates too. Karpathy later refactors it into micrograd as a Module.zero_grad() base method precisely to mirror nn.Module (nn.py L6–L8).
The subtlest part is what happens after the fix: training gets slower. At 132:33 the corrected loop descends in a visibly more controlled, more gradual way and needs more steps to reach a low loss. The buggy version had, by accident, been taking enormous steps on a problem so easy that enormous steps happened to work. Four examples, 41 parameters, a target the network can fit essentially exactly — there was no ravine to fall into. On any real problem the same bug would have destabilized training or quietly capped how good the model could get, and the code would still have run without error.
The summary chapter: what he claims a neural net is
From 134:03 Karpathy compresses the whole lecture into one paragraph, and it is a good paragraph to be able to reproduce from memory. A neural net is a mathematical expression — in the multi-layer-perceptron case a fairly simple one — that takes the data and the parameters as inputs and produces predictions. Attached to it is a loss function that scores those predictions against the targets, arranged so that low loss means the network is doing what you want. Backpropagation through that whole expression yields the gradient of the loss with respect to every parameter. Gradient descent then follows that gradient downhill, repeatedly. That is the entire algorithm.
we just have a blob of neural stuff and we can make it do arbitrary things Andrej Karpathy · 135:04
He then draws the line to the frontier explicitly, and the claim is stronger than it first sounds. This network has 41 parameters; a GPT-class model has hundreds of billions. The learning problem changes — instead of four hand-written examples with ±1 targets, you take a large corpus of internet text and train the network to predict the next token in a sequence — and Karpathy notes that at that scale the resulting networks display genuinely surprising emergent behaviour. But the machinery does not change. The Value abstraction is there (as tensors rather than scalars), the gradient is there, backpropagation is there, gradient descent is there.
He names two honest differences at 136:04. The update rule in practice is not this plain stochastic gradient descent — production training uses variants such as Adam, and things like learning-rate decay, which the micrograd demo notebook already hints at. And the loss for next-token prediction is not mean squared error but cross-entropy. Both are swaps of a component, not of the architecture of the idea. If you want to watch the same skeleton get filled in at the next size up, that is makemore part 1, and eventually let's build GPT — same forward/backward/update loop, bigger expression.
The code at the end of this part
xs = [
[2.0, 3.0, -1.0],
[3.0, -1.0, 0.5],
[0.5, 1.0, 1.0],
[1.0, 1.0, -1.0],
]
ys = [1.0, -1.0, -1.0, 1.0] # desired targets
for k in range(20):
# forward pass
ypred = [n(x) for x in xs]
loss = sum((yout - ygt)**2 for ygt, yout in zip(ys, ypred))
# backward pass
for p in n.parameters():
p.grad = 0.0
loss.backward()
# update
for p in n.parameters():
p.data += -0.1 * p.grad
print(k, loss.data)
Twelve lines, and every one of them is load-bearing:
- ypred = [n(x) for x in xs] — the forward pass. Calling the MLP on a plain list of Python floats builds a fresh expression graph of Value nodes, four of them (one per example), each rooted in the same 41 parameter Values. The parameters persist across iterations; the graph is rebuilt from scratch every step. That is why nothing needs to be "reset" except the gradients.
- loss = sum((yout - ygt)**2 …) — mean squared error without the mean: a sum of four squared errors. It is itself a Value, and it is the single root that the four sub-graphs feed into, which is what makes one backward() call sufficient. - and **2 exist only because of the operator work in P5 (__pow__ at engine.py L35, __sub__ at L78).
- for p in n.parameters(): p.grad = 0.0 — the fix from 132:02. It reaches every weight and bias through the nested list comprehensions of MLP.parameters() → Layer.parameters() → Neuron.parameters(). In the finished library this is hoisted to Module.zero_grad, nn.py L6. Note it only zeroes the parameters, not every node in the graph — the intermediate nodes are new objects each iteration and start at grad = 0 from the constructor anyway.
- loss.backward() — one call, and every one of the 41 parameters ends up holding dloss/dp. Internally it topologically sorts the graph, seeds loss.grad = 1.0, and walks the nodes in reverse calling each _backward closure: engine.py L54–L70. The topological order is what guarantees a node's own .grad is complete before it is used to push gradient to its children.
- p.data += -0.1 * p.grad — the update, and the only new idea in this part. It touches .data only. It builds no graph and creates no Values; it is ordinary float arithmetic that mutates the leaves the next forward pass will read.
- print(k, loss.data) — the whole diagnostic apparatus. If this number is not going down, something above it is wrong.
Where people get stuck
- "Why minus? The gradient is what I want, isn't it?" — The gradient answers "which way makes the loss go up fastest." You want the loss to go down, so you negate. Sanity check it on one parameter with a negative gradient, as Karpathy does: negative gradient means increasing this weight decreases the loss, so a correct update must increase it, and -lr × (negative) is positive. If your loss climbs steadily from step one, a flipped sign is the first thing to check — you have implemented gradient ascent.
- "If += was the correct fix in P4, why do I now have to zero the gradients?" — These are the same fact seen from two sides. Within one backward pass, accumulation is required, because a node reached by several downstream paths must sum its contributions. Between passes, accumulation is wrong, because the contract of .grad is "derivative of the current loss." Value has no way to know where one pass ends and another begins, so the caller must say so, by writing zeros. Every autograd framework makes the same trade and hands you the same responsibility.
- "My loss got worse after I fixed the zero-grad bug — did I break it?" — Almost certainly not; expect this. The buggy loop was taking effectively huge steps (the accumulated sum of every past gradient), which on a four-example toy problem is a free lunch. Once fixed, you are taking honest lr-sized steps and need more of them. Compare the two by number of steps to reach a given loss, and raise the learning rate or the step count rather than reverting the fix.
- "I ran the notebook and my losses start at 0.002, not 4.8." — The checked-in micrograd_lecture_second_half_roughly.ipynb has stored output from a re-run of the training cell on an already-trained network: the loop prints 0.00205 down to 0.00179 over its 20 steps, because n was never re-initialized. To reproduce the video's trajectory, re-run the n = MLP(3, [4, 4, 1]) cell first. This is also a small lesson in why notebook outputs are not evidence.
- "Is that it? Where is the rest of training?" — That is genuinely it, for this problem. What real training adds — batching, better optimizers, learning-rate schedules, regularization, validation splits — are refinements of the loop, not additions to the theory. Karpathy shows several of them in demo.ipynb immediately after this part.
Go deeper, verified
- micrograd/nn.py L6–L8 — Module.zero_grad — Andrej Karpathy · the three-line refactor of this part's fix, deliberately shaped like torch.nn.Module.
- micrograd/engine.py L54–L70 — Value.backward — Andrej Karpathy · the topological sort plus reverse walk that the one loss.backward() line in the training loop invokes.
- micrograd demo.ipynb — the moons classifier — Andrej Karpathy · the same loop at slightly larger scale, with batching, a max-margin loss, L2 regularization, and learning_rate = 1.0 - 0.9*k/100 decay.
- torch.optim.Optimizer.zero_grad — PyTorch docs — PyTorch · the production form of the fix, including the set_to_none default that trips people who inspect .grad after zeroing. Field map extra.
- torch.optim.SGD — PyTorch docs — PyTorch · p.data += -lr * p.grad with the options this part does not have: momentum, weight decay, Nesterov. Field map extra.
- A Recipe for Training Neural Networks — Andrej Karpathy (2019) · the long-form version of the "most common neural net mistakes" list he pulls up at 130:31; the zero-grad omission is item three on that list. Field map extra.
- CS231n: Optimization — stochastic gradient descent — Andrej Karpathy / Stanford CS231n · the step-size discussion in written form, with the "effective learning rate" picture and finite-difference gradient checking. Field map extra.
- Learning representations by back-propagating errors — Rumelhart, Hinton & Williams, Nature (1986) · the original: squared-error loss, gradients by the chain rule, weights nudged along the negative gradient. The loop in this part is that paper's algorithm in Python. Field map extra.
Exercises
- Reproduce the bug, then measure itcode — Take the loop above and make two runs from the same seed: one with the zero-grad lines, one without. (a) Set random.seed(1337) and re-create the MLP identically for both. (b) Record loss at every step for 30 steps. (c) In the buggy run, also record the norm of p.grad for one fixed parameter each step. A good answer shows the buggy run's gradient magnitude growing roughly linearly while the correct run's stays bounded, notes that the buggy run reaches a lower loss sooner, and explains why that is an artifact of the problem being trivially fittable rather than a virtue.
- Find the learning rate cliffcode — Sweep lr over [0.001, 0.01, 0.05, 0.1, 0.5, 1.0, 2.0], 50 steps each, correct zero-grad, same initialization every time. Plot final loss against lr on a log x-axis. A good answer identifies three regimes — too slow to converge, a working band, and divergence — reports roughly where the boundary falls for this network, and says why that boundary is a property of this loss surface and does not transfer to a different model.
- Write zero_grad the way the library doescode — Add a Module base class with zero_grad() and a default parameters() returning [], and make Neuron, Layer and MLP inherit from it, so the loop becomes n.zero_grad(). Then check your version against nn.py L4–L11. A good answer notices that the base zero_grad works for all three classes without overriding, because each one already implements parameters().
- The official exercise notebook, section 3 — The exercise Colab from the video description ends by building a softmax and negative-log-likelihood loss on top of Value and checking the gradients against PyTorch. That verification step is this part's discipline applied to a different loss: swap mean squared error for cross-entropy and confirm the same training loop still drives it down.