Why this loss and not another one
Transcript: this stretch, timestamped
P12 built the pre-training loss from scratch and got to a formula that is embarrassingly simple: average −log q(true token) over every position in the corpus. Nowhere in that derivation did two distributions appear, and cross-entropy is a two-distribution quantity. This page closes that gap, and it does so twice over — first by showing where the second distribution was hiding all along (it appears the moment you group the corpus by repeated context), then by proving that no other choice of loss would have worked. That second half is where the argument does real work, and it is also where Grant deliberately puts the details on screen rather than narrating them, so this page reconstructs them in full. What follows sets up P14, where the hidden distribution stops being hypothetical and becomes an actual second model.
Outline, with timestamps
- 20:38 — The naming puzzle: a loss built from logs, called cross-entropy loss.
- 20:50 — The one-hot answer: cross-entropy against a spike collapses to a single log term.
- 21:20 — Why that answer is unsatisfying: if the formula evaporates, why keep the name?
- 21:52 — Restart with an unknown loss F: the only obvious constraint is that it decrease.
- 22:23 — “My name is ___”: a context common enough to appear thousands of times.
- 23:26 — Weighting by p, the empirical frequency: the sum becomes a weighted average.
- 23:56 — Set F = −log and the weighted average is the cross-entropy formula.
- 24:30 — The implication run backwards: demand honesty, and F must be a log.
- 25:02 — The on-screen details: minimize subject to Σ q = 1, tangent contours, equal gradients — boiling down at 25:32 to F′(q) = c/q, so your hand is forced.
- 26:02 — A second footnote on overfitting, and the handoff to distillation.
The name that doesn’t fit yet, and the distribution that was hiding
Start with the standard explanation, because you will meet it everywhere and you should know exactly what is wrong with it. Fix a position in the training text whose true next token is t. Define a target distribution y that is one-hot at t: y(t) = 1, y(x) = 0 for every other token. Then write down the cross-entropy of the model's distribution q relative to that target:
H(y, q) = Σ y(x) · log( 1 / q(x) )
x
= 1 · log(1/q(t)) + 0 · log(1/q(x₁)) + 0 · log(1/q(x₂)) + …
= −log q(t)
Every term but one is annihilated by a zero weight, and what survives is exactly the per-token loss from P12. So the name is defensible. But notice what the derivation costs: you invented a distribution whose only job was to disappear. If the cross-entropy machinery contributes nothing that survives to the final expression, calling the result cross-entropy loss is decoration, not explanation.
"This does nothing to explain why you're using cross-entropy in the first place, and if it all just evaporates away anyway, why not just be straightforward? Why not just call it log loss, or information loss?"— Grant Sanderson, 21:20
Two better explanations follow. The first says the second distribution was real, you were just looking at one training example at a time and could not see it. The second says that even if you had never heard of cross-entropy, you would have been forced to invent it.
Take the first. Pick a context that recurs: "My name is ___". In a web-scale corpus that prefix appears an enormous number of times, followed by different names in proportion to how common those names are in the data. The model does not get to see a distribution — it sees one name per occurrence. But the loss is a sum over occurrences, and a sum over occurrences of the same context is a weighted average over what followed.
Write cᵢ for the context at position i and xᵢ for the token that actually followed. Let p̂(x | c) be the empirical frequency of x among all the positions sharing context c, and N_c the number of those positions. Then the corpus loss regroups exactly:
average loss = (1/N) · Σ −log q(xᵢ | cᵢ) one term per token position
i
group positions by their context c
= Σ (N_c / N) · Σ p̂(x | c) · log( 1 / q(x | c) )
c x
= Σ (N_c / N) · H( p̂(·|c) , q(·|c) )
c
That is the whole trick, and it is an identity, not an approximation — no distribution was estimated, only bookkeeping was rearranged. The empirical conditional p̂(·|c) never appears in any line of training code; it exists only in the aggregate, spread across thousands of gradient steps. But it is what the loss is really scoring against. Cross-entropy loss is a cross-entropy — averaged over contexts, weighted by how often each context occurs.
And now the property from the first half of the video earns its keep. For each context c, the term H(p̂(·|c), q(·|c)) is minimized precisely when q(·|c) = p̂(·|c). Minimizing the total therefore drives the model toward the data's own conditional statistics, context by context. Concretely, with four possible names and their empirical frequencies fixed, here is what four candidate models pay (in bits per token; multiply by ln 2 ≈ 0.693 for nats):
| model's distribution q over (Maria, Wei, Ana, Yusuf) | H(p, q) | D(p‖q) |
|---|---|---|
| truthful — (½, ¼, ⅛, ⅛), equal to p | 1.750 bits | 0.000 |
| hedged — (0.40, 0.30, 0.15, 0.15), flattened toward uniform | 1.779 bits | 0.029 |
| uniform — (0.25, 0.25, 0.25, 0.25), no opinion at all | 2.000 bits | 0.250 |
| bluffing — (0.90, 0.05, 0.025, 0.025), right ranking, wrong confidence | 2.487 bits | 0.737 |
The row worth staring at is the last one. That model has the ordering of the names exactly right — it would score a perfect 100% on top-1 accuracy against the mode — and it is punished harder than the model with no opinion whatsoever. Overconfidence in the correct direction is worse than agnosticism. Hedging is punished too, only more gently. The truthful report is the unique minimum, and every deviation from it costs, in either direction. Hold that thought; the next section explains why it is a theorem rather than an accident of these four numbers.
The Lagrange argument, run forwards
First the easy direction: verify that F = −log really does have the honesty property. Fix a distribution p over outcomes. Among all distributions q — that is, all non-negative vectors summing to 1 — which one minimizes the expected reported loss Σ p(x)·log(1/q(x))? The constraint is what makes this interesting: without Σ q(x) = 1 you would just send every q(x) to infinity. Grant shows this on screen as two contours becoming tangent; the algebra behind that picture is a Lagrange multiplier, using natural logs so the derivative is clean.
minimize L(q) = Σ p(x) · ln( 1 / q(x) ) subject to Σ q(x) = 1
x x
Lagrangian: 𝓛(q, λ) = Σ p(x) · ln(1/q(x)) + λ · ( Σ q(x) − 1 )
x x
since d/dq [ −ln q ] = −1/q :
∂𝓛/∂q(x) = − p(x)/q(x) + λ = 0 ⟹ q(x) = p(x) / λ
impose the constraint:
Σ q(x) = (1/λ) · Σ p(x) = 1/λ = 1 ⟹ λ = 1 ⟹ q(x) = p(x)
The multiplier lands on λ = 1, which is a useful sanity check: it says the marginal cost of the last unit of probability mass is the same everywhere at the optimum. And the stationary point is a genuine minimum rather than a saddle, because the objective is strictly convex in q — the Hessian is diagonal with entries ∂²L/∂q(x)² = p(x)/q(x)² > 0 for every outcome with p(x) > 0, and the feasible set (the probability simplex) is convex. A strictly convex function on a convex set has at most one minimizer, so q = p is not just a minimum, it is the minimum. (That convexity check is my addition; the video asserts uniqueness from the graph.)
The Lagrange argument, run backwards — the part that matters
Now the direction that does the real work, and the one Grant flashes on screen for the curious. Suppose you knew nothing about logs. You have some decreasing per-token loss F, and you make one demand: for every possible data distribution p, the expected loss Σ p(x)·F(q(x)) should be minimized at q = p and nowhere else. What functions F survive?
stationarity for a generic F: p(x) · F′(q(x)) + λ = 0 for every x so at the optimum, p(x) · F′(q(x)) = −λ, the SAME number for every x. demand the optimum be q = p: p(x) · F′(p(x)) = −λ(p) for every x write φ(t) = t · F′(t). The demand says: φ takes the same value on every entry of p — and this must hold for EVERY distribution p. pick any u, v in (0,1). Let w = ½ · min(1−u, 1−v), so w > 0 and u+w < 1, v+w < 1. (u, w, 1−u−w) is a valid distribution ⟹ φ(u) = φ(w) (v, w, 1−v−w) is a valid distribution ⟹ φ(v) = φ(w) hence φ(u) = φ(v) for all u, v: φ ≡ c, a constant. t · F′(t) = c ⟹ F′(t) = c / t ⟹ F(t) = c · ln t + d F decreasing forces c < 0. Taking c = −1, d = 0: F(t) = −ln t
Read the last two lines slowly, because they are the whole point of the chapter. The condition F′(t) = c/t is Grant's “some constant over q,” and the functions whose derivative is a constant over the variable are exactly the logarithms, up to an additive constant. So the surviving family is F(t) = a·(−ln t) + b with a > 0 — the negative log, rescaled and shifted. A positive rescaling is absorbed into the learning rate (which is exactly why base-2 versus base-e does not matter, as noted at 19:13), and an additive constant has zero gradient. Every member of the family is the same loss wearing different units.
"…if you want a loss function with the property that it's minimized only when the model matches the statistics of the data, your hand is forced. You actually have to use the negative log."— Grant Sanderson, 25:32
My addition, and a caveat worth knowing: the step where I built three-outcome distributions is not decorative. On a sample space with only two outcomes the same argument yields only φ(t) = φ(1−t), which is far weaker than constancy, and the logarithm is genuinely not the unique answer there. The uniqueness theorem needs at least three possible outcomes. For a language model with a vocabulary in the tens of thousands this is not a constraint anyone notices, but it is the reason the corresponding result is usually stated for categorical variables with n ≥ 3.
The name for this property: a strictly proper scoring rule
What has actually been proved is a statement from decision theory that predates deep learning by decades. Score a forecaster by S(q, x) — the penalty they pay for having reported distribution q when outcome x occurred. The rule is proper if reporting your true beliefs is among the expected-loss-minimizing reports, and strictly proper if it is the only one. Cross-entropy — under this name, the logarithmic score — is strictly proper. So is the Brier score, the other famous member of the family. The linear score S(q, x) = −q(x) is not proper at all.
Strict propriety has a clean information-theoretic reading that ties this page back to the first half of the video. Split the expected loss:
H(p, q) = Σ p(x) · log( 1 / q(x) )
x
= Σ p(x) · log( 1 / p(x) ) + Σ p(x) · log( p(x) / q(x) )
x x
= H(p) + D(p ‖ q)
The first term is the entropy of the data — irreducible, independent of the model, the floor that P12 described as “the entropy of language.” The second is the Kullback–Leibler divergence, and Gibbs' inequality says D(p‖q) ≥ 0 with equality if and only if p = q. So the regret for reporting anything other than your true beliefs is exactly the KL divergence between the truth and your report. Strict propriety and Gibbs' inequality are the same statement in two vocabularies, and the D(p‖q) column in the names table above is literally the price of dishonesty in bits. (The video returns to KL as its closing footnote, at 31:35.)
Why the obvious alternatives lose
Accuracy is not strictly proper. Score a model by whether its argmax matched the true token, and the expected score is p(argmax q) — maximized by any distribution whose mode is the data's mode. The truthful report is among the winners, so accuracy is weakly proper, but so is a model that reports 99.9% confidence on the mode and rounding error everywhere else. Accuracy cannot tell those apart. It is also piecewise constant, hence has zero gradient almost everywhere, which disqualifies it as a training objective before the propriety argument even starts.
A rule that rewards confidence rewards lying. Take the linear score, paying the model q(x) when x occurs. Expected payoff is Σ p(x)·q(x), which is linear in q — and a linear functional over a simplex is maximized at a vertex. For p = (0.6, 0.4), honest reporting pays 0.6² + 0.4² = 0.52, while claiming certainty pays 0.6. The rule pays a premium for bluffing. This is not a subtle failure mode; it is the generic behavior of any score linear in the report.
Squared error is proper but trains badly. The Brier score Σₓ (q(x) − y(x))² against a one-hot target is strictly proper — it is not disqualified on honesty grounds. It is disqualified on gradients. Differentiate both losses with respect to the pre-softmax logits z, where q = softmax(z), in the situation that matters most (confident and wrong: probability q_t on the true token, essentially all remaining mass on one wrong token):
cross-entropy: ∂L/∂z_j = q_j − y_j so |∂L/∂z_t| = 1 − q_t → 1 Brier: ∂L/∂z_t = −4 · q_t · (1 − q_t)² → 0
| probability on the true token | cross-entropy loss (nats) | Brier loss | |∂L/∂z_t|, cross-entropy | |∂L/∂z_t|, Brier |
|---|---|---|---|---|
| 0.5 | 0.693 | 0.500 | 0.500 | 0.500 |
| 0.1 | 2.303 | 1.620 | 0.900 | 0.324 |
| 0.01 | 4.605 | 1.960 | 0.990 | 0.0392 |
| 0.001 | 6.908 | 1.996 | 0.999 | 0.00399 |
The two agree at q_t = ½ and then diverge violently. Cross-entropy's loss is unbounded and its gradient strengthens toward the maximum as the model gets more confidently wrong; the softmax factor that would damp it cancels exactly against the 1/q from the log. Brier's loss is bounded above by 2 no matter how wrong the model is, and its gradient collapses like 4·q_t — at q_t = 0.001 it is 250× weaker than cross-entropy's. A confidently-wrong model under Brier is nearly a fixed point. That is the practical difference, and it has nothing to do with propriety: both losses are honest at the optimum, but only one of them can find its way there from a bad initialization.
The same statement in likelihood language
Everything above can be said without mentioning entropy at all, in the vocabulary most readers already have. Because a language model factors the probability of a document by the chain rule, the sum of per-token log-probabilities is the log-probability of the document:
ln P(x₁ … x_N | θ) = Σ ln q_θ(xᵢ | x₁ … xᵢ₋₁)
i
so the training loss = −(1/N) · ln P(corpus | θ) = −(1/N) · log-likelihood
Minimizing average cross-entropy on one-hot targets is exactly maximum likelihood estimation on the corpus. The two are the same objective up to the factor 1/N. Equivalently, since H(p̂, q) = H(p̂) + D(p̂‖q) and H(p̂) does not depend on the parameters, MLE is minimizing D(p̂ ‖ q_θ) — the KL divergence from the empirical distribution to the model. The three framings (average information per token, maximum likelihood, minimum KL to the data) are one objective seen from three directions, which is precisely why the same formula keeps turning up in unrelated corners of the literature.
One practical consequence: perplexity, the number quoted in every language-model paper, is just this loss exponentiated. perplexity = exp(mean cross-entropy in nats) = 2^(mean cross-entropy in bits). A perplexity of 12 means the model is, on average, as uncertain as if it were choosing uniformly among 12 tokens — and the bits-per-token reading is the direct bridge to the compression story the series picks up next.
What propriety does and does not buy you
Strict propriety is a statement about the global minimum of the expected loss over all distributions. It says: if your model could represent any distribution, and you found the true optimum, and you had infinite data, then the loss-minimizing model is the perfectly calibrated one. Every one of those three conditions fails in practice. Real networks have finite capacity, are trained by a stochastic optimizer that stops early, and see a finite corpus — and trained deep networks are in fact reliably overconfident relative to their accuracy, which is why post-hoc temperature scaling is a standard fix. Propriety guarantees that the loss is not pushing the model toward dishonesty; it guarantees nothing about where the model actually lands.
That gap is also, quietly, the subject of Grant's second on-screen footnote at 26:02 about overfitting: the empirical conditional p̂(·|c) that the grouping argument leans on only exists for contexts that recur, and the overwhelming majority of contexts in a large corpus are seen exactly once. The honest version of the argument is that the loss is a cross-entropy against a distribution the model must generalize its way to, not one it can ever observe. Which is exactly the pressure P14 relieves: replace the sparse empirical target with a dense one supplied by a bigger model, and every single example carries a full distribution's worth of signal.
Where people get stuck
“But there is only one distribution in the code.” Correct, and this is the single most common source of confusion about the name. Open any training loop and you will find loss = −logsoftmax(logits)[target], one distribution and one integer index. The second distribution is not materialized anywhere; it exists only after you sum over many positions that share a context, and it is empirical — defined by the corpus's own frequencies. Cross-entropy loss is a cross-entropy in expectation, not per example.
Which distribution goes in which slot. In H(p, q) = Σ p(x)·log(1/q(x)), the first argument supplies the weights (reality, the data) and the second supplies the code lengths (the model, the thing you optimize). The video's phrasing is “the cross-entropy of q relative to p” for exactly this expression — the code came from q, the world is p. Notation in the wild is genuinely inconsistent, so the safe habit is to check which argument the sum is weighted by rather than trusting the argument order. The quantity is asymmetric: swapping the roles gives a different number, as the 90/10 versus 50/50 example at 09:22 shows.
The sign in the Lagrangian. Written as Σ p·ln(1/q) + λ(Σ q − 1), the derivative of the first term with respect to q(x) is −p(x)/q(x) — negative, because pushing more probability onto any outcome always lowers the loss. The multiplier is what pushes back, and it must therefore come out positive; getting λ = 1 rather than λ = −1 is a cheap check that you differentiated correctly. If you write the objective as −Σ p·ln q instead, the same derivative appears with the same sign; the two forms are identical, not merely similar.
“Strictly proper” does not mean “best.” The Brier score is strictly proper too. Propriety is a necessary condition for a sane probabilistic loss, not a sufficient one — it rules out losses that reward lying, and says nothing about optimization dynamics, numerical stability, or robustness to label noise. The uniqueness theorem in this chapter is stronger than propriety alone: it says the log score is the only local strictly proper rule, where local means the score depends only on the probability assigned to the outcome that actually happened. Brier is strictly proper but not local — it looks at the probabilities assigned to outcomes that did not occur. If you want a loss that reads off a single number from the model's output, you get the logarithm and nothing else.
Going deeper, verified
- Lagrange multipliers and constrained optimization — Grant Sanderson, Khan Academy · The exact link from this video's description; the tangent-contours picture the derivation above turns into algebra.
- Strictly Proper Scoring Rules, Prediction, and Estimation — Tilmann Gneiting and Adrian E. Raftery (2007), JASA 102(477) · The canonical survey; the logarithmic, quadratic (Brier) and spherical scores, and the Savage representation linking propriety to convexity.
- Expected Information as Expected Utility — José M. Bernardo (1979), Annals of Statistics 7(3) · The characterization result: the logarithmic score is essentially the only local proper scoring rule, which is the theorem this chapter derives by hand.
- Verification of Forecasts Expressed in Terms of Probability — Glenn W. Brier (1950), Monthly Weather Review 78(1) · Three pages that introduced the other famous strictly proper rule, from meteorology rather than information theory.
- On Calibration of Modern Neural Networks — Guo, Pleiss, Sun and Weinberger (2017) · The empirical counterweight: a strictly proper loss does not produce a calibrated network, and temperature scaling as the standard repair.
Exercises
- Reproduce the names table, then break it. With p = (½, ¼, ⅛, ⅛), compute H(p, q) in bits for the four models in the table and confirm all four numbers to three decimals, plus H(p) = 1.75. Then add a fifth model that assigns probability 0 to Yusuf and renormalizes the rest. A good answer reports H(p, q) = ∞, explains that the culprit is a single term 0.125 · log(1/0) , and connects it to why implementations clamp probabilities or work in log-space throughout.
- Find the stationary point numerically and read off the multiplier. For the same p, minimize Σ p(x)·ln(1/q(x)) over the simplex by projected gradient descent (or by parameterizing q = softmax(z) and doing unconstrained descent on z). A good answer converges to q = p to several decimals, reports the minimum as H(p) = 1.75 · ln 2 = 1.213 nats, and verifies that p(x)/q(x) = 1 for every x at convergence — the numerical fingerprint of λ = 1.
- Demonstrate impropriety, then repair it. Take p = (0.6, 0.4). Under the linear score S(q, x) = q(x), compute the expected payoff for the honest report and for the certain report (1, 0); confirm 0.52 versus 0.60, so lying wins. Now score the same two reports with the log score and confirm the ranking flips: 0.6·ln(1/0.6) + 0.4·ln(1/0.4) = 0.673 nats for the honest report, versus ∞ for the certain one. Extra credit: sweep q₁ from 0 to 1 and plot both scores, so you can see the linear one sloping monotonically to the vertex while the log one bottoms out at q₁ = 0.6.