Intuition: what a wrong model costs you
Transcript: this stretch, timestamped
The previous page built the definition: widths from one distribution, heights from another, H(p,q) = Σ pᵢ·log₂(1/qᵢ). A definition you can evaluate is not yet a definition you trust. This stretch is where Grant shrinks the world down to two outcomes — small enough that the whole function fits on one graph with one free variable per distribution — and reads the shape off the picture. What comes out is the single structural fact the rest of the series leans on: cross-entropy, viewed as a function of the model, is minimised at the truth. Every downstream claim — that the language-tree zipping trick was measuring something real (P11), that next-token pre-training is a well-posed objective (P12), that distillation works (P15) — is a corollary of the inequality proved here.
Outline, with timestamps
- 08:26 — Shrink the world: two outcomes, q even at 50/50, p skewed 90/10.
- 08:52 — First surprise: every symbol costs one bit under a uniform code, so the weights cancel and cross-entropy is 1 bit too.
- 09:22 — Swap the roles: skewed q, even p. Now the rare symbol carries a lot of bits and gets hit half the time.
- 09:52 — The number: ≈1.74 bits, against an entropy of 1. Order matters; the two arguments play different roles.
- 10:22 — Two outcomes means one free parameter each, p₁ and q₁, so the whole thing is graphable.
- 10:54 — Hold p fixed, plot against q: one clear minimum, located at q = p.
- 11:24 — Slide p around and trace where the minimum goes. The traced curve is drawn in green.
- 11:56 — The green curve is the binary entropy of p. The headline fact, stated.
- 12:28 — Back to the bar diagram for the general case: widths pinned to p, heights free, area bottoming out at H(p).
Two outcomes, because two outcomes fit on a page
The generalisation on the previous page bought reach at the price of feel. Σ pᵢ·log₂(1/qᵢ) over an arbitrary alphabet is a thing you can compute, but not a thing you can picture. So the move here is to collapse the alphabet to two symbols. A distribution over two outcomes has exactly one degree of freedom — pick q₁ and q₂ = 1 − q₁ is forced — which means the whole cross-entropy surface is a function of two numbers, and you can draw it.
The first case (08:26): the code is built for q = (½, ½) and reality turns out to be p = (0.9, 0.1). Under a 50/50 model every outcome carries log₂(1/0.5) = 1 bit, so the code spends one bit per symbol, always, no matter which symbol shows up. That is the whole reason this example is the right one to start with: when all the heights are equal, the widths cannot matter. The weighted average of a constant is that constant. Cross-entropy comes out at exactly 1 bit — the same as H(q) — even though the model is badly wrong about reality.
That is the first thing to notice, and it is easy to blow past. Cross-entropy did not detect the error at all in the raw number; the error only shows up when you compare against what a correct model would have achieved. And a correct model would have done much better: H(p) = 0.9·log₂(1/0.9) + 0.1·log₂(1/0.1) = 0.9·0.1520 + 0.1·3.3219 = 0.4690 bits. So the 1 bit you actually spend is 0.531 bits of pure waste per symbol. Reality here is nearly deterministic, and a code that refuses to notice pays more than double.
Swap the roles and the arithmetic changes
Now flip which distribution is which (09:22). The code is built for the skewed q = (0.9, 0.1), and reality arrives even, p = (½, ½). The code's heights are now log₂(1/0.9) = 0.1520 bits for the outcome it expects and log₂(1/0.1) = 3.3219 bits for the one it doesn't. Averaged under the code's own beliefs those heights are cheap — that's H(q) = 0.469 bits, below one bit, which is the payoff for being confident. But reality serves up the expensive symbol half the time rather than a tenth of the time, so:
H(p,q) = ½·log₂(1/0.9) + ½·log₂(1/0.1)
= ½·0.1520 + ½·3.3219
= 0.0760 + 1.6610
= 1.7370 bits per symbol
Grant quotes this as "around 1.74 bits"; the animation carries the digits, the captions don't, so the table below is my own arithmetic, computed to four places. Against H(p) = 1 bit, the model is burning 0.737 extra bits per symbol — a 74% overhead on the compressed file. Note what changed and what didn't: the same two distributions, the same formula, and yet 1.000 in one order and 1.737 in the other. The two slots of H(p,q) genuinely do different jobs. p sets the widths — how often each case actually happens. q sets the heights — how many bits your code committed to spending on it. Swapping them is not a relabelling, it is a different question.
| Configuration | H(q) — code's own entropy | H(p,q) — bits actually spent | H(p) — best possible | Excess D(p‖q) |
|---|---|---|---|---|
| code for q=(½,½), reality p=(0.9,0.1) | 1.0000 | 1.0000 | 0.4690 | 0.5310 |
| code for q=(0.9,0.1), reality p=(½,½) | 0.4690 | 1.7370 | 1.0000 | 0.7370 |
All figures in bits per symbol, computed by hand from log₂(1/0.9) = 0.15200 and log₂(1/0.1) = 3.32193. The last column is the Kullback–Leibler divergence, D(p‖q) = H(p,q) − H(p); it differs across the two rows, which is the cleanest possible demonstration that KL is not symmetric.
Fix p, sweep q, and the curve has exactly one bottom
With one free parameter per distribution, the payoff is that you can draw the thing (10:22). Pin p — decide what reality is and stop touching it — and plot H(p,q) as q₁ slides from 0 to 1. The graph dives to a single clear minimum, and the minimum sits at q₁ = p₁ (10:54).
"Always remember that question cross entropy is asking. How well does a code optimized for one setting, Q, perform in a different setting, P? That compression efficiency will obviously be at its best when both settings align."— Grant Sanderson, 10:54
My addition, since the video shows the graph rather than differentiating it: in the two-outcome case the claim is a one-line calculus exercise. Write f(q₁) = −p₁·log₂ q₁ − (1−p₁)·log₂(1−q₁). Then
f′(q₁) = (1/ln 2) · [ −p₁/q₁ + (1−p₁)/(1−q₁) ]
f′(q₁) = 0 ⟺ p₁(1−q₁) = q₁(1−p₁) ⟺ p₁ = q₁
f″(q₁) = (1/ln 2) · [ p₁/q₁² + (1−p₁)/(1−q₁)² ] > 0 for all q₁ ∈ (0,1)
The second derivative is strictly positive everywhere on the open interval, so f is strictly convex and the stationary point is the unique global minimum. Its value there is f(p₁) = H(p), by inspection: substituting q = p into Σ pᵢ·log₂(1/qᵢ) gives back the definition of entropy. Both ends blow up — as q₁ → 0 the term p₁·log₂(1/q₁) → ∞, and symmetrically at 1 — which is the graphical signature of the asymmetry we come to below.
The green curve is the entropy of p
Then the good bit (11:24). Slide p around. The whole cross-entropy curve reshapes as you go, and its minimum travels; trace the path of that travelling minimum and you get a second curve, drawn in green. What is it? It is the best achievable compression when reality is p — which is precisely the definition of H(p). In this two-outcome setting the green curve is the binary entropy function H₂(p₁) = −p₁·log₂ p₁ − (1−p₁)·log₂(1−p₁): zero at both ends, peaking at 1 bit at p₁ = ½. Geometrically, entropy is the lower envelope of the whole family of cross-entropy curves — each cross-entropy curve touches it at exactly one point, and that point is where the model tells the truth.
"…if you think of P as fixed and Q as the variable, it takes on its smallest possible value when Q is equal to P, and more specifically, that smallest value is the entropy of P."— Grant Sanderson, 11:56
The general statement, beyond two outcomes, is H(p,q) ≥ H(p) for every q, with equality if and only if q = p. This is Gibbs' inequality, and it deserves a proof rather than a graph, because everything downstream rests on it. The one-line version uses the elementary bound ln x ≤ x − 1, which holds for all x > 0 with equality only at x = 1 (the line x − 1 is the tangent to the concave ln at that point):
H(p) − H(p,q) = Σᵢ pᵢ·log₂(qᵢ/pᵢ)
= (1/ln 2) · Σᵢ pᵢ·ln(qᵢ/pᵢ)
≤ (1/ln 2) · Σᵢ pᵢ·(qᵢ/pᵢ − 1) [ln x ≤ x − 1]
= (1/ln 2) · ( Σᵢ qᵢ − Σᵢ pᵢ )
= (1/ln 2) · (1 − 1) = 0
⟹ H(p,q) ≥ H(p), i.e. D(p‖q) = H(p,q) − H(p) ≥ 0
Equality forces qᵢ/pᵢ = 1 for every i with pᵢ > 0, so q = p on the support of p; since both sum to 1, they agree everywhere. The same result drops out of Jensen's inequality applied to the concave logarithm: Σ pᵢ·log₂(qᵢ/pᵢ) ≤ log₂(Σ pᵢ·qᵢ/pᵢ) = log₂(Σ qᵢ) = log₂ 1 = 0, with strictness unless qᵢ/pᵢ is constant. Both routes are standard; I checked each derivation line by line rather than quoting it from memory.
In words, without symbols: you cannot beat the code that was designed for the truth. The optimal-code result from Part 1 says the best you can ever do against reality p is H(p) bits per symbol. A code built for some other q is a legitimate code — it obeys Kraft's inequality, it decodes fine — it is just one of the many codes that isn't the best one. So it lands above the floor. Gibbs' inequality is the formal statement that the floor is real and that only the truth touches it.
The two ways to be wrong are not priced alike
Gibbs tells you the penalty is non-negative. It does not tell you how it is distributed, and that distribution is the practically important part. The per-outcome price is log₂(1/q), and that function is brutally lopsided: it is gentle near q = 1 and unbounded near q = 0. Some numbers to hold:
| Probability the model assigned | Bits charged if it happens · log₂(1/q) |
|---|---|
| 0.5 | 1.00 |
| 0.25 | 2.00 |
| 0.1 | 3.32 |
| 0.01 | 6.64 |
| 0.001 | 9.97 |
| 0.000001 | 19.93 |
| 0 | ∞ |
Each factor-of-10 drop in assigned probability adds a flat log₂ 10 = 3.32 bits. There is no ceiling. Now contrast the two failure modes, both measured against the same truth p = (0.9, 0.1), whose entropy is 0.4690 bits:
| Model q₁ | Failure mode | H(p,q) | Excess over H(p) |
|---|---|---|---|
| 0.9 | correct | 0.4690 | 0.0000 |
| 0.5 | maximally vague | 1.0000 | 0.5310 |
| 0.99 | overconfident, right direction | 0.6774 | 0.2084 |
| 0.999 | very overconfident, right direction | 0.9979 | 0.5289 |
| 0.1 | confidently wrong | 3.0049 | 2.5359 |
| 0.01 | very confidently wrong | 5.9809 | 5.5119 |
| 0.001 | catastrophically wrong | 8.9694 | 8.5004 |
Computed by me from H(p,q) = 0.9·log₂(1/q₁) + 0.1·log₂(1/(1−q₁)), bits per symbol, four decimal places.
The shape of that table is the point. Vagueness has a hard ceiling. A model that gives up and spreads its mass uniformly over n outcomes pays exactly log₂ n bits per symbol, no matter what p is — for two outcomes, 1.00 bit, so the worst vagueness can ever cost you here is 1 − H(p) ≤ 1 bit. Even for a 50,000-token vocabulary the uniform model's bill is a finite 15.6 bits. Confident error has no ceiling at all. Assign a thousandth to something that happens a tenth of the time and you are already 8.5 bits over the floor, and the column keeps going. Note also that the two directions of overconfidence are not equal: pushing q₁ to 0.999 in the correct direction costs 0.53 bits, the same order as giving up entirely, because the 10% of the time you are wrong you pay 9.97 bits. Overconfidence hurts wherever it points; it is only ruinous when it points the wrong way.
This asymmetry is why cross-entropy training yields hedged, roughly calibrated probabilities rather than bravado. A model tempted to sharpen a 0.9 into a 0.999 saves 0.14 bits per symbol averaged over the 90% branch, and pays 6.64 extra bits on every occurrence of the 10% branch — 0.66 bits per symbol once weighted. The net trade is a loss of 0.53 bits, and gradient descent can see that. The unbounded left tail of log₂(1/q) acts as an insurance premium against ever saying "impossible" about something that can happen.
Holding the fact in your head for many outcomes
The graph trick only works with one free parameter per distribution, so for the general case Grant goes back to the bar diagram (12:28): one bar per symbol, width pᵢ set by reality, height log₂(1/qᵢ) set by your code, total area equal to H(p,q). Freeze the widths. Now the heights are the only thing you control, and they are not independent — pulling one bar down forces others up, because the qᵢ must still sum to 1. That constraint is exactly what makes the problem interesting: there is a cheapest legal assignment of heights, and it is the one where each height matches its own width, log₂(1/pᵢ).
Two footnotes that the two-outcome picture can hide. First, H(p,q) ≥ H(p) says nothing about H(q). Take q = (½, ¼, ¼), so H(q) = 1.5 bits, and let reality be the point mass p = (1, 0, 0). Then H(p,q) = 1·log₂(1/½) = 1 bit, comfortably below H(q) — and still above H(p) = 0, as it must be. The bound is against the entropy of the distribution in the width slot, always. Second, this is where the constrained-optimisation flavour of the argument comes from: minimising Σ pᵢ·log₂(1/qᵢ) subject to Σ qᵢ = 1 is a textbook Lagrange-multiplier problem, and Grant's video description points at his own Lagrange-multiplier material for readers who want to run it that way. It gives the same answer as the Gibbs argument above, by a longer road.
That is the whole payload of this chapter. One inequality, one picture, and one asymmetry. The next page cashes the inequality in against the opening puzzle: what, exactly, was gzip measuring when it "recognised" a language.
Where people get stuck
"Which distribution goes in which slot?" — the most common and most consequential confusion, made worse by the fact that the notation in the wild is genuinely inconsistent (Grant flags this himself). The reliable mnemonic is the diagram, not the letters: reality sets the widths, the model sets the heights. The distribution you are averaging over — the one supplying the frequencies, the one your data is drawn from — is the width. The distribution supplying the code lengths log₂(1/·) is the height. In training, your data is the width and your network's output is the height. If you find yourself writing H(model, data), you have the arguments backwards and the inequality will point the wrong way.
"The first example gave 1 bit and the model was wrong — did cross-entropy fail?" — no, and this is worth sitting with. A raw cross-entropy number is not a measure of error; it is a bit count. It mixes together the irreducible cost of reality's own randomness, H(p), and the cost of your model being wrong, D(p‖q). Only the second is a mistake. That is why an LLM's loss of, say, 2.1 nats tells you very little on its own about model quality: you don't know how much of it is H(p). Comparisons between models on the same data are meaningful; the absolute number is not.
"log₂(1/q) or −log₂ q?" — identical, and picking one and sticking to it removes a class of sign errors. log₂(1/q) = −log₂ q, and it is non-negative because q ≤ 1. Reading it as "how many halvings from 1 down to q" keeps it manifestly positive and matches the code-length intuition; reading it as "minus log" is more compact but invites a dropped sign. Grant makes the same remark in the recap earlier in this video (04:39), that he half-wishes the field had written it as log₍½₎ q.
"What happens at zero?" — the two zeros behave completely differently. If pᵢ = 0 the term contributes nothing, whatever the model said, because 0·log₂(1/qᵢ) = 0: you are never charged for a possibility that never occurs. But if pᵢ > 0 and qᵢ = 0, cross-entropy is +∞ — one occurrence of something your model declared impossible and the average is destroyed. This is not a mathematical curiosity; it is why every practical implementation clips, smooths, or uses a softmax that cannot output an exact zero, and why the same asymmetry shows up as numerical NaNs in undertrained models.
"Why isn't KL symmetric, if it's a 'distance'?" — because cross-entropy isn't, and the table above shows it numerically: D(p‖q) = 0.531 in one direction and 0.737 in the other for the very same pair. D is a divergence, not a metric — non-negative, zero only when the arguments agree, but neither symmetric nor obedient to the triangle inequality. That is a feature: the question "how badly does a code for q serve reality p" is not the same question as its reverse, and no symmetric quantity could answer both.
Going deeper, verified
- But what is entropy? | Compression is Intelligence Part 1 — 3Blue1Brown (2026) · linked from this video's own description; the optimal-code result that makes H(p) the floor this page's inequality is measured against.
- A Mathematical Theory of Communication — Claude E. Shannon (1948) · the source of the entropy definition and of the coding theorem that says H(p) is achievable and unbeatable; §6 and the appendix carry the convexity arguments this page reproduces.
- On Information and Sufficiency — Solomon Kullback and Richard A. Leibler (1951), Annals of Mathematical Statistics 22(1):79–86 · the divergence D(p‖q) = H(p,q) − H(p) in its original setting, as a statistical discrimination measure rather than a compression overhead.
- Gibbs' inequality — Wikipedia · a compact statement and proof of exactly the H(p,q) ≥ H(p) result, including the equality condition and the ln x ≤ x − 1 route used above.
- Strictly Proper Scoring Rules, Prediction, and Estimation — Tilmann Gneiting and Adrian E. Raftery (2007), JASA 102(477):359–378 · the forecasting-theory frame for the callout: the logarithmic score is strictly proper precisely because of Gibbs' inequality, which is the formal reason cross-entropy training rewards honesty.
Exercises
- Redraw the green curve — for p₁ ∈ {0.1, 0.3, 0.5}, plot H(p,q) against q₁ on one set of axes, then overlay the binary entropy function H₂(p₁). A good answer shows three convex curves each touching the entropy curve at exactly one point, that point being q₁ = p₁, and notes that the curves diverge to +∞ at both ends while the entropy curve stays bounded by 1.
- Price the tail — take p = (0.9, 0.1) and find, numerically, the value of q₁ > 0.9 at which overconfidence in the right direction costs the same as giving up entirely at q₁ = ½. A good answer lands near q₁ ≈ 0.999 (excess 0.5289 bits versus 0.5310), and explains the mechanism: the gain of 0.13 bits on the 90% branch is swamped by 9.97 bits charged on the 10% branch.
- Break the wrong bound — construct a p and q over three outcomes with H(p,q) < H(q), then verify H(p,q) ≥ H(p) still holds. A good answer explains why the second inequality can never be broken while the first has no reason to hold, and identifies the width slot as the one the bound is about. (One worked case is in the section above; find a different one, ideally with p not a point mass.)
- Reproduce both tables — write ten lines of code computing H(p), H(p,q) and D(p‖q) for arbitrary discrete distributions, and regenerate every number on this page. A good answer handles pᵢ = 0 correctly (contributes 0) and qᵢ = 0 correctly (returns infinity only when the matching pᵢ is positive).