ENTROPY // FIELD MAP
← field map
P10 · CROSS-ENTROPY3Blue1Brown · 08:26–12:59 · 5 min

Intuition: what a wrong model costs you

"But what is cross-entropy? | Compression is Intelligence Part 2" — the Intuition and examples chapter, where the formula from the previous stretch gets pushed through two-outcome cases until it has a shape you can feel.

Transcript: this stretch, timestamped

TL;DR — Cross-entropy H(p,q) = Σ pᵢ·log₂(1/qᵢ) is the average bits you spend when reality is p but your code was built for q. Push it through two-outcome examples and one fact falls out: hold p fixed, vary q, and the function has a single minimum, sitting exactly at q = p, with minimum value H(p). That is Gibbs' inequality, H(p,q) ≥ H(p), and it is the entire licence for using cross-entropy as a training loss — the thing the loss is trying to reach is the truth. The second fact worth carrying: the penalty is wildly asymmetric. Being vague costs a bounded amount; being confidently wrong costs unboundedly much, because log₂(1/q) blows up as q → 0.

The previous page built the definition: widths from one distribution, heights from another, H(p,q) = Σ pᵢ·log₂(1/qᵢ). A definition you can evaluate is not yet a definition you trust. This stretch is where Grant shrinks the world down to two outcomes — small enough that the whole function fits on one graph with one free variable per distribution — and reads the shape off the picture. What comes out is the single structural fact the rest of the series leans on: cross-entropy, viewed as a function of the model, is minimised at the truth. Every downstream claim — that the language-tree zipping trick was measuring something real (P11), that next-token pre-training is a well-posed objective (P12), that distillation works (P15) — is a corollary of the inequality proved here.

Outline, with timestamps

Two outcomes, because two outcomes fit on a page

The generalisation on the previous page bought reach at the price of feel. Σ pᵢ·log₂(1/qᵢ) over an arbitrary alphabet is a thing you can compute, but not a thing you can picture. So the move here is to collapse the alphabet to two symbols. A distribution over two outcomes has exactly one degree of freedom — pick q₁ and q₂ = 1 − q₁ is forced — which means the whole cross-entropy surface is a function of two numbers, and you can draw it.

The first case (08:26): the code is built for q = (½, ½) and reality turns out to be p = (0.9, 0.1). Under a 50/50 model every outcome carries log₂(1/0.5) = 1 bit, so the code spends one bit per symbol, always, no matter which symbol shows up. That is the whole reason this example is the right one to start with: when all the heights are equal, the widths cannot matter. The weighted average of a constant is that constant. Cross-entropy comes out at exactly 1 bit — the same as H(q) — even though the model is badly wrong about reality.

That is the first thing to notice, and it is easy to blow past. Cross-entropy did not detect the error at all in the raw number; the error only shows up when you compare against what a correct model would have achieved. And a correct model would have done much better: H(p) = 0.9·log₂(1/0.9) + 0.1·log₂(1/0.1) = 0.9·0.1520 + 0.1·3.3219 = 0.4690 bits. So the 1 bit you actually spend is 0.531 bits of pure waste per symbol. Reality here is nearly deterministic, and a code that refuses to notice pays more than double.

Swap the roles and the arithmetic changes

Now flip which distribution is which (09:22). The code is built for the skewed q = (0.9, 0.1), and reality arrives even, p = (½, ½). The code's heights are now log₂(1/0.9) = 0.1520 bits for the outcome it expects and log₂(1/0.1) = 3.3219 bits for the one it doesn't. Averaged under the code's own beliefs those heights are cheap — that's H(q) = 0.469 bits, below one bit, which is the payoff for being confident. But reality serves up the expensive symbol half the time rather than a tenth of the time, so:

H(p,q) = ½·log₂(1/0.9) + ½·log₂(1/0.1)
       = ½·0.1520  + ½·3.3219
       = 0.0760    + 1.6610
       = 1.7370 bits per symbol

Grant quotes this as "around 1.74 bits"; the animation carries the digits, the captions don't, so the table below is my own arithmetic, computed to four places. Against H(p) = 1 bit, the model is burning 0.737 extra bits per symbol — a 74% overhead on the compressed file. Note what changed and what didn't: the same two distributions, the same formula, and yet 1.000 in one order and 1.737 in the other. The two slots of H(p,q) genuinely do different jobs. p sets the widths — how often each case actually happens. q sets the heights — how many bits your code committed to spending on it. Swapping them is not a relabelling, it is a different question.

ConfigurationH(q) — code's own entropyH(p,q) — bits actually spentH(p) — best possibleExcess D(p‖q)
code for q=(½,½), reality p=(0.9,0.1)1.00001.00000.46900.5310
code for q=(0.9,0.1), reality p=(½,½)0.46901.73701.00000.7370

All figures in bits per symbol, computed by hand from log₂(1/0.9) = 0.15200 and log₂(1/0.1) = 3.32193. The last column is the Kullback–Leibler divergence, D(p‖q) = H(p,q) − H(p); it differs across the two rows, which is the cleanest possible demonstration that KL is not symmetric.

Fix p, sweep q, and the curve has exactly one bottom

With one free parameter per distribution, the payoff is that you can draw the thing (10:22). Pin p — decide what reality is and stop touching it — and plot H(p,q) as q₁ slides from 0 to 1. The graph dives to a single clear minimum, and the minimum sits at q₁ = p₁ (10:54).

"Always remember that question cross entropy is asking. How well does a code optimized for one setting, Q, perform in a different setting, P? That compression efficiency will obviously be at its best when both settings align."— Grant Sanderson, 10:54

My addition, since the video shows the graph rather than differentiating it: in the two-outcome case the claim is a one-line calculus exercise. Write f(q₁) = −p₁·log₂ q₁ − (1−p₁)·log₂(1−q₁). Then

f′(q₁) = (1/ln 2) · [ −p₁/q₁ + (1−p₁)/(1−q₁) ]

f′(q₁) = 0  ⟺  p₁(1−q₁) = q₁(1−p₁)  ⟺  p₁ = q₁

f″(q₁) = (1/ln 2) · [ p₁/q₁² + (1−p₁)/(1−q₁)² ]  >  0   for all q₁ ∈ (0,1)

The second derivative is strictly positive everywhere on the open interval, so f is strictly convex and the stationary point is the unique global minimum. Its value there is f(p₁) = H(p), by inspection: substituting q = p into Σ pᵢ·log₂(1/qᵢ) gives back the definition of entropy. Both ends blow up — as q₁ → 0 the term p₁·log₂(1/q₁) → ∞, and symmetrically at 1 — which is the graphical signature of the asymmetry we come to below.

The green curve is the entropy of p

Then the good bit (11:24). Slide p around. The whole cross-entropy curve reshapes as you go, and its minimum travels; trace the path of that travelling minimum and you get a second curve, drawn in green. What is it? It is the best achievable compression when reality is p — which is precisely the definition of H(p). In this two-outcome setting the green curve is the binary entropy function H₂(p₁) = −p₁·log₂ p₁ − (1−p₁)·log₂(1−p₁): zero at both ends, peaking at 1 bit at p₁ = ½. Geometrically, entropy is the lower envelope of the whole family of cross-entropy curves — each cross-entropy curve touches it at exactly one point, and that point is where the model tells the truth.

"…if you think of P as fixed and Q as the variable, it takes on its smallest possible value when Q is equal to P, and more specifically, that smallest value is the entropy of P."— Grant Sanderson, 11:56

The general statement, beyond two outcomes, is H(p,q) ≥ H(p) for every q, with equality if and only if q = p. This is Gibbs' inequality, and it deserves a proof rather than a graph, because everything downstream rests on it. The one-line version uses the elementary bound ln x ≤ x − 1, which holds for all x > 0 with equality only at x = 1 (the line x − 1 is the tangent to the concave ln at that point):

H(p) − H(p,q) = Σᵢ pᵢ·log₂(qᵢ/pᵢ)
              = (1/ln 2) · Σᵢ pᵢ·ln(qᵢ/pᵢ)
              ≤ (1/ln 2) · Σᵢ pᵢ·(qᵢ/pᵢ − 1)        [ln x ≤ x − 1]
              = (1/ln 2) · ( Σᵢ qᵢ − Σᵢ pᵢ )
              = (1/ln 2) · (1 − 1)  =  0

⟹  H(p,q) ≥ H(p),  i.e.  D(p‖q) = H(p,q) − H(p) ≥ 0

Equality forces qᵢ/pᵢ = 1 for every i with pᵢ > 0, so q = p on the support of p; since both sum to 1, they agree everywhere. The same result drops out of Jensen's inequality applied to the concave logarithm: Σ pᵢ·log₂(qᵢ/pᵢ) ≤ log₂(Σ pᵢ·qᵢ/pᵢ) = log₂(Σ qᵢ) = log₂ 1 = 0, with strictness unless qᵢ/pᵢ is constant. Both routes are standard; I checked each derivation line by line rather than quoting it from memory.

In words, without symbols: you cannot beat the code that was designed for the truth. The optimal-code result from Part 1 says the best you can ever do against reality p is H(p) bits per symbol. A code built for some other q is a legitimate code — it obeys Kraft's inequality, it decodes fine — it is just one of the many codes that isn't the best one. So it lands above the floor. Gibbs' inequality is the formal statement that the floor is real and that only the truth touches it.

Why this licenses the loss function. Suppose you have samples from reality p and a model q with knobs on it. Minimising average −log₂ q(observed outcome) over the data is minimising an empirical estimate of H(p,q). Gibbs says that objective's global minimum over all distributions q sits at q = p — and nowhere else. The loss is not merely correlated with truth; truth is its exact argmin. In forecasting language this makes log loss a strictly proper scoring rule. Every practical consequence follows: a model minimising cross-entropy has no incentive to shade its probabilities in any direction, so training pushes toward calibration rather than toward confidence.

The two ways to be wrong are not priced alike

Gibbs tells you the penalty is non-negative. It does not tell you how it is distributed, and that distribution is the practically important part. The per-outcome price is log₂(1/q), and that function is brutally lopsided: it is gentle near q = 1 and unbounded near q = 0. Some numbers to hold:

Probability the model assignedBits charged if it happens · log₂(1/q)
0.51.00
0.252.00
0.13.32
0.016.64
0.0019.97
0.00000119.93
0∞

Each factor-of-10 drop in assigned probability adds a flat log₂ 10 = 3.32 bits. There is no ceiling. Now contrast the two failure modes, both measured against the same truth p = (0.9, 0.1), whose entropy is 0.4690 bits:

Model q₁Failure modeH(p,q)Excess over H(p)
0.9correct0.46900.0000
0.5maximally vague1.00000.5310
0.99overconfident, right direction0.67740.2084
0.999very overconfident, right direction0.99790.5289
0.1confidently wrong3.00492.5359
0.01very confidently wrong5.98095.5119
0.001catastrophically wrong8.96948.5004

Computed by me from H(p,q) = 0.9·log₂(1/q₁) + 0.1·log₂(1/(1−q₁)), bits per symbol, four decimal places.

The shape of that table is the point. Vagueness has a hard ceiling. A model that gives up and spreads its mass uniformly over n outcomes pays exactly log₂ n bits per symbol, no matter what p is — for two outcomes, 1.00 bit, so the worst vagueness can ever cost you here is 1 − H(p) ≤ 1 bit. Even for a 50,000-token vocabulary the uniform model's bill is a finite 15.6 bits. Confident error has no ceiling at all. Assign a thousandth to something that happens a tenth of the time and you are already 8.5 bits over the floor, and the column keeps going. Note also that the two directions of overconfidence are not equal: pushing q₁ to 0.999 in the correct direction costs 0.53 bits, the same order as giving up entirely, because the 10% of the time you are wrong you pay 9.97 bits. Overconfidence hurts wherever it points; it is only ruinous when it points the wrong way.

This asymmetry is why cross-entropy training yields hedged, roughly calibrated probabilities rather than bravado. A model tempted to sharpen a 0.9 into a 0.999 saves 0.14 bits per symbol averaged over the 90% branch, and pays 6.64 extra bits on every occurrence of the 10% branch — 0.66 bits per symbol once weighted. The net trade is a loss of 0.53 bits, and gradient descent can see that. The unbounded left tail of log₂(1/q) acts as an insurance premium against ever saying "impossible" about something that can happen.

Holding the fact in your head for many outcomes

The graph trick only works with one free parameter per distribution, so for the general case Grant goes back to the bar diagram (12:28): one bar per symbol, width pᵢ set by reality, height log₂(1/qᵢ) set by your code, total area equal to H(p,q). Freeze the widths. Now the heights are the only thing you control, and they are not independent — pulling one bar down forces others up, because the qᵢ must still sum to 1. That constraint is exactly what makes the problem interesting: there is a cheapest legal assignment of heights, and it is the one where each height matches its own width, log₂(1/pᵢ).

Two footnotes that the two-outcome picture can hide. First, H(p,q) ≥ H(p) says nothing about H(q). Take q = (½, ¼, ¼), so H(q) = 1.5 bits, and let reality be the point mass p = (1, 0, 0). Then H(p,q) = 1·log₂(1/½) = 1 bit, comfortably below H(q) — and still above H(p) = 0, as it must be. The bound is against the entropy of the distribution in the width slot, always. Second, this is where the constrained-optimisation flavour of the argument comes from: minimising Σ pᵢ·log₂(1/qᵢ) subject to Σ qᵢ = 1 is a textbook Lagrange-multiplier problem, and Grant's video description points at his own Lagrange-multiplier material for readers who want to run it that way. It gives the same answer as the Gibbs argument above, by a longer road.

That is the whole payload of this chapter. One inequality, one picture, and one asymmetry. The next page cashes the inequality in against the opening puzzle: what, exactly, was gzip measuring when it "recognised" a language.

Where people get stuck

"Which distribution goes in which slot?" — the most common and most consequential confusion, made worse by the fact that the notation in the wild is genuinely inconsistent (Grant flags this himself). The reliable mnemonic is the diagram, not the letters: reality sets the widths, the model sets the heights. The distribution you are averaging over — the one supplying the frequencies, the one your data is drawn from — is the width. The distribution supplying the code lengths log₂(1/·) is the height. In training, your data is the width and your network's output is the height. If you find yourself writing H(model, data), you have the arguments backwards and the inequality will point the wrong way.

"The first example gave 1 bit and the model was wrong — did cross-entropy fail?" — no, and this is worth sitting with. A raw cross-entropy number is not a measure of error; it is a bit count. It mixes together the irreducible cost of reality's own randomness, H(p), and the cost of your model being wrong, D(p‖q). Only the second is a mistake. That is why an LLM's loss of, say, 2.1 nats tells you very little on its own about model quality: you don't know how much of it is H(p). Comparisons between models on the same data are meaningful; the absolute number is not.

"log₂(1/q) or −log₂ q?" — identical, and picking one and sticking to it removes a class of sign errors. log₂(1/q) = −log₂ q, and it is non-negative because q ≤ 1. Reading it as "how many halvings from 1 down to q" keeps it manifestly positive and matches the code-length intuition; reading it as "minus log" is more compact but invites a dropped sign. Grant makes the same remark in the recap earlier in this video (04:39), that he half-wishes the field had written it as log₍½₎ q.

"What happens at zero?" — the two zeros behave completely differently. If pᵢ = 0 the term contributes nothing, whatever the model said, because 0·log₂(1/qᵢ) = 0: you are never charged for a possibility that never occurs. But if pᵢ > 0 and qᵢ = 0, cross-entropy is +∞ — one occurrence of something your model declared impossible and the average is destroyed. This is not a mathematical curiosity; it is why every practical implementation clips, smooths, or uses a softmax that cannot output an exact zero, and why the same asymmetry shows up as numerical NaNs in undertrained models.

"Why isn't KL symmetric, if it's a 'distance'?" — because cross-entropy isn't, and the table above shows it numerically: D(p‖q) = 0.531 in one direction and 0.737 in the other for the very same pair. D is a divergence, not a metric — non-negative, zero only when the arguments agree, but neither symmetric nor obedient to the triangle inequality. That is a feature: the question "how badly does a code for q serve reality p" is not the same question as its reverse, and no symmetric quantity could answer both.

Going deeper, verified

Exercises

  1. Redraw the green curve — for p₁ ∈ {0.1, 0.3, 0.5}, plot H(p,q) against q₁ on one set of axes, then overlay the binary entropy function H₂(p₁). A good answer shows three convex curves each touching the entropy curve at exactly one point, that point being q₁ = p₁, and notes that the curves diverge to +∞ at both ends while the entropy curve stays bounded by 1.
  2. Price the tail — take p = (0.9, 0.1) and find, numerically, the value of q₁ > 0.9 at which overconfidence in the right direction costs the same as giving up entirely at q₁ = ½. A good answer lands near q₁ ≈ 0.999 (excess 0.5289 bits versus 0.5310), and explains the mechanism: the gain of 0.13 bits on the 90% branch is swamped by 9.97 bits charged on the 10% branch.
  3. Break the wrong bound — construct a p and q over three outcomes with H(p,q) < H(q), then verify H(p,q) ≥ H(p) still holds. A good answer explains why the second inequality can never be broken while the first has no reason to hold, and identifies the width slot as the one the bound is about. (One worked case is in the section above; find a different one, ideally with p not a point mass.)
  4. Reproduce both tables — write ten lines of code computing H(p), H(p,q) and D(p‖q) for arbitrary discrete distributions, and regenerate every number on this page. A good answer handles pᵢ = 0 correctly (contributes 0) and qᵢ = 0 correctly (returns infinity only when the matching pᵢ is positive).
Previous: P9 Defining cross-entropy · Next: P11 Back to the language trees: what zip was measuring · Back to the map.