3BLUE1BROWN · 2020–20266 videos · 2 h 25 min · 20 pages here
Entropy, mapped
Grant Sanderson's Compression is Intelligence series rebuilds information theory from one question — how short can you make a message? — and walks it all the way to the loss function every language model is trained on. This map splits the two core videos chapter by chapter, adds the Wordle videos where entropy gets used on a real problem and the Hamming videos where redundancy gets added back on purpose, and links every claim into the video at the second it is made.
Tick a page when you have worked through it; this browser remembers.
00
The spine
Six ideas, in the order the series earns them. If you read nothing else, read this list — everything below is one of these six unpacked at length.
- 1 · SurpriseA message carries information in proportion to how unlikely it was. Something you already expected tells you nothing. The amount of surprise in an outcome of probability p turns out to be forced to be log₂(1/p) bits — not chosen, forced, once you insist that independent surprises add. P4
- 2 · CodesA short code word is a bet that its symbol is common. Prefix-free codes cost you a budget that must sum to 1, so making one word shorter makes another longer. The best you can do is spend log₂(1/p) bits on a symbol of probability p. P2 · P3
- 3 · EntropyEntropy is the average surprise: H(p) = Σ p(x)·log₂(1/p(x)). It is the floor on bits per symbol that no encoder can beat, and the yardstick against which every compressor is judged. P6
- 4 · Cross-entropyCross-entropy is what you actually pay when your model of the world is wrong. Encode with a code built for q while reality is p and you pay H(p,q) = Σ p(x)·log₂(1/q(x)) — never less than H(p), and equal only when q = p. P9 · P10
- 5 · KL divergenceThe gap between them is the fine you pay for your wrong beliefs: D(p‖q) = H(p,q) − H(p). Zero if and only if you were right, asymmetric, and not a distance. P15
- 6 · Where the loss comes fromTraining a language model minimises cross-entropy, which is the same thing as building the best possible text compressor. That is not an analogy — it is the same number, and it is why the objective is the log of the probability of the right token and nothing else. P12 · P13
01 · PART 1 · 32 min
Reinventing entropy
A robot on a distant moon needs instructions in as few bits as possible. Six pages take that toy problem to Shannon's definition of entropy, then point it at English and ask how compressible a language actually is. Watch Part 1 ↗ · transcript
P1
The thesis of the whole series: a good predictor and a good compressor are the same object, so "how well does this model understand text" and "how small can it make a file" are one question asked twice.
P2
Encode up/down/left/right with probabilities ½, ¼, ⅛, ⅛. The naive answer is 2 bits; the clever answer is 1.75. Why the receiver can still tell where one code word ends, and what constraint that imposes.
P3
Code lengths are a zero-sum budget: shortening one word lengthens another. The signature of an optimal code is that the output bitstream looks like fair coin flips — no pattern left to exploit.
P4
The definition is not a convention. Demand that the information in two independent events adds, and that certainty carries zero, and the logarithm is the only function left standing.
P5
English is not a stream of independent letters, so counting letter frequencies badly overestimates its entropy. How Shannon measured the real number using human predictions, and why the answer keeps falling as your model gets better.
P6
Average the surprise and you have entropy: a hard floor on bits per symbol, reachable in the limit and beatable never. What the quantity does and does not have to do with thermodynamics.
02 · PART 2 · 34 min
Cross-entropy, and where the loss function comes from
Nine pages on the quantity you get when your model of the world is wrong: what it measures, why it is the training objective for every language model, why no other objective would do, and the divergence hiding inside it. Watch Part 2 ↗ · transcript
P7
Append a document to a corpus, zip it, and see how much it grew. That number alone recovers the tree of European languages — the hook that the rest of the video explains.
P8
The half of Part 1 you need in cache before cross-entropy makes sense, restated compactly and with the fractional-bit caveat spelled out.
P9
Two distributions, two roles: one supplies the code words, the other supplies the frequencies. Which is which is the entire content of the definition, and the source of every sign error people make with it.
P10
Concrete numbers for a mismatched code, the punishment curve for confident wrong predictions, and the inequality that makes cross-entropy usable as a loss at all.
P11
The 2002 result decoded: gzip's dictionary is a crude language model, and the extra bytes are an estimate of the cross-entropy of one language under another's model.
P12
The line of PyTorch every language model is trained with, read as information theory: the loss is bits per token, the model is a compressor, and gradient descent is shopping for a shorter code book.
P13
Cross-entropy is uniquely minimised by telling the truth about your uncertainty. The constrained-optimisation argument for that, and what goes wrong with the obvious alternatives.
P14
Swap the one-hot target for a big model's full distribution and the same formula becomes distillation. Why a teacher's second-best guesses teach more than its best one.
P15
Subtract the unavoidable cost from the cost you paid and what is left is the price of being wrong. Why it is asymmetric, why that is a feature, and where you meet it outside compression.
03 · 2022 · 41 min
Entropy in action: Wordle
The 2022 pair, where the same definitions get pointed at a real optimisation problem — and then at the author's own bug. This is where entropy stops being a formula and becomes a decision rule. Watch ↗ · transcript
P16
Every guess splits the remaining words into buckets by colour pattern. The entropy of that split is the expected number of bits you learn, which turns "which word should I play" into an arithmetic problem.
P17
Maximising information is a proxy, not the goal. Adding a prior over which words are plausible answers and a second ply of search, and the gap between the proxy and the real objective.
P18
A week later, a confession: a small bug in the pattern-matching code had been distorting every number. What changed, what did not, and why this is the most useful ten minutes in the pair.
transcript
04 · 2020 · 37 min
The other direction: redundancy on purpose
Compression removes redundancy. Error correction puts it back, deliberately and efficiently. Shannon's two theorems are the two ends of the same channel, and Hamming codes are the cleanest place to see the second one. Watch ↗ · transcript 1 · transcript 2
P19
Eleven message bits, four parity bits, and a scheme where the failing parity checks spell out the index of the flipped bit in binary. Derived rather than stated.
P20
The whole encoder and decoder collapse to an XOR over positions. Why that is the same algorithm as the parity groups, seen from a different angle.
05
Work it out yourself
Each page sets its own exercises. These are the ones worth doing even if you only skim.
- Build the robot's encoder. Take the ½, ¼, ⅛, ⅛ distribution from P2, implement the prefix code, encode ten thousand samples and measure the bits per symbol. Confirm you get 1.75 and not 2. Then change the distribution to be uniform and watch the advantage vanish.
- Measure the entropy of English yourself. Take any book from Project Gutenberg. Compute the entropy of the letter distribution, then of letter pairs given the previous letter, then of triples. Watch the bits-per-character number fall with each order, and compare against the estimates in P5.
- Show the compressor and the model are the same object. Take any small language model, compute its average cross-entropy in bits per character on a held-out text, multiply by the character count, and compare against what gzip, xz and zstd produce on the same file. Details in P12.
- Play the Wordle solver. Implement the entropy-of-the-split scoring from P16 over the official word list, look at your top openers, then re-read P18 and check your pattern-matching code against the duplicate-letter cases that bit Grant.
- Convince yourself cross-entropy is a proper scoring rule. Fix a true distribution over three outcomes, then grid-search over reported distributions and plot the expected loss. Confirm the minimum sits exactly at the truth, and that squared error does not have the same property in the same way. Argument in P13.
06
Sources
Reinventing Entropy3Blue1Brown, 7 Jun 2026, 32 min — Compression is Intelligence, Part 1
But what is cross-entropy?3Blue1Brown, 16 Jul 2026, 34 min — Compression is Intelligence, Part 2
Solving Wordle using information theory3Blue1Brown, 6 Feb 2022, 31 min
The best Wordle opener is not "crane"3Blue1Brown, 13 Feb 2022, 11 min — the addendum
But what are Hamming codes?3Blue1Brown, 4 Sep 2020, 20 min
Hamming codes, part 23Blue1Brown, 4 Sep 2020, 17 min
A Mathematical Theory of CommunicationClaude Shannon, 1948 — the paper the series rebuilds
Prediction and Entropy of Printed EnglishClaude Shannon, 1951 — the guessing-game measurement in P5
Visual Information TheoryChris Olah, 2015 — the visualisation Grant credits in the description
The Wordle solver code3b1b/videos — the actual scripts behind P16–P18
Bayes' theorem, the geometry of changing beliefs3Blue1Brown, 2019 — not a unit here, but the prior/update background for P17
CS336, mappedwhere the cross-entropy of P12 becomes a training run you write yourself