ENTROPY // FIELD MAP
← all maps
3BLUE1BROWN · 2020–20266 videos · 2 h 25 min · 20 pages here

Entropy, mapped

Grant Sanderson's Compression is Intelligence series rebuilds information theory from one question — how short can you make a message? — and walks it all the way to the loss function every language model is trained on. This map splits the two core videos chapter by chapter, adds the Wordle videos where entropy gets used on a real problem and the Hamming videos where redundancy gets added back on purpose, and links every claim into the video at the second it is made.

Tick a page when you have worked through it; this browser remembers.

00

The spine

Six ideas, in the order the series earns them. If you read nothing else, read this list — everything below is one of these six unpacked at length.

01 · PART 1 · 32 min

Reinventing entropy

A robot on a distant moon needs instructions in as few bits as possible. Six pages take that toy problem to Shannon's definition of entropy, then point it at English and ask how compressible a language actually is. Watch Part 1 ↗ · transcript

P1
▶ 00:003 minprediction ≡ compression · why cross-entropy shows up in pre-training
The thesis of the whole series: a good predictor and a good compressor are the same object, so "how well does this model understand text" and "how small can it make a file" are one question asked twice.
P2
▶ 03:287 minvariable-length codes · the prefix-free property · decoding as walking a tree
Encode up/down/left/right with probabilities ½, ¼, ⅛, ⅛. The naive answer is 2 bits; the clever answer is 1.75. Why the receiver can still tell where one code word ends, and what constraint that imposes.
P3
▶ 10:464 minthe code-length budget · why a good code makes every bit a coin flip
Code lengths are a zero-sum budget: shortening one word lengthens another. The signature of an optimal code is that the output bitstream looks like fair coin flips — no pattern left to exploit.
P4
▶ 14:473 minsurprise · additivity over independent events · the bit as a unit
The definition is not a convention. Demand that the information in two independent events adds, and that certainty carries zero, and the logarithm is the only function left standing.
P5
▶ 17:407 minletter frequencies · context · Shannon 1951 · bits per character
English is not a stream of independent letters, so counting letter frequencies badly overestimates its entropy. How Shannon measured the real number using human predictions, and why the answer keeps falling as your model gets better.
P6
▶ 24:297 minH(p) = Σ p log(1/p) · the noiseless coding theorem · von Neumann's naming joke
Average the surprise and you have entropy: a hard floor on bits per symbol, reachable in the limit and beatable never. What the quantity does and does not have to do with thermodynamics.
02 · PART 2 · 34 min

Cross-entropy, and where the loss function comes from

Nine pages on the quantity you get when your model of the world is wrong: what it measures, why it is the training objective for every language model, why no other objective would do, and the divergence hiding inside it. Watch Part 2 ↗ · transcript

P7
▶ 00:003 minBenedetto–Caglioti–Loreto 2002 · compression as a similarity measure
Append a document to a corpus, zip it, and see how much it grew. That number alone recovers the tree of European languages — the hook that the rest of the video explains.
P8
▶ 03:022 minlog(1/p) code lengths · entropy as the floor · the one-slide version of Part 1
The half of Part 1 you need in cache before cross-entropy makes sense, restated compactly and with the fractional-bit caveat spelled out.
P9
▶ 05:203 minH(p,q) = Σ p log(1/q) · true distribution vs. model distribution · asymmetry
Two distributions, two roles: one supplies the code words, the other supplies the frequencies. Which is which is the entire content of the definition, and the source of every sign error people make with it.
P10
▶ 08:265 minworked examples · confident and wrong · why H(p,q) ≥ H(p)
Concrete numbers for a mismatched code, the punishment curve for confident wrong predictions, and the inequality that makes cross-entropy usable as a loss at all.
P11
▶ 12:592 minthe compressor as a model · cross-entropy as a similarity score
The 2002 result decoded: gzip's dictionary is a crude language model, and the extra bytes are an estimate of the cross-entropy of one language under another's model.
P12
▶ 14:556 minnext-token prediction · softmax · one-hot targets · training as compression
The line of PyTorch every language model is trained with, read as information theory: the loss is bits per token, the model is a compressor, and gradient descent is shopping for a shorter code book.
P13
▶ 20:386 minproper scoring rules · the Lagrange multiplier argument · honesty as the optimum
Cross-entropy is uniquely minimised by telling the truth about your uncertainty. The constrained-optimisation argument for that, and what goes wrong with the obvious alternatives.
P14
▶ 26:134 minsoft targets · teacher and student · why the wrong answers carry the signal
Swap the one-hot target for a big model's full distribution and the same formula becomes distillation. Why a teacher's second-best guesses teach more than its best one.
P15
▶ 31:352 minD(p‖q) = H(p,q) − H(p) · not a distance · where it shows up elsewhere
Subtract the unavoidable cost from the cost you paid and what is left is the price of being wrong. Why it is asymmetric, why that is a feature, and where you meet it outside compression.
03 · 2022 · 41 min

Entropy in action: Wordle

The 2022 pair, where the same definitions get pointed at a real optimisation problem — and then at the author's own bug. This is where entropy stops being a formula and becomes a decision rule. Watch ↗ · transcript

P16
▶ 00:0018 minpattern distributions · entropy of a guess · greedy one-step search
Every guess splits the remaining words into buckets by colour pattern. The entropy of that split is the expected number of bits you learn, which turns "which word should I play" into an arithmetic problem.
P17
▶ 18:1512 minword-frequency priors · uncertainty vs. expected score · measured performance
Maximising information is a proxy, not the goal. Adding a prior over which words are plausible answers and a second ply of search, and the gap between the proxy and the real objective.
P18
▶ 00:0011 minthe addendum video · the corrected opener · what the error teaches
A week later, a confession: a small bug in the pattern-matching code had been distorting every number. What changed, what did not, and why this is the most useful ten minutes in the pair. transcript
04 · 2020 · 37 min

The other direction: redundancy on purpose

Compression removes redundancy. Error correction puts it back, deliberately and efficiently. Shannon's two theorems are the two ends of the same channel, and Hamming codes are the cleanest place to see the second one. Watch ↗ · transcript 1 · transcript 2

P19
▶ 00:0020 minparity · binary search over parity groups · (15,11) and the extended code
Eleven message bits, four parity bits, and a scheme where the failing parity checks spell out the index of the flipped bit in binary. Derived rather than stated.
P20
▶ 00:0017 minXOR of the indices of the on bits · one algorithm, several perspectives
The whole encoder and decoder collapse to an XOR over positions. Why that is the same algorithm as the parity groups, seen from a different angle.
05

Work it out yourself

Each page sets its own exercises. These are the ones worth doing even if you only skim.

  1. Build the robot's encoder. Take the ½, ¼, ⅛, ⅛ distribution from P2, implement the prefix code, encode ten thousand samples and measure the bits per symbol. Confirm you get 1.75 and not 2. Then change the distribution to be uniform and watch the advantage vanish.
  2. Measure the entropy of English yourself. Take any book from Project Gutenberg. Compute the entropy of the letter distribution, then of letter pairs given the previous letter, then of triples. Watch the bits-per-character number fall with each order, and compare against the estimates in P5.
  3. Show the compressor and the model are the same object. Take any small language model, compute its average cross-entropy in bits per character on a held-out text, multiply by the character count, and compare against what gzip, xz and zstd produce on the same file. Details in P12.
  4. Play the Wordle solver. Implement the entropy-of-the-split scoring from P16 over the official word list, look at your top openers, then re-read P18 and check your pattern-matching code against the duplicate-letter cases that bit Grant.
  5. Convince yourself cross-entropy is a proper scoring rule. Fix a true distribution over three outcomes, then grid-search over reported distributions and plot the expected loss. Confirm the minimum sits exactly at the truth, and that squared error does not have the same property in the same way. Argument in P13.
06

Sources