ENTROPY // FIELD MAP
← field map
P18 · ENTROPY IN ACTION3Blue1Brown · 00:00–10:52 · 11 min

The bug: when the objective you optimise is not the one you meant

The whole of the Wordle addendum — the correction video Grant posted a week after the main one, retracting a number and keeping the method.

Transcript: this stretch, timestamped

TL;DR — The code that reproduced Wordle's colouring used a shortcut for repeated letters: tag the k-th copy of a letter as a distinct symbol, then colour each square independently. That rule asks "does the answer contain a k-th copy of this letter?" where Wordle asks "is there a copy left after the greens have taken theirs?" The two agree almost always and disagree in one specific shape — a green sitting to the right of a non-green duplicate — which paints a grey square yellow. Measured on the video's own word lists, that is 2.1% of guess/answer pairs, every error in the same direction, and it quietly understated the information content of every guess with a repeated letter. The entropy formulas, the bucket picture and the two-step search were all correct; the environment they were run against was not, and so the reported winner changed. The thing to carry: an optimiser's answer is a claim about your simulator, and only becomes a claim about the world to the extent that the simulator is right.

The previous page walked through the machinery: word frequencies as a prior, expected information as a one-step heuristic, an exhaustive two-step search when one step isn't enough, and the honest admission that "optimal" is defined relative to a word list and an objective. This page is what happened next. A week after that video went up, Grant posted an addendum: the pattern function at the very bottom of the stack — the piece of code that says which squares turn grey, yellow and green — did not match Wordle in a small class of cases involving repeated letters. Everything above it inherited the error. Nothing about entropy changed; the ranking of openers did. It is the most quietly instructive ten minutes in the series, because the failure mode is the one that actually bites people in machine learning: not a wrong formula, but a right formula evaluated against a subtly wrong world.

Outline, with timestamps

Wordle's colouring rule is a matching, not a per-letter test 00:31

Colouring a guess against an answer looks like five independent questions — is this letter here, is it somewhere, is it nowhere — and it is not. It is a two-pass assignment in which each letter of the answer can be used up at most once:

pass 1 (greens):  for each position i
                      if guess[i] == answer[i]:
                          colour[i] = GREEN
                          strike out answer[i] and guess[i]

pass 2 (yellows): for each surviving position i, left to right
                      if some surviving answer letter equals guess[i]:
                          colour[i] = YELLOW
                          strike out that answer letter and guess[i]
                      else:
                          colour[i] = GREY

The striking-out is the whole content of the rule. SPEED against ABIDE: one E in the answer, no greens, so the guess's first E claims it and goes yellow while the second finds the cupboard bare and goes grey — greyness here is a positive statement, "there is no second E". Against ERASE, which holds two Es, both go yellow. Same guess, different colours, because what is rationed is the count in the answer. Structurally it is a matching between guess and answer positions sharing a letter, with positional matches given absolute priority — awkward to vectorise over thirteen thousand words squared, which is why it got shortcut.

The bug, exactly 01:32

The shortcut is still legible in the video's code repository, and it is genuinely clever, which is why it survived. Rather than track consumption, relabel letters so duplicates become distinct symbols — the first E of a word stays e, the second becomes e+26, the third e+52 — and then every square colours independently, with no bookkeeping at all:

# the relabelling: the k-th copy of a letter gets an offset of 26·k
for i in range(n):
    for j in range(i):
        arr[:, i] += (arr[:, i] == arr[:, j]) * 26

# then, per square, with no state carried between positions:
if   guess[i] % 26 == answer[i] % 26:   colour = GREEN    # same base letter, same slot
elif guess[i] in answer:                colour = YELLOW   # tagged: "is there a k-th copy?"
else:                                   colour = GREY

The % 26 in the green test strips the tag, so greens are right. The yellow test does not strip it — that is the trick: the guess's second E goes yellow only if the answer also has a second E. Both of Grant's narrated examples come out correct under it; the shortcut passes every test a person would write by hand in thirty seconds.

What it actually computes is this. For a letter c occurring A times in the answer, the shortcut paints the guess's k-th copy yellow whenever k ≤ A, with k counted in guess order. Wordle paints the r-th non-green copy yellow whenever r ≤ A − G, where G is how many copies came up green. The tags have no idea that the greens already spent some of the answer's supply. So the rules part company in exactly one shape: a green copy of a letter sitting to the right of a non-green copy of the same letter.

Writing · for grey, Y for yellow, G for green — every word below is on the official answer list, so any row can be checked in the browser:

guessanswerWordlethe shortcut
SPEEDABIDE· · Y · Y· · Y · Yagrees
SPEEDERASEY · Y Y ·Y · Y Y ·agrees
ARRAYAGREEG · G · ·G Y G · ·differs
GEESEABIDE· · · · G· Y · · Gdiffers

Take ARRAY against AGREE. The guess's second R lands on the R of AGREE and goes green, consuming the answer's only R, so Wordle greys the first R. The shortcut asks that first R a different question — "does the answer have a first R?" — gets yes, and paints it yellow, telling the player there is another R elsewhere in the word. There is not.

The error has a sign. Every disagreement is grey-in-Wordle, yellow-in-the-simulator; the shortcut can never lose a yellow. One line: if the r-th non-green copy is a true yellow then r ≤ A − G, and its guess-order index is k = r + g with g ≤ G the greens to its left, so k ≤ A − G + g ≤ A and the shortcut calls it yellow too. The two rules differ precisely when g < G — some green sitting to the right. (This argument and the measurements below are mine, not the video's; I checked the claim exhaustively against the repository's word lists — no exception in 636,483 mis-coloured squares.)

How big was it, in numbers 00:00

The video says "a very small percentage of cases" and leaves it there; the size of the effect is worth pinning down, because it is the whole lesson. These are my measurements, from implementing both rules and running them over every pair of words in the lists checked into the video's repository — 12,953 allowed guesses × 2,309 answers, 29.9 million pairs. (Grant quotes 2,315 answers and "13,000" guesses 07:17, so the last digit moves with the list.)

quantityvalue
guess/answer pairs coloured differently633,137 of 29,908,477 — 2.12%
individual squares mis-coloured636,483, of which grey→yellow
… of which any other kind of error0
guesses affected for at least one answer4,639 — exactly the guesses with a repeated letter
… whose expected information changed4,444 lower, 0 higher, 195 unchanged
mean shift in expected information−0.032 bits (worst: COOEE, 4.245 → 4.091)
top ten openers by one-step informationidentical under both rules — none has a repeated letter

That last row sharpens what went wrong. The bug cannot touch a guess with five distinct letters, so it never touched the one-step entropy of SOARE, SLANE, SALET, CRATE, TRACE or CRANE; the damage arrives one level up. The two-step search evaluates thousands of second guesses, plenty of which repeat letters, and marked every one of them down. The full-game simulation is worse off still: there a wrong pattern function does not merely misprice a candidate, it hands the solver a wrong posterior, so the simulated game drifts from the real one move by move. The average score that came out was an honest measurement of a game that is not Wordle.

Why 2% was enough to flip the answer 07:17

Because the quantity being optimised is a mean over 2,315 games and the contenders are separated by less than a hundredth of a guess — Grant calls it "a really tight race" as the leaderboard appears. A ranking is a discrete output of a continuous computation, and its fragility is set not by the size of the error but by the gaps between the things ranked: when the top ten differ in the third decimal place, a perturbation in the third decimal place reorders them. It is the same reason a 0.2-point benchmark gap between two models is not evidence of anything.

And the error is systematic, not noise. Noise of this magnitude would largely cancel across 2,315 games; a bias that always points one way does not. Every repeated-letter guess was marked down, so the comparison was tilted rather than blurred, and no amount of averaging recovers the truth.

What this video reports as the best openers 05:07

The second half of the video is the analysis that never got shown the first time: three objectives, three winners. All three assume what the earlier video deliberately refused — the official answer list, taken as known, uniform over its members. Grant is explicit that this is cheating and that he is doing it on purpose:

"To my taste that feels a bit like overfitting to a test set, and what's more fun is building something that's resilient."— Grant Sanderson, 03:35

With a known uniform answer list the probability computation collapses into counting: sort the 2,315 answers into the 3⁵ = 243 pattern buckets and each pattern's probability is its bucket's share 04:36. The three objectives, in increasing order of cost and of closeness to the thing you actually want:

objectivewhat it computeswinner in this videoH₁, recomputed
one-step informationE[I] = Σ p(pattern)·log₂(1/p(pattern)) over the 243 bucketsSOARE 05:075.885 bits
two-step informationthat, plus the bucket-weighted average of the best second guess's information in each bucketSLANE 06:43 (SOARE falls to 14th)5.769 bits
lowest average scoreplay all 2,315 games with the full solver; average the number of guessesSALET 07:50, with TRACE and CRATE all but tied5.836 / 5.830 / 5.835 bits

The H₁ column is my recomputation of one-step expected information under the corrected rule, uniform over the repository's 2,309 answers — included to show how little separates these words, and that the ordering by information is not the ordering by score. CRANE, the word from the original thumbnail, is 32nd on that metric at 5.741 bits: never the information winner, only the winner of the simulated-score contest in the buggy game. SOARE is an obsolete term for a young hawk, SLANE a turf-cutting spade, SALET a variant of sallet, a light medieval helmet. Grant finds that last too fake to use, and points instead at TRACE and CRATE — near-identical score, real words, both on the answer list 07:50.

The caveat is the shape of the whole page: these are the winners this video reports, for this word list, under these objectives, with a solver of this depth. Change the list — the New York Times had just bought the game — and the result evaporates 09:26. "Optimal" was never a property of the English language.

The general lesson: an optimiser reports on your simulator 02:02

No part of the information theory was implicated. H(p) = Σ p(x)·log₂(1/p(x)) was correct; so was reading it as "how many times do you expect to halve the space of possibilities"; so were the bucket picture, the greedy-versus-two-step comparison, and the insistence that expected information is a heuristic for score rather than score itself. What was wrong sat strictly beneath all of it, in the function mapping (guess, answer) to an observation. The optimiser did its job perfectly on the model it was handed.

That failure mode is endemic in machine learning and it never announces itself: a leaked test set, an eval harness that normalises whitespace differently from the scorer, an off-by-one between tokens and labels, a reward model scoring fluency when you meant correctness, a simulator whose dynamics drift from the deployed ones. Each produces numbers that are precise, reproducible, internally consistent and about the wrong universe. Confidence is not evidence: a stronger optimiser pointed at a mis-specified objective gets you a more confidently wrong answer, faster.

The connection to the scoring-rule material on P13 is sharp. Cross-entropy is a proper scoring rule — uniquely minimised by reporting your true beliefs, which is what licenses training against it. But propriety is a statement about a pair: the score, and the distribution you score against. H(p, q) is minimised at q = p for the p that actually generates your data, so if that data comes from a buggy simulator, the p in the guarantee is the simulator's distribution and not the world's. Minimising a proper score against the wrong p converges beautifully to the wrong beliefs. Propriety buys honesty about your model; it buys nothing about the measurement apparatus.

How would you have caught this 01:32

Concretely, in rough order of cost.

Unit-test the duplicate-letter cases — including the one nobody thinks of. The natural hand-written tests are SPEED/ABIDE and SPEED/ERASE, and the shortcut passes both. What fails is the third shape, a duplicate going green to the right of a non-green copy: ARRAY/AGREE, GEESE/ABIDE, SEVER/OLDER. Enumerate shapes rather than examples — (guess repeats a letter) × (answer has 0, 1, 2 copies) × (green on the first copy / on a later copy / none). A dozen cells, one of which is the bug.

Differential-test the fast version against a slow, obviously-correct one. Highest value on the list, and nearly free. Write the two-pass rule the boring way, with a list you strike letters out of, then run both over every pair of words — 30 million pairs is under a minute in NumPy — and assert equality. Every number on this page came out of that exercise. When you replace a readable implementation with an optimised one, keep the readable one as the oracle; the reason to trust the fast path is that the slow path exists.

Assert invariants that need no oracle. Property-based testing catches this class without your knowing what the bug is. Colouring a word against itself must be all green. The green-plus-yellow count for a letter c can never exceed the copies of c in the answer — that one property fails on ARRAY/AGREE, and is the bug. Best of all: the answer must survive its own pattern, so filtering the word list by pattern(guess, answer) must always leave the answer standing. Cheap, oracle-free, and violated the moment your filter and your colourer disagree.

Check against the real game, and against a second implementation. Play a few hands in the browser with deliberately duplicate-heavy guesses and diff the colours: ten minutes, and it is the only test that validates your reading of the specification rather than your implementation of your reading. Then, when a result is the headline — "the best opener is X" — re-derive it a second way, or against someone else's published solver, before publishing. Anything hinging on the third decimal place earns that routinely.

Invalidate your caches. A trap specific to this design and to much ML tooling: the pattern function's output is precomputed once into a matrix on disk and thereafter only looked up. Fix the source and the stale artefact silently outlives the bug. Cache keys want a hash of the code that produced them — otherwise the fix ships and the numbers don't move, which is worse than the original error.

Where people get stuck

"So Wordle's rule is just per-letter presence?" No, and this is the misconception the bug is made of. Presence is a property of the letter; the colouring rations copies. The clean mental model is a supply-and-demand ledger: the answer supplies some number of each letter, greens are served first with absolute priority, and the leftovers are handed to non-green positions from left to right. Grey means "no copy left for you", not "this letter does not occur".

"If the pattern function was wrong, weren't the entropies wrong too?" Only for guesses that repeat a letter — 4,639 of the 12,953 allowed words, and none of the top openers. The reason the published conclusion moved anyway is that the two-step search and the score simulation both consult the pattern function thousands of times per candidate, on words that do repeat letters. The moral is that a bug's blast radius is set by how many layers sit on top of it, not by how often it fires.

"Which of SOARE, SLANE and SALET is the answer?" The question is malformed until you name an objective. They are winners of three different competitions — most expected information after one guess, after two, and lowest average number of guesses in full simulated play. The three disagree because expected information is a proxy for score and proxies are not the thing. That gap is itself the most transferable idea in the video: greedy optimisation of a one-step heuristic is not global optimisation, and the honest way to find out how big the gap is, is to run the real objective.

"The captions spell it 'soar' and 'salé'." The words are soare, slane and salet — all five letters, all present in the allowed-guess list checked into the video's code repository, which is the check to run when a caption spells an unfamiliar word by ear.

Going deeper, verified

Exercises

  1. Write both colourers and diff them — implement the two-pass rule with explicit striking-out, and the tag-by-26 shortcut. Run both over the word lists from the repository and report the number of disagreeing pairs. A good answer gets a figure near 2% of pairs, notes that the affected guesses are exactly those with a repeated letter, and — the real prize — verifies the direction claim: every mis-coloured square is grey in Wordle and yellow in the shortcut, no exceptions.
  2. Price the bug in bits — for one repeated-letter guess, say COOEE, compute the 243-bucket distribution over the answer list under both rules and evaluate E[I] = Σ p·log₂(1/p) for each. A good answer reports both numbers (mine: 4.245 bits correct, 4.091 buggy) and explains the sign: the spurious yellows merge answers that Wordle would have separated more often than they split ones it would have joined, so the guess looks flatter and less informative than it is.
  3. Find the smallest failing case by hand — without running any code, construct a guess and an answer, both real English words, where the shortcut and Wordle disagree, and justify it from the rule k ≤ A versus r ≤ A − G. A good answer names the required shape — a repeated letter whose later copy is green and whose earlier copy has no supply left — and gives a pair such as ARRAY/AGREE, together with the two colourings and which square differs.
Previous: P17 Priors, two-step search, and what "optimal" actually means · Next: P19 Reinventing Hamming codes · Back to the map.