ENTROPY // FIELD MAP
← field map
CONCEPT INDEXevery term, one place

Concept index

The vocabulary of the series, defined tightly, each pointing at the page that earns it
How to use this — the pages tell the story in order; this page is for when you already know the story and need the definition. Everything is in bits (base-2 logarithms) unless said otherwise. Throughout, p is the distribution reality draws from and q is the distribution your model asserts — keeping those two straight is most of the subject.

The five that matter

TermDefinitionReads asPage
Information content
surprisal
I(x) = log₂(1/p(x)) = −log₂ p(x) How surprised you should be by one outcome, in bits. Also the length of the code word an optimal encoder would spend on it. P4
Entropy
H(p)
H(p) = Σ p(x)·log₂(1/p(x)) Average surprise. The floor on bits per symbol that no encoder can beat, and the yardstick for every compressor. P6
Cross-entropy
H(p, q)
H(p, q) = Σ p(x)·log₂(1/q(x)) What you actually pay: q's prices at p's frequencies. Always at least H(p), equal only when q = p. P9 · P10
KL divergence
D(p‖q)
D(p‖q) = H(p, q) − H(p) = Σ p(x)·log₂(p(x)/q(x)) The surcharge for being wrong — the bits you paid over the unavoidable minimum. Non-negative, zero only at q = p, asymmetric. P15
Perplexity PP = 2^H (or e^H if H is in nats) Entropy exponentiated back into "number of options". A perplexity of 20 means the model is as uncertain as someone choosing uniformly among 20 things. P12
The one identity to carry: cross-entropy = entropy + KL. The first term is the cost reality imposes and no model can remove; the second is the cost your model adds. Training can only ever attack the second, which is why minimising cross-entropy and minimising KL are the same optimisation.

Coding and compression

Where it lands in machine learning

The other side of the channel

Names and papers

Back to the field map, or start at P1.