Data 2
Transcript: cleaned auto-captions with timestamps
This is the second half of the data unit, and it is deliberately the least glamorous lecture in the course. Lecture 13 was a tour of datasets — what went into BERT, what went into the Pile, what went into Dolma — and its punchline was that data does not fall from the sky: someone crawls a live service, someone converts HTML to text, someone filters, someone dedupes. Lecture 14 opens the black box on those last two steps. Percy is explicit that the point is mechanics, not intuition: by the end you know the algorithms that every open pretraining pipeline runs, and you still have no idea what "good data" means, because that only comes from staring at the data yourself. It is also the last lecture before the course pivots to alignment, so it is the last time the course is about what goes into the model rather than what comes out.
Outline, with timestamps
- 00:04 — Where lecture 13 left off: the crawl → text → filter → dedupe pipeline, and which two steps get opened today.
- 01:14 — The filtering primitive: given target T and raw R, find T′ ⊂ R that looks like T. Two requirements: generalize, and be fast.
- 02:20 — KenLM: Kneser-Ney n-grams as a scorer, live perplexity demos, and CCNet's keep-the-top-third rule.
- 08:34 — fastText: why a factorized linear classifier over hashed n-grams beat the neural models of 2016, and still wins here.
- 13:14 — DSIR: importance resampling, and why matching a distribution is not the same as classifying membership in it.
- 18:59 — One framework, three scoring functions — plus the student question about n-grams only seeing local context.
- 23:31 — Language identification, and OpenWebMath as a case study: lid.176's failure modes, then treating "math" as a language for 14.7B tokens that beat 20× more generic data.
- 29:52 — Quality filtering in the wild: GPT-3, LLaMA, phi-1 — and the turn where the target set stops being a corpus and becomes a prompt.
- 35:00 — Toxicity filtering: the Jigsaw dataset, and Dolma's two fastText heads.
- 36:47 — Duplicates: mirrors, licenses, templated spam, and the paragraph that appears 61,036 times in C4.
- 41:28 — The design space (item / match / action), the linear-time constraint, and exact dedup by hash grouping.
- 46:42 — Bloom filters: the k-hash trick, and the false-positive analysis that gives the optimal k.
- 58:34 — Jaccard and MinHash: making collision probability equal similarity, and the permutation argument for why.
- 64:52 — LSH bands: sharpening a probability into a threshold, tuning b and r, and the closing Q&A on semantic dedup.
The one primitive, and the budget that constrains it
Strip away the paper names and every filtering step in every open pretraining pipeline has the same shape. You have T, a target set: a modest pile of text you wish you had more of. You have R, a raw set: Common Crawl, or a whole GitHub dump, orders of magnitude bigger and mostly junk. You want T′ ⊆ R that resembles T. Two requirements fall out immediately (01:14): it has to generalize — returning exactly T would be pointless, you already have T — and it has to be fast.
"Fast" is not a vague preference here, and the arithmetic is the most useful thing in the first half of the lecture. Suppose you filter the web down to 1%: your scorer touches 100 tokens for every 1 token you eventually train on. If the scorer were a transformer forward pass, you would spend roughly as much compute deciding what to train on as training — at which point, Percy notes, you may as well have trained on everything (12:38). So the filter's per-token cost must sit about two orders of magnitude below a training step's. That single inequality is why this lecture is about counting n-grams and hashing bigrams rather than anything with attention in it. You can use BERT or Llama as your quality classifier; you just have to earn the flops back, and usually you can't.
Three scorers, one skeleton
KenLM — a generative model of T. Fit a Kneser-Ney-smoothed n-gram model to the target text and score raw documents by their perplexity under it (02:20). Fitting is literally counting n-grams and normalizing; smoothing exists because raw counts are zero for most plausible n-grams, and Kneser-Ney backs off to shorter contexts to cover them. KenLM is the standard implementation, inherited from statistical machine translation, and popular here for no deeper reason than that it is what everyone reached for.
The live demo is worth watching precisely because it goes badly. A Wikipedia-trained model gives a Wikipedia sentence about Stanford's founding a perplexity around 187; the CS336 regrade policy scores worse, which is reasonable; "asdf asdf asdf" worse still, which is what you want; and then "the the the the the…" scores low (06:44) — exactly the failure you were hoping it wouldn't have. Percy's response is the honest one: it doesn't matter much, because this is a crude filter for removing true nonsense, not a quality oracle. CCNet built LLaMA's web data this way — sort paragraphs by increasing perplexity, keep the best third, ship it.
fastText — a discriminative classifier. The 2016 paper's contribution was mostly an engineering observation: for text classification an almost-linear model matched the neural architectures of the day and ran orders of magnitude faster. The trick is a factorization (09:43). Naive bag-of-words needs a V×K matrix, huge and sparse; fastText maps the vocabulary into a small hidden dimension H (16 in the lecture's example) and classifies from there, costing H·(V+K) parameters with no nonlinearity anywhere in the forward pass. To get n-gram features without an unbounded feature space it hashes each n-gram into a fixed number of bins — 10 million in practice. Collisions are simply tolerated: two unrelated bigrams sharing a bin get a weight that averages them, and the loss accounts for it. For filtering K = 2, so the whole thing degenerates to a fast binary linear classifier over hashed features.
DSIR — a ratio, and resampling. The third option treats the problem as distribution matching rather than classification (13:14). Importance resampling is the standard Monte Carlo move: draw from the proposal q, weight each sample by p(x)/q(x), normalize, resample. The obstacle here is that the target set Dp is small by construction — its smallness is the entire reason you are doing this — so you cannot fit a good model to it. DSIR's answer is fastText's hashing trick again: fit a bag-of-hashed-n-grams unigram model to each side, score by the ratio, resample proportionally. The reported GLUE gains over a heuristic classifier are real but modest, and Percy says so.
What makes this section a framework rather than a list is the collapse at 18:59. Fit something from T and R, derive a score, keep the top of R. The three methods differ only in the score:
| generative (KenLM) | score(x) = pT(x) | keep above a threshold, stochastically; R is never used |
| discriminative (fastText) | score(x) = p(T | x) | keep above a threshold; uses both sides, but only as labels |
| importance (DSIR) | score(x) = pT(x) / pR(x) | resample with probability ∝ score; targets the distribution, not the boundary |
That last row is the conceptual payoff. A classifier answers "is this document in the target region", and a set of documents can all be confidently in-region while collectively being far less diverse than T was; the ratio, used as a resampling weight, tries to reproduce T's shape. Whether that matters at scale is unsettled — the empirical gap is small — but it is why two methods that look interchangeable are not.
A student pushes on the obvious weakness at 21:16: an n-gram scorer only sees local context, so a document can be locally fluent and globally nonsense, and is trivially gameable. Percy concedes the construction — he had just shown one himself — and reframes the goal: these filters exist to remove true garbage, and on average that is enough. Hold onto that, because it is what stops you over-trusting a quality score later.
The same machinery, pointed at four different jobs
Language identification. The off-the-shelf fastText lid.176 model covers 176 languages and was trained on multilingual sites — Wikipedia plus a translation corpus and a Southeast European news corpus. Dolma simply keeps pages with p(English) ≥ 0.5. The reason to filter by language at all is the compute argument again (24:37): in a compute-limited regime, every token spent on another language is a token not spent on yours. BLOOM's corpus was only about 30% English and its English performance suffered accordingly. Frontier models are heavily multilingual — but they are not compute-limited in the same way, and get positive transfer instead of interference.
The demo again earns its place by misbehaving (26:24). "The quick brown fox…" gets only ~0.71 English; duplicating it doesn't move the probability, which is correct — saying something twice does not make it more English. LaTeX comes back weakly English, a line of C++ comes back Russian, a bare "Hello!" comes back Italian. The stated caveats are the ones that bite in production: short strings, low-resource languages, dialects of English filtered out as non-English, similar pairs like Malay and Indonesian, and code-switching, where ground truth isn't even well defined. A downloadable classifier that everyone uses is not thereby a correct classifier.
Math as a language. OpenWebMath is the clean case study (28:42): rule-based prefilters, then KenLM trained on ProofPile with a perplexity cutoff, then a fastText classifier for mathematical writing with deliberately asymmetric thresholds — a low bar (0.17) for pages the rules already flagged as math, a high bar (0.8) for pages they didn't. Hacky, and it works: 14.7B tokens that train 1.4B models to beat models trained on 20× more unspecialized data. This is the lecture's strongest argument for filtering as a first-class technique rather than hygiene — if you know the domain you care about, going and getting more of it beats hoping the general crawl contains enough.
Quality. Some pipelines (C4, Gopher, RefinedWeb, FineWeb, Dolma) deliberately avoid model-based quality filtering; others (GPT-3, LLaMA, DCLM) embrace it, and the second camp is becoming the norm. GPT-3 took positives from its non-Common-Crawl sources — Wikipedia, WebText2, Books1, Books2 — negatives from Common Crawl, trained a linear classifier over word features, and kept documents stochastically against a Pareto-distributed threshold rather than a hard cut. LLaMA used a sharper target: pages referenced by Wikipedia rather than Wikipedia itself, a proxy for "the kind of page an encyclopedia would cite".
phi-1 is where the shape of the problem changes (32:39). Its target set did not exist anywhere. The authors prompted GPT-4 to rate the educational value of ~100K documents from the Python subset of The Stack, used those ratings as labels, trained a random forest over embeddings from a pretrained code model, and ran that over the full raw set. On HumanEval, a 1.3B model on raw Stack-Python reached 12.19% after 96K steps; the same architecture on the filtered subset hit 17.68% after 36K — better score, a third of the steps.
Toxicity. Same machinery, different labels (35:00). The Jigsaw Toxic Comments dataset — Wikipedia talk-page comments annotated as toxic, severe toxic, obscene, threat, insult, identity hate — came out of a project on improving online discussion. Dolma trains two fastText heads on it, one for hate and one for NSFW. Nothing new algorithmically; that is the point of the section.
Deduplication is a structurally different problem
Filtering is a unary operation: look at a document, emit a score, decide. Deduplication is pairwise — you cannot tell whether a document is a duplicate without reference to other documents — and pairwise over a web-scale corpus is quadratic, a non-starter (42:36). The back half of the lecture is one idea worked out three ways: use hashing to turn a pairwise question into a per-document one, so the algorithm runs in linear time.
Why bother: the web is full of exact duplicates for structural reasons — mirrors, forks — and Common Crawl cannot know that twelve URLs serve the same Project Gutenberg page. Near-duplicates come from licenses pasted everywhere, from templated content where an ad generator swapped "Canada" for "USA", and from copy-paste that lost a comma. The number that lands is a single product-description paragraph appearing 61,036 times in C4 (39:37). The text is fine English. You just do not want 61,036 epochs of it.
"Quality filtering says this piece of data, I don't want to ever train on it. Deduplication says, well, this might be fine, but I only want a small number of these rather than 61,000 copies."— Percy Liang, 40:12
The measured benefits are two: fewer tokens for the same information, so training is cheaper; and less memorization, which mitigates copyright and privacy exposure, since a model that saw a passage once is far less likely to regurgitate it. Dedup is a partial mitigation, not a solution.
Any dedup scheme is three choices (41:28): what is an item (sentence, paragraph, document), what counts as a match (exact, or a fraction of shared subitems), and what action you take (remove all copies, or all but one). Exact dedup then falls out in three lines of MapReduce: hash every item, group by hash, keep one per group — simple, high precision, trivially parallel. C4 does exactly this over three-sentence spans, and Percy flags a consequence nobody seems to mind: excise a duplicated span from the middle of a document and what remains may not be coherent prose any more (46:03).
Bloom filters: buying accuracy with hash functions, not memory
A Bloom filter is an approximate set-membership structure: a bit array plus k hash functions, memory-efficient, updatable, no deletes, and one-sided in its error — "no" is always true, "yes" is usually true (46:42). Insert by setting the k bits an item hashes to; query by checking that all k are set.
The instructive part is the analysis, because the naive lever is the wrong one. With one hash function and m bins, the false-positive rate falls like 1/m — polynomial in memory, so a 10-10 rate costs an absurd array. Adding hash functions is exponentially better. Fix a test bin i and an item not in the set; the probability that a single insertion with k hashes misses i is (1 − 1/m)k, so after n insertions and accounting for the k chances the query itself has to miss:
f = ( 1 − (1 − 1/m)k·n )k
With m = 1000, n = 100 and k = 10 that is about 0.010, down from 0.63 with a single hash (55:37). Optimizing over k for a fixed m/n ratio gives k* = ln2 · m/n — here 6.93, not 10 — at which point exactly half the bits are set and f = 0.5k* ≈ 0.008. So k is genuinely non-monotonic, as a student notices at 57:23: too few hashes and each query is under-constrained; too many and you saturate the array. The practical reading is a three-way dial between compute (k), memory (m), and false-positive rate (f). Dolma sets f = 10-15 and runs its exact dedup over paragraphs.
MinHash and LSH: collisions you want
Exact dedup misses everything that is morally a duplicate but differs by a comma. To catch those you need a similarity measure, and the one used is Jaccard: |A ∩ B| / |A ∪ B| over the two documents' item sets. Two documents are near-duplicates if Jaccard exceeds a threshold — typically very high, ~0.99, meaning roughly one word in a hundred may differ. Nothing about this is semantic: drop the word "not" and the meaning inverts while the Jaccard barely moves (59:39). It is a superficial-similarity test by design.
The algorithmic problem is that Jaccard is pairwise. MinHash is the bridge: a family of hash functions with the property Pr[minhash(A) = minhash(B)] = Jaccard(A, B). Compute it by hashing every element of a set and keeping the minimum.
"Normally, you think of hash collisions as to be avoided at all costs… but here, it's not that you want more hash collisions, it's just that hash collisions are something that you want to control for. You want just the right level of hash collisions governed by the similarity."— Percy Liang, 60:46
The proof is a one-liner once you see it (61:51). A random hash function induces a uniformly random permutation of the universe of items. Walk the permutation and ask which item of A ∪ B you hit first; every item is equally likely to be first. If that first item lies in A ∩ B, then it is the argmin for both sets and the MinHashes agree. If it lies in only one of them, they disagree. So the collision probability is |A ∩ B| / |A ∪ B| exactly — the Jaccard, by construction.
That alone is not enough: a single collision tells you similarity was probably highish, not that it cleared a threshold. Locality sensitive hashing sharpens it with an and/or structure (64:52). Take n = b·r MinHash functions, split them into b bands of r each, and declare A and B candidates if all r hashes agree within some band. A fixed band matches with probability sr, so:
P[collide | similarity s] = 1 − (1 − sr)b
That is an S-curve, and the two parameters do different jobs: increasing r sharpens it and pushes it right (harder to match), increasing b pushes it back left. Tune them together and you get an arbitrarily crisp step at a threshold of your choosing, at the cost of n hashes per document. Lee et al.'s setting is n = 9000, b = 20, r = 450 (71:18): the implied threshold (1/b)1/r is ≈ 0.993, and at exactly the threshold the collision probability is 1 − (1 − 1/b)b → 1 − 1/e ≈ 0.64. A nice sanity check — at the threshold the answer is a coin flip, and the curve is steep enough that a hair either side resolves to near-certainty.
Where the lecturer hedges
Two open questions close the lecture, and both are worth more than the algorithms. Asked how you dedupe paraphrases now that synthetic data is everywhere, Percy points at embedding-based similarity — LSH was invented for approximate nearest neighbours, so embed the documents and the same framework applies — then warns that embeddings cost far more than MinHash, and that fuzzy matching at a permissive threshold deletes enormous amounts of legitimate data (74:21). The regime that works is the one where near-duplicates really are near-duplicates.
Asked whether high-quality data should sometimes be duplicated deliberately, he agrees outright (75:27): mid-training routinely takes multiple epochs over good sources, and dedup is a pretraining-hygiene tool for the long tail of junk. He then speculates that a high duplicate count is itself weak evidence of importance, so the right operation might be neither "keep one" nor "keep all" but something sublinear — square root or log of the count — and says explicitly that he does not know. One of the few places in the course where a standard practice is admitted to be convention rather than result.
"Now you have all the tools. Now, in some sense, if you ask how you teach data, this is only the beginning. You really have to spend time with the data, looking at it, filtering, and training models."— Percy Liang, 78:28
What you build with this
This lecture is the direct instruction manual for Assignment 4: Data (handout, leaderboard) — the assignment opened back at lecture 11, but lecture 13 gave you the corpus history and this one gives you the algorithms. Read the adapter functions in tests/adapters.py and the mapping is one-to-one: run_identify_language is the lid.176 section at 23:31; run_classify_nsfw and run_classify_toxic_speech are Dolma's two Jigsaw-trained fastText heads at 35:00; run_classify_quality is the GPT-3 / LLaMA positives-vs-Common-Crawl recipe at 29:52, where you pick your own target set and train your own classifier; run_exact_line_deduplication is the hash-group MapReduce at 44:19; and run_minhash_deduplication, which takes num_hashes, num_bands, ngrams and jaccard_threshold, is precisely the b/r construction from 64:52 — those four arguments are the four knobs the LSH section derives. The pieces this lecture does not cover — HTML-to-text extraction, PII masking, the Gopher heuristic quality rules — came from lecture 13 and the handout. You then train a model on your own filtered data with the assignment-1 codebase and submit to the leaderboard, which is the real lesson: the filtering choices show up as a loss number.
Supporting materials, verified
- KenLM — Kenneth Heafield · The n-gram implementation the filtering world standardized on, originally built for machine translation; the lecture's perplexity demo runs against a Wikipedia-trained KenLM model.
- Kneser–Ney smoothing — Wikipedia · The back-off scheme that makes n-gram counts usable at all; Percy links this rather than deriving it.
- CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data — Wenzek et al. (2019) · The paragraph-level "sort by perplexity, keep the top third" pipeline that produced LLaMA's web data.
- Bag of Tricks for Efficient Text Classification — Joulin, Grave, Bojanowski, Mikolov (2016) · The fastText paper: factorized bag-of-embeddings, hashed n-gram features, asynchronous SGD.
- fastText language identification (lid.176) — Facebook AI · The off-the-shelf 176-language classifier Dolma and most pipelines run; the lecture demos its failure cases live.
- Data Selection for Language Models via Importance Resampling (DSIR) — Xie, Santurkar, Ma, Liang (2023) · The hashed-n-gram importance-weight method; Percy is a co-author.
- The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset — Laurençon et al. (2023) · Where BLOOM's ~30%-English composition is documented; the lecture's example of multilinguality costing you English performance under a compute budget.
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text — Paster, Dos Santos, Azerbayev, Ba (2023) · The rules + KenLM + asymmetric-threshold fastText pipeline that yields 14.7B math tokens.
- Language Models are Few-Shot Learners (GPT-3) — Brown et al. (2020), Appendix A · The original quality classifier: curated positives, Common Crawl negatives, Pareto-based stochastic keeping.
- LLaMA: Open and Efficient Foundation Language Models — Touvron et al. (2023) · Uses pages referenced by Wikipedia as positives — a sharper target set than Wikipedia itself.
- Textbooks Are All You Need (phi-1) — Gunasekar et al. (2023) · GPT-4 labels 100K Stack-Python docs for educational value, a random forest distils it; 12.19% → 17.68% on HumanEval in a third of the steps.
- Dolma: an Open Corpus of Three Trillion Tokens — Soldaini et al. (2024) · The reference open pipeline for nearly every setting quoted here: p(English) ≥ 0.5, two Jigsaw fastText heads, Bloom-filter dedup at f = 1e-15 over paragraphs.
- Jigsaw Toxic Comment Classification Challenge — Jigsaw / Conversation AI (2018) · Wikipedia talk-page comments labelled toxic / severe toxic / obscene / threat / insult / identity hate; the training data behind Dolma's toxicity filters.
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5 / C4) — Raffel et al. (2019) · Source of the three-sentence-span exact dedup, and of the paragraph that survived 61,036 times.
- Deduplicating Training Data Makes Language Models Better — Lee et al. (2021) · The evidence that dedup helps, and the source of the n = 9000, b = 20, r = 450 MinHash-LSH setting worked through at 71:18.
- Mining of Massive Datasets, ch. 3: Finding Similar Items — Leskovec, Rajaraman, Ullman · The canonical treatment of shingling, MinHash and LSH banding; Percy links it directly for the derivation.
- Bloom filter — Wikipedia · The independence assumptions behind the false-positive formula used in the lecture.
- Berkeley CS170 notes on Bloom filters — David Wagner · The compute/memory/FPR tradeoff at more length than the lecture has time for.
- A Survey on Data Selection for Language Models — Albalak et al. (2024) · field map extra — the map of everything this lecture samples from; the natural next read if the taxonomy at 18:59 is the part that stuck.
- SemDeDup: Data-efficient learning at web-scale through semantic deduplication — Abbas, Tirumala, Simig, Ganguli, Morcos (2023) · field map extra — the canonical embedding-based answer to the closing question about paraphrase-sensitive dedup; removes 50% of a LAION subset with minimal loss.
- DataComp-LM: In search of the next generation of training sets for language models — Li et al. (2024) · field map extra — a controlled benchmark that isolates the effect of filtering choices; the empirical case that model-based filtering is the single highest-leverage knob.
- GLUE benchmark — Wang et al. · field map extra — the evaluation DSIR reports its (modest) gains on, if you want to calibrate how large "better" is here.
Exercises
- Rederive the Bloom-filter dial code — (1) implement insert/query with mmh3.hash(item, seed) for seeds 0…k−1 over a bit array of m bins; (2) insert n = 100 real strings and query 10,000 strings you know are absent, measuring the empirical false-positive rate; (3) sweep k from 1 to 30 at m = 1000 and overlay the empirical curve on f = (1 − (1 − 1/m)k·n)k; (4) check the empirical minimum lands near k* = ln2 · m/n ≈ 6.93 and f ≈ 0.008; (5) now solve the inverse problem — what m/n does Dolma's f = 1e-15 require, and how many bits per item is that? A good answer notes where the independence assumption in the formula starts to visibly fail at small m.
- Design an LSH curve to spec code — (1) write P(s; b, r) = 1 − (1 − s**r)**b and threshold(b, r) = (1/b)**(1/r); (2) find every (b, r) with b·r ≤ 256 satisfying P(0.85) < 0.01 and P(0.95) > 0.99; (3) among those, pick the one minimising b·r — that is your hashing cost per document; (4) plot the winner's S-curve against the Lee et al. setting (b = 20, r = 450) on the same axes. A good answer explains in one sentence why b and r pull the curve in different directions, and confirms that P at the threshold sits near 1 − 1/e regardless of which (b, r) you chose.
- Three scorers on one corpus code — working in the assignment 4 repo: (1) pick a target set T (Wikipedia-referenced pages, or arXiv abstracts) and take one Common Crawl WET shard as R; (2) score R three ways — KenLM perplexity under a model fit to T, a fastText binary classifier trained on T-vs-R, and a DSIR hashed-unigram ratio; (3) take the top 1% under each and measure the pairwise overlap; (4) read ten documents each method keeps that the other two reject. A good answer reports the three overlap numbers and characterises what the ratio-based method keeps that the classifier does not — the diversity claim from 18:59 is either visible here or it isn't.
- Break the quality filter, then decide if you care — construct three documents that a Wikipedia-fit KenLM assigns low perplexity but that carry no information (start from the lecture's "the the the…" case at 06:44 and get more adversarial: shuffled Wikipedia sentences, template-filled boilerplate, high-frequency function-word soup). Then estimate what fraction of a real crawl shard such text plausibly represents, and argue for or against Percy's position that filtering out true nonsense is enough. A good answer separates "this filter is gameable" from "this filter is inadequate for its actual job", and says which failure mode dedup would have caught anyway.