CS336 // FIELD MAP
← field map
LECTURE 13 · FEED ITPercy Liang · 2025-05-13 · 79 min

Data 1

Stanford CS336 · Spring 2025 · lecture 13 of 17

Transcript: cleaned auto-captions with timestamps

TL;DR — Every previous lecture asked how to train a model on a fixed dataset; this one asks where the dataset comes from, and the honest answer is that nobody has a principle — only a lineage of recipes. Percy walks the whole genealogy, BooksCorpus through Nemotron-CC, and the through-line is a single ladder: a live service (Reddit, GitHub, the web) becomes a raw snapshot (a Common Crawl dump, GH Archive) becomes processed text (HTML extraction, language ID, quality filtering, dedup) becomes an aggregate (The Pile, Dolma). The field's one real trend line is that rule-based filtering — the ethically-motivated default through 2023 — lost to model-based filtering, and DCLM is where it lost. Remember the ladder: when someone says "we trained on GitHub," none of the four rungs is specified, and every rung changes the model.

This is the lecture where CS336 admits that the thing that most determines a model's quality is the thing that the course cannot teach as theory. Architecture, optimizers, tokenization, scaling laws, parallelism — all of that assumed a dataset already existed. Percy's opening claim is that the dataset is the part that actually differentiates frontier models, and his evidence is negative space: open-weight labs publish architecture and training procedure in detail, then describe their corpus in one sentence. Llama 3's paper gives you rotary details and a parallelism strategy, and tells you the data comes "from a variety of data sources containing knowledge until the end of 2023." Two reasons for the silence, both stated bluntly: competitive advantage, and litigation. So instead of a formalism, the lecture gives you a genealogy — roughly twenty datasets in chronological order — and asks you to induce the intuitions from the pattern. That is the right way to receive it: not as a list to memorize, but as a record of which ideas kept getting reinvented.

Outline, with timestamps

Why this lecture has no theory in it

05:45

Percy sets expectations early and does not walk it back: there is no good formalism for choosing a data mixture. That is a strange thing to hear thirteen lectures into a course that has otherwise handed you closed-form scaling laws and cost models, but it has a structural cause. Data work is the one part of the LLM stack that scales with headcount. An architecture is defined once, by a small team, and then it is done; a data effort can absorb three hundred people working in parallel on multilinguality, code, math, document conversion, filtering heuristics, licensing. Data is the line item with the most linear returns to spending and the least need for a unifying theory. That is exactly why it stays heuristic.

"My hot take is that data is the most important thing in getting language models right."— Percy Liang, 00:04

He marks this as contested — Tatsu would put scaling laws first — and note the epistemic position: every number here comes from open-weight or academic releases, the only ones with published corpora. This is the genealogy of the open lineage; the right prior is that closed labs are ahead on the same road, not on a different one.

The quality ramp: pre-, mid-, post-training

02:55

A modern training run is a ramp from large amounts of low-quality data to small amounts of high-quality data. Pre-training is trillions of tokens of raw web text; mid-training a much smaller curated set installing specific capabilities (math, code, long context); post-training instruction and chat data plus RL, where safety behaviour usually lands. A base model is the checkpoint after pre-training and mid-training; an instruct or chat model is after post-training.

Hold the three stages loosely: Percy says the lines are blurry and modern runs have more stages than three, and he later collapses mid- and post-training into one section for exactly that reason. AI2's OLMo 2 makes the ramp concrete because every stage is published (04:39). Pre-training is ~3.9T tokens dominated by DCLM-baseline web text, plus code, papers, math and Wikipedia. Mid-training — the Dolmino mix — reuses the same sources, filtered much harder: DCLM-baseline drops from ~3.7T tokens to ~700B, Wikipedia stays, synthetic sets appear, and the GSM8K training set gets tossed in outright, for ~10B tokens total. Post-training is the separate Tülu 3 recipe. Note what "tossed in GSM8K" implies about the state of the art: the mid-training mix is not a principled distribution, it is a set of things someone believed would help, at ratios someone tuned.

Common Crawl is not the internet

13:31

This is the segment with the most durable payoff, because "trained on the internet" is a phrase you will hear constantly and it is false in several separable ways. Common Crawl is a non-profit founded in 2007 that has run roughly a hundred monthly crawls, each cheap relative to training — rent machines, finish in under two weeks. Mechanically it is a breadth-first traversal from hundreds of millions of seed URLs, run on Apache Nutch, with a frontier queue and four standing policies: which pages to download (selection), how hard to hit a server and whether robots.txt permits you (politeness), when to revisit a changed page, and how to cope with dynamic URLs where many addresses resolve to the same content (a major source of downstream duplication). Crucially, politeness is a deliberate coverage limit: Common Crawl is not trying to mirror the web, and not even every Wikipedia article is in it (17:43). That is why the frontier labs all run their own crawlers, and why the New York Times' robots.txt disallow list reads as a roster of LLM developers (19:22). It is guidance, not enforcement; nothing makes a crawler obey it.

Each crawl ships in two formats, and choosing between them is a real modelling decision. WARC is the raw HTTP response — usually HTML. WET is Common Crawl's own lossy HTML-to-text conversion. Most serious corpora since 2021 throw away WET and re-extract from WARC with a purpose-built tool, because the extractor materially changes downstream accuracy: the DCLM ablation puts naive WET about four points below Trafilatura (17:08). Four points from a boilerplate-stripping library is a larger effect than most architecture changes in this course, for a fraction of the engineering.

The ladder to carry away: live service → raw snapshot → processed text → aggregated dataset. GitHub the website, GH Archive the hourly event dump, The Stack's licence-filtered near-deduplicated 3.1 TB, and Dolma the mixture that contains it are four different objects. Every number that matters — token count, quality, legal exposure — is set on a rung, and a claim like "we trained on GitHub" names none of them.

Two filtering religions, and how one won

21:37

The most useful axis in the lecture is defined in 2019 by two papers solving the same problem in opposite ways.

CCNet (Meta) is model-based: deduplicate paragraphs, run a fastText language-ID classifier to keep a target language, then score every document under a KenLM 5-gram model trained on Wikipedia and keep what looks Wikipedia-like. Wikipedia stands in for "quality," which works and is also the ceiling — anything Wikipedia does not cover, the filter will not find. Its low-resource-language motivation is easy to miss: the point was Urdu as much as English.

C4 (Google, in the T5 paper) is rule-based: one April 2019 snapshot — 1.4T tokens — through hand-written heuristics. Keep lines ending in punctuation with ≥5 words, drop pages with fewer than three sentences, drop anything hitting a bad-words list, drop pages containing { (removing essentially all code), drop boilerplate, keep English at p > 0.99 under langdetect. Result: 806 GB, ~156B tokens.

Percy's framing of the trade-off is the part worth stealing (25:08). A classifier is only as good as its positive set, and when the goal is a broad corpus, curating positives with the coverage you want is circular — you needed the diverse corpus to build them. Rules fail complementarily: well-formed spam sails through, and a well-formed non-Wikipedia-ish sentence survives, which is sometimes exactly what you wanted. Neither dominates on paper.

What follows is a long rules-first era whose motivation was not purely technical. Gopher's MassiveText (2021) uses manual rules — the famous one, which you implement in Assignment 4, is that 80% of words must contain at least one alphabetic character — plus Google SafeSearch for toxicity; RefinedWeb (2023) and Dolma (2024) both state outright that they avoid ML-based filtering to avoid biases. Two legs to the argument: the only classifiers cheap enough to run over trillions of tokens were weak ones that do not understand the page, and a Wikipedia-shaped filter systematically discards text from communities that do not write like Wikipedia (43:02).

That era ends with DCLM (51:13). DataComp-LM's real contribution is infrastructure — a standard 240T-token Common Crawl pool so data-processing algorithms can be compared the way architectures are — but its headline result is a very aggressive quality classifier, and the positive set is the move. 200K examples from OpenHermes-2.5 (mostly GPT-4-generated instruction data) and the ELI5 subreddit; negatives are 200K samples of RefinedWeb, not garbage, just less curated. Train fastText on that, run it over the pool, keep the top slice: 240T tokens down to 3.8T, roughly 1.5%, beating RefinedWeb by ~3 points. Read that again — they used instruction data to select pre-training data, not to train on but to define which direction "useful" points in. It is the lecture's sharpest idea and the clearest evidence that the pre-/post-training boundary is a convenience rather than a fact. AI2 conceded by building OLMo 2 on DCLM-baseline (54:52).

Nemotron-CC (NVIDIA, late 2024) fixes DCLM's cost. Filtering to 1.5% leaves 3.8T tokens, which will not sustain a very large model trained for a long time, so NVIDIA optimized for tokens retained at fixed quality. They re-ran the HTML-to-text ablation with a token-yield objective and picked jusText over Trafilatura because it keeps more. They ensembled classifiers rather than taking one top slice: prompt Nemotron-4-340B-Instruct to score documents for educational value, distil it into a cheap model, run it alongside the DCLM classifier, bucket the scores and sample from every bucket — preserving coverage instead of trusting one notion of quality (57:10). And they used an LM to rewrite: low-quality documents rephrased, high-quality ones turned into synthetic QA and extraction pairs. Result: 6.3T tokens with a 1.1T high-quality subset, above DCLM, which is above FineWeb.

C4 (2019)rules only806 GB · ~156B tok
The Pile (2021)22 curated domains825 GB · ~275B tok
MassiveText (2021)rules + SafeSearch10.5 TB (Gopher used 300B)
LLaMA mix (2023)CCNet + Wikipedia-reference classifier1.2T tok
RefinedWeb (2023)rules, no ML, MinHash dedup5T tok (600B released)
Dolma (2024)rules + Jigsaw toxicity + Bloom-filter dedup3T tok
FineWeb (2024)rules, 95 CC dumps, PII scrub15T tok
DCLM-baseline (2024)fastText quality classifier3.8T tok (from 240T pool)
Nemotron-CC (2024)classifier ensemble + LM rewriting6.3T tok (1.1T HQ)

What the source names actually contain

31:37

The middle third is Percy opening each aggregate's ingredient list, and the lesson is that familiar names hide specific, contingent objects. BooksCorpus, which trained BERT, is 7K self-published Smashwords ebooks priced at $0, scraped in 2015 and since taken down for a terms-of-service violation. Project Gutenberg is ~75K copyright-cleared books, tiny next to LibGen's ~4M — that gap is the whole economic reason shadow libraries are in this story, and why Books3 (196K books from Bibliotik, in The Pile, in LLaMA) is now a lawsuit rather than a dataset. The Enron corpus is in The Pile because it is the only large public email corpus that exists; if your model has odd priors about email, that is why.

Two sources are singled out structurally. Stack Exchange matters because its native form is Q&A with vote metadata attached — pre-training data that already looks like the chat interaction you are trying to produce, which is Percy's argument that the pre-/post-training distinction is porous from the data side too (37:35). GitHub shows the ladder at full length: 28M public repositories in the live service; GH Archive as the hourly event snapshot; The Stack as the processed artefact — 137M repositories cloned, 51B files of which 5B unique, filtered to permissive licences with go-license-detector, near-deduplicated by MinHash, yielding 3.1 TB. Code is also the one domain where licensing is machine-readable, which is why licence filtering is possible there and nowhere else.

"When someone comes to you and says, I trained on GitHub, then you'll have to ask them: what exactly does that mean?"— Percy Liang, 41:13

The other recurring trick, appearing three times in different clothes, is borrowing someone else's curation signal. GPT-2's WebText takes outbound links from Reddit posts with ≥3 karma — 8M pages, 40 GB — using strangers' upvotes as a free quality classifier. GPT-3 trains a classifier to find text resembling WebText, Wikipedia and books. LLaMA sharpens it one turn: its classifier predicts not "does this look like a Wikipedia page" but "does this look like a page Wikipedia cites," reaching good documents that look nothing like an encyclopedia article (44:12). Human link structure labels for free. Attached caution: Percy keeps clicking "random article" and "random repository" on purpose, because your mental image of Wikipedia and GitHub comes from the pages you visit, which are wildly unrepresentative of the ones you would train on (39:56).

The Wikipedia digression carries the lecture's one security result: because Wikipedia ships periodic dumps and reverts vandalism asynchronously, an attacker can time a malicious edit to land inside a dump before rollback, injecting behaviour like negative sentiment on a trigger phrase into any model trained on it (10:30). That hole is patched; the generalization is not. Your training data is written by the open internet, where people with incentives can reach it, and you have almost no oversight.

60:31

Fifteen minutes of intellectual-property law, and the most reusable fifteen minutes in the lecture. Copyright's premise is incentive: a monopoly to encourage creating intellectual goods. US law since the Copyright Act of 1976 protects "original works of authorship fixed in any tangible medium of expression." Three consequences do the work: it covers expression, not ideas (you cannot copyright quicksort, only an implementation); it requires originality, so a telephone directory is unprotected absent creativity in selection or arrangement; and registration is not required for protection, unlike patents, though it is required — a $65 filing — before you can sue.

So "how much of the web is copyrighted" answers to essentially all of it, and was never the interesting question. The interesting one is whether you may use it, and there are two routes. Licence: a promise not to sue, negotiated (Google–Reddit, OpenAI–Shutterstock, OpenAI–Stack Exchange) or granted in advance, as with Creative Commons — created by Lessig and Eldred in 2001 to let people opt into public-domain-like reuse without waiting out the term. CC works remain copyrighted; the licence makes them usable. The structural problem: you cannot licence the web, because for a random page there is no one to negotiate with.

Which leaves fair use and its four factors: purpose and character (educational and transformative favoured over commercial and reproductive); nature of the work (factual over creative); amount used (a snippet over the whole); and effect on the market. Applied to training they split awkwardly. Copying the corpus is already the violation, before you do anything with it — a real obstacle to open data releases, since hosting the data is the copying. Training is plausibly transformative, and you can argue the system is after the idea rather than the expression. But the amount factor is hostile, because you want the whole work; models demonstrably memorize; and on market effect Percy is direct — language models affect the market for writers and artists regardless of how copyright resolves. He also punctures the assumption that this is about verbatim text: plots and characters are protectable, so a Harry Potter continuation can infringe with near-zero n-gram overlap, while a parody may be fair use. Copyright is semantics and economics, not string matching.

Finally, the layer people forget: terms of service bind independently. YouTube hosts many Creative Commons videos, and scripting downloads of them still violates YouTube's terms. Licence, fair use and ToS are three separate gates and you must clear all three. Two caveats on the law as presented: Percy simplifies the term to "75 years," where the US rule is life plus 70 (95 years from publication for works made for hire), and he is describing May 2025 while the active litigation is what will settle it. Treat this as a map of the arguments, not advice.

Buying capabilities with small data

69:51

The last section is compressed — Percy is watching the clock — and treats mid- and post-training together because the boundary does not survive contact. The framing shifts from "more quality" to "install a specific capability."

Long context is a scheduling argument first and a data argument second. Attention is quadratic in sequence length, so training long from the start wastes compute on a model that is not yet good; you extend late. From the data side you need genuine long-range dependencies, and the two natural sources are books and mathematics — LongLoRA extends Llama 2 7B from 4K to 100K tokens on PG-19 and Proof-Pile.

Task collections were the 2022 answer: reformat every existing NLP dataset as an instruction. Super-NaturalInstructions assembled 1.6K+ community-contributed tasks; the Flan Collection 1.8K+, with zero-shot, few-shot and chain-of-thought variants. You get cross-task transfer and a model that handles your favourite benchmarks. The failure mode is visible in the artefact: the prompts are templated, so a model tuned on them learns the template as much as the intent.

Instruction and chat data is where "task" dissolves and the open community went almost entirely synthetic. Alpaca generated 52K examples from text-davinci-003 via Self-Instruct. Vicuna fine-tuned on 70K real ChatGPT conversations shared to ShareGPT. WizardLM's Evol-Instruct evolves seed questions to be broader and harder. MAmmoTH2 goes back to Common Crawl, trains a fastText classifier to find quiz sites, and extracts 10M QA pairs with GPT-4 and Mixtral — the borrow-the-web's-structure trick again, now for post-training. Percy sorts all of it by provenance risk rather than quality, which is the practitioner's sort order (76:23): distilling from GPT-4 is fine for research but violates OpenAI's terms if you are building a competitor; distilling from open-weight models (Llama, Mixtral, DeepSeek-R1, Qwen — what the Llama-Nemotron set does, reasoning traces included) is commercially viable; hiring annotators is the paranoid option, with the delicious failure mode that they may quietly be using GPT-4 anyway. The counter-example is Llama 2 chat: 27,540 vendor-annotated examples Meta claimed beat millions of open ones — and which Percy thinks should have been fewer still, with the savings spent on RLHF data, where lectures 15–17 pick up.

"If you think that this whole field is a mess, you're right. It's very heuristic, which means that there's many opportunities to hopefully improve that."— Percy Liang, 78:41

What you build with this

This lecture is the conceptual backing for Assignment 4: Data (handout PDF, leaderboard), which is formally released alongside lecture 11 but is mostly an implementation of the pipeline described here. Percy flags the pieces as he goes. You do HTML-to-text extraction yourself rather than taking Common Crawl's WET files, and you compare extractors — the DCLM ablation at 17:08 is the reason that step exists and the reason it is worth points. You implement language identification and Gopher-style quality rules — the "80% of words contain at least one alphabetic character" rule named at 43:02 is literally on the assignment. You implement deduplication, exact and fuzzy, which every corpus in the genealogy needs and which the lecture keeps mentioning in passing (MinHash for RefinedWeb and The Stack, Bloom filters for Dolma). And you train a quality classifier in the DCLM mould, at which point the choice of positive set from 52:24 stops being trivia and becomes your design decision. The leaderboard metric is downstream model quality at a fixed token budget, so the whole lecture's trade-off — aggressive filtering versus token yield, exactly Nemotron-CC's complaint about DCLM — is what you are actually optimizing. L14 Data 2 continues into filtering and deduplication at implementation depth.

Supporting materials, verified

Exercises

  1. Extractor bake-off code — Reproduce the ablation behind the four-point gap at 17:08. (1) Pull one Common Crawl WARC segment and its matching WET file. (2) Run Trafilatura, Resiliparse and jusText over the WARC records. (3) For each of the four texts, record tokens retained per source document and a boilerplate rate — a cheap proxy is the fraction of lines that are navigation-like, under five words with no terminal punctuation. (4) Sample 30 documents and eyeball where the extractors disagree. (5) Report the yield/quality frontier. A good answer shows WET losing on cleanliness while jusText wins on retained tokens, and names which of the two objectives you would optimize for a 400B-parameter run versus a 1B one.
  2. Rebuild the DCLM classifier, then break it code — (1) Take positives from OpenHermes-2.5 and ELI5 and negatives from FineWeb, ~50K each. (2) Train a fastText classifier. (3) Score a held-out FineWeb shard and keep the top 1.5%, matching DCLM's ratio. (4) Characterize what survived: domain distribution, mean document length, share of Q&A-shaped text. (5) Now swap the positive set for Wikipedia alone — the 2019 CCNet choice — and re-run. A good answer quantifies how much of DCLM's advantage comes from the classifier architecture versus from choosing instruction data as the positive class, and names one document type the Wikipedia-positive filter throws away that the OpenHermes-positive filter keeps.
  3. Gopher rules on real garbage code — Implement the MassiveText manual rules against Assignment 4's harness: mean word length in [3, 10], at least 80% of words containing an alphabetic character, symbol-to-word ratio, bullet-line and ellipsis-line fractions, minimum stop-word count. Run them over 10K raw Common Crawl documents and log which rule fires per rejection. A good answer includes the per-rule rejection histogram plus 5 documents each rule rejected that you believe it should have kept — that false-positive set is the entire empirical case for model-based filtering.
  4. Provenance audit — Pick one open model with a published corpus (OLMo 2, Falcon, or a Pythia checkpoint) and write a one-page memo placing every component on the ladder — live service, snapshot, processed text, aggregate — and, for each, which of the three legal gates it clears: licence, fair use, terms of service. Flag the components you cannot resolve. A good answer identifies at least one source whose legal status changed after the model shipped (Books3 is the easy one; Reddit and Stack Exchange's 2023 API changes are the interesting ones) and says what a lab should have done differently at collection time.
Next: L14 Data 2 · Back to the map.