Data 1
Transcript: cleaned auto-captions with timestamps
This is the lecture where CS336 admits that the thing that most determines a model's quality is the thing that the course cannot teach as theory. Architecture, optimizers, tokenization, scaling laws, parallelism — all of that assumed a dataset already existed. Percy's opening claim is that the dataset is the part that actually differentiates frontier models, and his evidence is negative space: open-weight labs publish architecture and training procedure in detail, then describe their corpus in one sentence. Llama 3's paper gives you rotary details and a parallelism strategy, and tells you the data comes "from a variety of data sources containing knowledge until the end of 2023." Two reasons for the silence, both stated bluntly: competitive advantage, and litigation. So instead of a formalism, the lecture gives you a genealogy — roughly twenty datasets in chronological order — and asks you to induce the intuitions from the pattern. That is the right way to receive it: not as a list to memorize, but as a record of which ideas kept getting reinvented.
Outline, with timestamps
- 00:04 — The hot take: data is what you can't read about, so it's what matters
- 02:55 — Pre-, mid-, post-training: the quality ramp, worked through OLMo 2
- 06:22 — BERT (2018): BooksCorpus, Wikipedia, and a poisoning aside
- 12:15 — GPT-2's WebText: Reddit karma as a free quality signal
- 13:31 — Common Crawl, mechanically: seeds, politeness, WARC vs. WET
- 17:43 — Q&A: does the crawler filter, can a site opt out, how much is copyrighted
- 21:37 — CCNet vs. C4: the two filtering religions, defined in the same year
- 27:28 — GPT-3 and The Pile: quality classifiers, and 22 curated domains
- 31:37 — Inside the Pile: Gutenberg, Books3, Stack Exchange, GitHub, The Stack
- 41:49 — Gopher, LLaMA, RefinedWeb, FineWeb: the rules-only era and its argument
- 48:45 — Dolma, DCLM, Nemotron-CC: classifiers win, and then get ensembled
- 59:56 — Interlude: what about non-English data?
- 60:31 — Copyright: what it covers, licences, the four fair-use factors, terms of service
- 69:51 — Mid- and post-training: long context, task collections, instruction and chat data
- 76:57 — Summary: data does not fall from the sky
Why this lecture has no theory in it
Percy sets expectations early and does not walk it back: there is no good formalism for choosing a data mixture. That is a strange thing to hear thirteen lectures into a course that has otherwise handed you closed-form scaling laws and cost models, but it has a structural cause. Data work is the one part of the LLM stack that scales with headcount. An architecture is defined once, by a small team, and then it is done; a data effort can absorb three hundred people working in parallel on multilinguality, code, math, document conversion, filtering heuristics, licensing. Data is the line item with the most linear returns to spending and the least need for a unifying theory. That is exactly why it stays heuristic.
"My hot take is that data is the most important thing in getting language models right."— Percy Liang, 00:04
He marks this as contested — Tatsu would put scaling laws first — and note the epistemic position: every number here comes from open-weight or academic releases, the only ones with published corpora. This is the genealogy of the open lineage; the right prior is that closed labs are ahead on the same road, not on a different one.
The quality ramp: pre-, mid-, post-training
A modern training run is a ramp from large amounts of low-quality data to small amounts of high-quality data. Pre-training is trillions of tokens of raw web text; mid-training a much smaller curated set installing specific capabilities (math, code, long context); post-training instruction and chat data plus RL, where safety behaviour usually lands. A base model is the checkpoint after pre-training and mid-training; an instruct or chat model is after post-training.
Hold the three stages loosely: Percy says the lines are blurry and modern runs have more stages than three, and he later collapses mid- and post-training into one section for exactly that reason. AI2's OLMo 2 makes the ramp concrete because every stage is published (04:39). Pre-training is ~3.9T tokens dominated by DCLM-baseline web text, plus code, papers, math and Wikipedia. Mid-training — the Dolmino mix — reuses the same sources, filtered much harder: DCLM-baseline drops from ~3.7T tokens to ~700B, Wikipedia stays, synthetic sets appear, and the GSM8K training set gets tossed in outright, for ~10B tokens total. Post-training is the separate Tülu 3 recipe. Note what "tossed in GSM8K" implies about the state of the art: the mid-training mix is not a principled distribution, it is a set of things someone believed would help, at ratios someone tuned.
Common Crawl is not the internet
This is the segment with the most durable payoff, because "trained on the internet" is a phrase you will hear constantly and it is false in several separable ways. Common Crawl is a non-profit founded in 2007 that has run roughly a hundred monthly crawls, each cheap relative to training — rent machines, finish in under two weeks. Mechanically it is a breadth-first traversal from hundreds of millions of seed URLs, run on Apache Nutch, with a frontier queue and four standing policies: which pages to download (selection), how hard to hit a server and whether robots.txt permits you (politeness), when to revisit a changed page, and how to cope with dynamic URLs where many addresses resolve to the same content (a major source of downstream duplication). Crucially, politeness is a deliberate coverage limit: Common Crawl is not trying to mirror the web, and not even every Wikipedia article is in it (17:43). That is why the frontier labs all run their own crawlers, and why the New York Times' robots.txt disallow list reads as a roster of LLM developers (19:22). It is guidance, not enforcement; nothing makes a crawler obey it.
Each crawl ships in two formats, and choosing between them is a real modelling decision. WARC is the raw HTTP response — usually HTML. WET is Common Crawl's own lossy HTML-to-text conversion. Most serious corpora since 2021 throw away WET and re-extract from WARC with a purpose-built tool, because the extractor materially changes downstream accuracy: the DCLM ablation puts naive WET about four points below Trafilatura (17:08). Four points from a boilerplate-stripping library is a larger effect than most architecture changes in this course, for a fraction of the engineering.
Two filtering religions, and how one won
The most useful axis in the lecture is defined in 2019 by two papers solving the same problem in opposite ways.
CCNet (Meta) is model-based: deduplicate paragraphs, run a fastText language-ID classifier to keep a target language, then score every document under a KenLM 5-gram model trained on Wikipedia and keep what looks Wikipedia-like. Wikipedia stands in for "quality," which works and is also the ceiling — anything Wikipedia does not cover, the filter will not find. Its low-resource-language motivation is easy to miss: the point was Urdu as much as English.
C4 (Google, in the T5 paper) is rule-based: one April 2019 snapshot — 1.4T tokens — through hand-written heuristics. Keep lines ending in punctuation with ≥5 words, drop pages with fewer than three sentences, drop anything hitting a bad-words list, drop pages containing { (removing essentially all code), drop boilerplate, keep English at p > 0.99 under langdetect. Result: 806 GB, ~156B tokens.
Percy's framing of the trade-off is the part worth stealing (25:08). A classifier is only as good as its positive set, and when the goal is a broad corpus, curating positives with the coverage you want is circular — you needed the diverse corpus to build them. Rules fail complementarily: well-formed spam sails through, and a well-formed non-Wikipedia-ish sentence survives, which is sometimes exactly what you wanted. Neither dominates on paper.
What follows is a long rules-first era whose motivation was not purely technical. Gopher's MassiveText (2021) uses manual rules — the famous one, which you implement in Assignment 4, is that 80% of words must contain at least one alphabetic character — plus Google SafeSearch for toxicity; RefinedWeb (2023) and Dolma (2024) both state outright that they avoid ML-based filtering to avoid biases. Two legs to the argument: the only classifiers cheap enough to run over trillions of tokens were weak ones that do not understand the page, and a Wikipedia-shaped filter systematically discards text from communities that do not write like Wikipedia (43:02).
That era ends with DCLM (51:13). DataComp-LM's real contribution is infrastructure — a standard 240T-token Common Crawl pool so data-processing algorithms can be compared the way architectures are — but its headline result is a very aggressive quality classifier, and the positive set is the move. 200K examples from OpenHermes-2.5 (mostly GPT-4-generated instruction data) and the ELI5 subreddit; negatives are 200K samples of RefinedWeb, not garbage, just less curated. Train fastText on that, run it over the pool, keep the top slice: 240T tokens down to 3.8T, roughly 1.5%, beating RefinedWeb by ~3 points. Read that again — they used instruction data to select pre-training data, not to train on but to define which direction "useful" points in. It is the lecture's sharpest idea and the clearest evidence that the pre-/post-training boundary is a convenience rather than a fact. AI2 conceded by building OLMo 2 on DCLM-baseline (54:52).
Nemotron-CC (NVIDIA, late 2024) fixes DCLM's cost. Filtering to 1.5% leaves 3.8T tokens, which will not sustain a very large model trained for a long time, so NVIDIA optimized for tokens retained at fixed quality. They re-ran the HTML-to-text ablation with a token-yield objective and picked jusText over Trafilatura because it keeps more. They ensembled classifiers rather than taking one top slice: prompt Nemotron-4-340B-Instruct to score documents for educational value, distil it into a cheap model, run it alongside the DCLM classifier, bucket the scores and sample from every bucket — preserving coverage instead of trusting one notion of quality (57:10). And they used an LM to rewrite: low-quality documents rephrased, high-quality ones turned into synthetic QA and extraction pairs. Result: 6.3T tokens with a 1.1T high-quality subset, above DCLM, which is above FineWeb.
| C4 (2019) | rules only | 806 GB · ~156B tok |
| The Pile (2021) | 22 curated domains | 825 GB · ~275B tok |
| MassiveText (2021) | rules + SafeSearch | 10.5 TB (Gopher used 300B) |
| LLaMA mix (2023) | CCNet + Wikipedia-reference classifier | 1.2T tok |
| RefinedWeb (2023) | rules, no ML, MinHash dedup | 5T tok (600B released) |
| Dolma (2024) | rules + Jigsaw toxicity + Bloom-filter dedup | 3T tok |
| FineWeb (2024) | rules, 95 CC dumps, PII scrub | 15T tok |
| DCLM-baseline (2024) | fastText quality classifier | 3.8T tok (from 240T pool) |
| Nemotron-CC (2024) | classifier ensemble + LM rewriting | 6.3T tok (1.1T HQ) |
What the source names actually contain
The middle third is Percy opening each aggregate's ingredient list, and the lesson is that familiar names hide specific, contingent objects. BooksCorpus, which trained BERT, is 7K self-published Smashwords ebooks priced at $0, scraped in 2015 and since taken down for a terms-of-service violation. Project Gutenberg is ~75K copyright-cleared books, tiny next to LibGen's ~4M — that gap is the whole economic reason shadow libraries are in this story, and why Books3 (196K books from Bibliotik, in The Pile, in LLaMA) is now a lawsuit rather than a dataset. The Enron corpus is in The Pile because it is the only large public email corpus that exists; if your model has odd priors about email, that is why.
Two sources are singled out structurally. Stack Exchange matters because its native form is Q&A with vote metadata attached — pre-training data that already looks like the chat interaction you are trying to produce, which is Percy's argument that the pre-/post-training distinction is porous from the data side too (37:35). GitHub shows the ladder at full length: 28M public repositories in the live service; GH Archive as the hourly event snapshot; The Stack as the processed artefact — 137M repositories cloned, 51B files of which 5B unique, filtered to permissive licences with go-license-detector, near-deduplicated by MinHash, yielding 3.1 TB. Code is also the one domain where licensing is machine-readable, which is why licence filtering is possible there and nowhere else.
"When someone comes to you and says, I trained on GitHub, then you'll have to ask them: what exactly does that mean?"— Percy Liang, 41:13
The other recurring trick, appearing three times in different clothes, is borrowing someone else's curation signal. GPT-2's WebText takes outbound links from Reddit posts with ≥3 karma — 8M pages, 40 GB — using strangers' upvotes as a free quality classifier. GPT-3 trains a classifier to find text resembling WebText, Wikipedia and books. LLaMA sharpens it one turn: its classifier predicts not "does this look like a Wikipedia page" but "does this look like a page Wikipedia cites," reaching good documents that look nothing like an encyclopedia article (44:12). Human link structure labels for free. Attached caution: Percy keeps clicking "random article" and "random repository" on purpose, because your mental image of Wikipedia and GitHub comes from the pages you visit, which are wildly unrepresentative of the ones you would train on (39:56).
The Wikipedia digression carries the lecture's one security result: because Wikipedia ships periodic dumps and reverts vandalism asynchronously, an attacker can time a malicious edit to land inside a dump before rollback, injecting behaviour like negative sentiment on a trigger phrase into any model trained on it (10:30). That hole is patched; the generalization is not. Your training data is written by the open internet, where people with incentives can reach it, and you have almost no oversight.
Copyright, in the shape a practitioner needs it
Fifteen minutes of intellectual-property law, and the most reusable fifteen minutes in the lecture. Copyright's premise is incentive: a monopoly to encourage creating intellectual goods. US law since the Copyright Act of 1976 protects "original works of authorship fixed in any tangible medium of expression." Three consequences do the work: it covers expression, not ideas (you cannot copyright quicksort, only an implementation); it requires originality, so a telephone directory is unprotected absent creativity in selection or arrangement; and registration is not required for protection, unlike patents, though it is required — a $65 filing — before you can sue.
So "how much of the web is copyrighted" answers to essentially all of it, and was never the interesting question. The interesting one is whether you may use it, and there are two routes. Licence: a promise not to sue, negotiated (Google–Reddit, OpenAI–Shutterstock, OpenAI–Stack Exchange) or granted in advance, as with Creative Commons — created by Lessig and Eldred in 2001 to let people opt into public-domain-like reuse without waiting out the term. CC works remain copyrighted; the licence makes them usable. The structural problem: you cannot licence the web, because for a random page there is no one to negotiate with.
Which leaves fair use and its four factors: purpose and character (educational and transformative favoured over commercial and reproductive); nature of the work (factual over creative); amount used (a snippet over the whole); and effect on the market. Applied to training they split awkwardly. Copying the corpus is already the violation, before you do anything with it — a real obstacle to open data releases, since hosting the data is the copying. Training is plausibly transformative, and you can argue the system is after the idea rather than the expression. But the amount factor is hostile, because you want the whole work; models demonstrably memorize; and on market effect Percy is direct — language models affect the market for writers and artists regardless of how copyright resolves. He also punctures the assumption that this is about verbatim text: plots and characters are protectable, so a Harry Potter continuation can infringe with near-zero n-gram overlap, while a parody may be fair use. Copyright is semantics and economics, not string matching.
Finally, the layer people forget: terms of service bind independently. YouTube hosts many Creative Commons videos, and scripting downloads of them still violates YouTube's terms. Licence, fair use and ToS are three separate gates and you must clear all three. Two caveats on the law as presented: Percy simplifies the term to "75 years," where the US rule is life plus 70 (95 years from publication for works made for hire), and he is describing May 2025 while the active litigation is what will settle it. Treat this as a map of the arguments, not advice.
Buying capabilities with small data
The last section is compressed — Percy is watching the clock — and treats mid- and post-training together because the boundary does not survive contact. The framing shifts from "more quality" to "install a specific capability."
Long context is a scheduling argument first and a data argument second. Attention is quadratic in sequence length, so training long from the start wastes compute on a model that is not yet good; you extend late. From the data side you need genuine long-range dependencies, and the two natural sources are books and mathematics — LongLoRA extends Llama 2 7B from 4K to 100K tokens on PG-19 and Proof-Pile.
Task collections were the 2022 answer: reformat every existing NLP dataset as an instruction. Super-NaturalInstructions assembled 1.6K+ community-contributed tasks; the Flan Collection 1.8K+, with zero-shot, few-shot and chain-of-thought variants. You get cross-task transfer and a model that handles your favourite benchmarks. The failure mode is visible in the artefact: the prompts are templated, so a model tuned on them learns the template as much as the intent.
Instruction and chat data is where "task" dissolves and the open community went almost entirely synthetic. Alpaca generated 52K examples from text-davinci-003 via Self-Instruct. Vicuna fine-tuned on 70K real ChatGPT conversations shared to ShareGPT. WizardLM's Evol-Instruct evolves seed questions to be broader and harder. MAmmoTH2 goes back to Common Crawl, trains a fastText classifier to find quiz sites, and extracts 10M QA pairs with GPT-4 and Mixtral — the borrow-the-web's-structure trick again, now for post-training. Percy sorts all of it by provenance risk rather than quality, which is the practitioner's sort order (76:23): distilling from GPT-4 is fine for research but violates OpenAI's terms if you are building a competitor; distilling from open-weight models (Llama, Mixtral, DeepSeek-R1, Qwen — what the Llama-Nemotron set does, reasoning traces included) is commercially viable; hiring annotators is the paranoid option, with the delicious failure mode that they may quietly be using GPT-4 anyway. The counter-example is Llama 2 chat: 27,540 vendor-annotated examples Meta claimed beat millions of open ones — and which Percy thinks should have been fewer still, with the savings spent on RLHF data, where lectures 15–17 pick up.
"If you think that this whole field is a mess, you're right. It's very heuristic, which means that there's many opportunities to hopefully improve that."— Percy Liang, 78:41
What you build with this
This lecture is the conceptual backing for Assignment 4: Data (handout PDF, leaderboard), which is formally released alongside lecture 11 but is mostly an implementation of the pipeline described here. Percy flags the pieces as he goes. You do HTML-to-text extraction yourself rather than taking Common Crawl's WET files, and you compare extractors — the DCLM ablation at 17:08 is the reason that step exists and the reason it is worth points. You implement language identification and Gopher-style quality rules — the "80% of words contain at least one alphabetic character" rule named at 43:02 is literally on the assignment. You implement deduplication, exact and fuzzy, which every corpus in the genealogy needs and which the lecture keeps mentioning in passing (MinHash for RefinedWeb and The Stack, Bloom filters for Dolma). And you train a quality classifier in the DCLM mould, at which point the choice of positive set from 52:24 stops being trivia and becomes your design decision. The leaderboard metric is downstream model quality at a fixed token budget, so the whole lecture's trade-off — aggressive filtering versus token yield, exactly Nemotron-CC's complaint about DCLM — is what you are actually optimizing. L14 Data 2 continues into filtering and deduplication at implementation depth.
Supporting materials, verified
- Executable lecture: lecture_13.py trace — Percy Liang (2025) · the lecture's own script, with every dataset link and number; the source file is the authoritative version of anything the captions garble
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Devlin et al. (2018) · the BooksCorpus + Wikipedia mix, and the shift from sentences to documents
- Aligning Books and Movies — Zhu et al. (2015) · the vision-language paper that scraped Smashwords and produced BooksCorpus, since taken down
- Poisoning Web-Scale Training Datasets is Practical — Carlini et al. (2023) · the Wikipedia dump-timing attack; pair with Concealed Data Poisoning Attacks on NLP Models (Wallace et al., 2020) for the trigger-phrase exploit
- Language Models are Unsupervised Multitask Learners — Radford et al. (2019) · GPT-2 and WebText: Reddit karma as a quality proxy · open replication at OpenWebTextCorpus
- Common Crawl — non-profit, since 2007 · the monthly crawl, WARC and WET formats, and the coverage limits the whole lecture works around
- CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data — Wenzek et al. (2019) · the model-based branch: language ID plus a KenLM Wikipedia 5-gram quality score
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer — Raffel et al. (2019) · famous for T5, but C4 is the contribution this lecture cares about: the rule-based branch
- Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus — Dodge et al. (2021) · what is actually inside C4, and what the bad-words filter removed
- Language Models are Few-Shot Learners — Brown et al. (2020) · §2.2 is the GPT-3 data section: the WebText/Wikipedia/books quality classifier and fuzzy dedup
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling — Gao et al. (2021) · the 22-domain grassroots corpus; also where WARC-plus-jusText beat WET in public
- The Stack: 3 TB of permissively licensed source code — Kocetkov et al. (2022) · the full live-service → snapshot → processed ladder, done for GitHub via GH Archive
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher — Rae et al. (2021) · MassiveText and the manual quality rules you reimplement in Assignment 4
- LLaMA: Open and Efficient Foundation Language Models — Touvron et al. (2023) · the "pages Wikipedia references" classifier · reproduced as RedPajama v1 and deduplicated to Cerebras's SlimPajama (627B tokens)
- The RefinedWeb Dataset for Falcon LLM — Penedo et al. (2023) · the strongest statement of "web data only," with an explicit refusal of ML-based filtering · successor FineWeb (15T tokens) is the best lightly-filtered base to build on
- Dolma: an Open Corpus of Three Trillion Tokens — Soldaini et al. (2024) · AI2's fully-documented mix; the rules-only position immediately before it collapsed
- DataComp-LM: In search of the next generation of training sets for language models — Li et al. (2024) · the pivot point: a fastText classifier trained on instruction data, 240T → 3.8T tokens, plus the HTML-extractor ablation
- Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset — Su et al. (2024) · classifier ensembling and LM rewriting, optimizing token yield rather than quality alone
- 2 OLMo 2 Furious — OLMo team, AI2 (2024) · the worked three-stage example at 04:39, and the model that adopted DCLM-baseline
- Tülu 3: Pushing Frontiers in Open Language Model Post-Training — Lambert et al. (2024) · the post-training half of the OLMo 2 example
- Super-NaturalInstructions — Wang et al. (2022) · 1.6K+ community tasks in one prompt format · and The Flan Collection (Longpre et al., 2023), 1.8K+ tasks with zero-shot / few-shot / CoT variants
- Alpaca — Taori et al., Stanford CRFM (2023) · 52K synthetic examples via Self-Instruct · see also WizardLM / Evol-Instruct and MAmmoTH2 (10M QA pairs mined from Common Crawl quiz sites)
- Llama 2: Open Foundation and Fine-Tuned Chat Models — Touvron et al. (2023) · 27,540 vendor-annotated examples claimed to beat millions of open ones · contrast the Llama-Nemotron post-training dataset, synthesized from open-weight models
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models — Chen et al. (2023) · 4K → 100K context on PG-19 and Proof-Pile, the concrete version of the long-context data argument
- Foundation Models and Fair Use — Henderson et al. (2023) · Percy's own further-reading pick · with Fair Learning (Lemley & Casey, 2021) and The Files are in the Computer (Lee, Cooper & Grimmelmann, 2024)
- Deduplicating Training Data Makes Language Models Better — Lee et al. (2021) · field map extra — the lecture mentions dedup at every stop without ever justifying it; this is the justification, and it is what Assignment 4's dedup section is graded against
- Scaling Data-Constrained Language Models — Muennighoff et al. (2023) · field map extra — makes Nemotron-CC's "we need more tokens" quantitative: how far repeated epochs get you before returns collapse, which is the real cost of aggressive filtering
- Extracting Training Data from Large Language Models — Carlini et al. (2020) · field map extra — the memorization result that the fair-use argument has to survive; read it directly after the copyright section
Exercises
- Extractor bake-off code — Reproduce the ablation behind the four-point gap at 17:08. (1) Pull one Common Crawl WARC segment and its matching WET file. (2) Run Trafilatura, Resiliparse and jusText over the WARC records. (3) For each of the four texts, record tokens retained per source document and a boilerplate rate — a cheap proxy is the fraction of lines that are navigation-like, under five words with no terminal punctuation. (4) Sample 30 documents and eyeball where the extractors disagree. (5) Report the yield/quality frontier. A good answer shows WET losing on cleanliness while jusText wins on retained tokens, and names which of the two objectives you would optimize for a 400B-parameter run versus a 1B one.
- Rebuild the DCLM classifier, then break it code — (1) Take positives from OpenHermes-2.5 and ELI5 and negatives from FineWeb, ~50K each. (2) Train a fastText classifier. (3) Score a held-out FineWeb shard and keep the top 1.5%, matching DCLM's ratio. (4) Characterize what survived: domain distribution, mean document length, share of Q&A-shaped text. (5) Now swap the positive set for Wikipedia alone — the 2019 CCNet choice — and re-run. A good answer quantifies how much of DCLM's advantage comes from the classifier architecture versus from choosing instruction data as the positive class, and names one document type the Wikipedia-positive filter throws away that the OpenHermes-positive filter keeps.
- Gopher rules on real garbage code — Implement the MassiveText manual rules against Assignment 4's harness: mean word length in [3, 10], at least 80% of words containing an alphabetic character, symbol-to-word ratio, bullet-line and ellipsis-line fractions, minimum stop-word count. Run them over 10K raw Common Crawl documents and log which rule fires per rejection. A good answer includes the per-rule rejection histogram plus 5 documents each rule rejected that you believe it should have kept — that false-positive set is the entire empirical case for model-based filtering.
- Provenance audit — Pick one open model with a published corpus (OLMo 2, Falcon, or a Pythia checkpoint) and write a one-page memo placing every component on the ladder — live service, snapshot, processed text, aggregate — and, for each, which of the three legal gates it clears: licence, fair use, terms of service. Flag the components you cannot resolve. A good answer identifies at least one source whose legal status changed after the model shipped (Books3 is the easy one; Reddit and Stack Exchange's 2023 API changes are the interesting ones) and says what a lab should have done differently at collection time.