CS336 // FIELD MAP
← field map
ASSIGNMENT 4 · DATA65 pts · v1.0.4 · Spring 2025

Filtering Language Modeling Data

Stanford CS336 · Spring 2025 · assignment 4 of 5 · walkthrough of the official handout, everything linked

Lectures behind it: L11 Scaling laws 2 (opens it), L12 Evaluation (why the metric is perplexity), L13 Data 1 (the corpus history), L14 Data 2 (the algorithms).

TL;DR — You build a Common Crawl processing pipeline from scratch: HTML-to-text extraction, language identification, PII masking, NSFW and toxicity classification, Gopher-style heuristic quality rules, a trained fastText quality classifier, exact line deduplication, and MinHash+LSH fuzzy document deduplication. Then you point that pipeline at 5,000 Common Crawl WET files (~375 GB compressed), tokenize whatever survives with the GPT-2 tokenizer, and train a frozen, staff-supplied GPT-2-small-shaped model on it. You hand in a write-up, a code zip, and a leaderboard PR carrying one number: validation loss on the C4-100-domains split of Paloma. Eleven adapter functions and six test files cover the primitives; the last 14 points — filter_data, inspect_filtered_data, tokenize_data, train_model — have no unit tests at all and are graded on prose and on the loss. Budget days, not hours: the staff describe the leaderboard training run alone as a multi-hour job on 2 GPUs, and it sits at the end of a pipeline you have to run over hundreds of gigabytes first.

A4 is the assignment where the course stops letting you change the model. Every prior assignment gave you an architecture or a system to make better; this one hands you a training script you are explicitly forbidden to touch and tells you that the only remaining free variable is the corpus. That constraint is the whole pedagogical point. The literature it draws on — Gopher's heuristics, WebText's Reddit-karma trick, DCLM's classifier bake-offs, RefinedWeb's insistence that filtered web data alone is enough — all argues that data curation is where modern pretraining gains actually come from, and the only honest way to demonstrate that is to pin everything else and see whether your filters move the number. Practically, this means two things people underestimate: (1) most of your time goes into engineering throughput over hundreds of gigabytes, not into clever filter design, and (2) the write-up problems, where you read your own data and say what you see, are worth 8 points and are the ones where you actually learn what your filters did.

Map of the assignment

§ProblemPtsDeliverableGraded byLecture
2.1look_at_cc4write-upwrite-up onlyL13
2.2extract_text3code + write-uptest_extract.pyL13
2.3language_identification6code + write-uptest_langid.pyL14
2.4mask_pii3code + write-uptest_pii.pyL13
2.5harmful_content6code + write-uptest_toxicity.pyL14
2.6gopher_quality_filters3code + write-uptest_quality.pyL13
2.7quality_classifier15code (trained model)test_quality.pyL14
3.1exact_deduplication3codetest_deduplication.pyL14
3.2minhash_deduplication8codetest_deduplication.pyL14
4filter_data6script + write-upwrite-up onlyL13 · L14
4inspect_filtered_data4write-upwrite-up onlyL13
4tokenize_data2script + numberwrite-up onlyL01
4train_model2write-up + leaderboard PRleaderboardL12
Total65

1 · Assignment overview: what the repo gives you

The repo is unusual for this course in that the module you are supposed to write is empty. cs336_data/__init__.py is a blank file and the README's directory sketch is mostly TODO(you). There is no scaffolding to fill in, no half-written class hierarchy — you get a set of adapter stubs that raise NotImplementedError and total freedom about what sits behind them. That freedom is real: nothing in the tests constrains your module layout, your CLI, or how you parallelize.

PathWhat it isDo you edit it?
cs336_data/Empty module — your filtering and dedup codeYes, all of it
tests/adapters.py11 stub functions the graders callYes — wire them to your code
tests/test_*.py6 test files, the provided suiteNo (your code must pass them as given)
cs336-basics/Staff A1 model, optimized, plus a DDP training scriptNo — leaderboard runs must use it as-is
configs/experiment/your_data.yamlHydra experiment config for your runOnly paths.train_bin and the two wandb keys
get_assets.shSymlinks or downloads the two Dolma classifiers into cs336_data/assetsNo, just run it
test_and_make_submission.shRuns pytest, writes test_results.xml, zips the submissionNo, just run it

Note what test_and_make_submission.sh excludes from the zip: *.bin, *.pt, *.txt, *.json, *.pkl. Your trained quality classifier is a .bin and will not be in your submission. That is intentional — the graders re-run the suite from your zip, so anything binary has to be reproducible or fetchable, which is also why the leaderboard README asks you to snapshot your pipeline.

Setup: environment, data, tests

Dependencies are managed by uv as in every CS336 assignment; uv run pytest resolves the environment on first call. The pyproject.toml already pins the libraries the handout expects you to use, so you should not need to add much: resiliparse and fastwarc for HTML and WARC handling, fasttext for the three classifiers, mmh3 for MinHash, nltk for word tokenization, tldextract for domain extraction, and xopen for transparent gzip I/O (the dedup tests open your outputs with xopen, so writing plain text or gzip both work).

The test harness is the same indirection you know from A1 and A2. Each tests/test_*.py imports a run_* function from tests/adapters.py; those stubs raise NotImplementedError until you make them call into cs336_data. Fixture files live under tests/fixtures/ and are resolved through tests/common.py. Run a single problem's tests with the -k selector the handout names for it, e.g. uv run pytest -k test_mask_ips. "Passes" here is weaker than it looks: several tests are explicitly sanity checks with a TODO comment telling you to adjust the expected label to whatever your system returns. Passing the suite does not mean your filter is good — it means your adapter has the right shape.

Everything else is data you have to fetch or find on the cluster. The table below is the full asset list the handout names, with both the Together cluster path and a public download that resolves today.

AssetCluster pathPublic source
Sample WARC (Apr 2025 crawl)/data/CC/example.warc.gzdata.commoncrawl.org · …00065.warc.gz
Matching WET/data/CC/example.warc.wet.gzdata.commoncrawl.org · …00065.warc.wet.gz
Leaderboard corpus — 5,000 WETs, ~375 GB compressed/data/CC/CC*.warc.wet.gzCommon Crawl · get started
fastText language ID, 176 languages/data/classifiers/lid.176.binfasttext.cc · language identification · lid.176.bin
Dolma NSFW classifier (Jigsaw bigrams)/data/classifiers/dolma_fasttext_nsfw_jigsaw_model.bindolma-artifacts.org · HF mirror
Dolma hate-speech classifier (Jigsaw bigrams)/data/classifiers/dolma_fasttext_hatespeech_jigsaw_model.bindolma-artifacts.org · HF mirror
Wikipedia outbound URLs — 43.5 M links, enwiki 2024-04-20/data/wiki/enwiki-20240420-extracted_urls.txt.gznlp.stanford.edu · enwiki-20240420-extracted_urls.txt.gz
Paloma C4-100-domains validation, GPT-2 tokenized/data/paloma/tokenized_paloma_c4_100_domains_validation.binallenai/paloma

The validation set is a raw uint16 dump, not a HuggingFace dataset: read it with np.fromfile(path, dtype=np.uint16) and decode with the GPT-2 tokenizer. Do read some of it. Knowing what C4's hundred most common domains look like is the single most useful piece of information for designing filters, and the handout permits you to use it — see the rule quoted in filter_data.

2.1 · Looking at the data

The section that gets skipped and shouldn't. Common Crawl ships three parallel views of every crawl: WARC (raw response bytes plus HTTP metadata), WAT (extracted metadata as JSON — links, titles), and WET (Common Crawl's own plain-text extraction). Most of the assignment lives in the gap between WARC and WET: WET is what you get if someone else makes the extraction decisions for you, and §2.2 exists to show you those decisions are not good enough. The mental model to carry forward is that a crawl is not a corpus. It is a pile of HTTP responses, unsorted, in every language, most of them navigation chrome, parked domains, and boilerplate.

Problem (look_at_cc): 4 points

Deliverable: four short written answers — 2–3 sentences on the first WARC record, 3–4 sentences comparing it to the WET extraction, 1–2 sentences on domain-dependence of "good", and annotations of 25 WET records plus a count of how many you read before hitting something you'd call high quality.

There is no code and no test here; all four points are write-up. Part (d) is where the value is and where people cut corners — annotating 25 documents by hand is tedious, and the temptation is to write three and generalize. Don't: the count you produce ("it took me N records to find one good page") is the number that calibrates every threshold you pick later. If your English fraction and your quality-classifier precision estimates disagree with what you saw by hand in these 25, one of them is wrong. Practically, zcat file.warc.gz | less works but is painful for 25 records; it is faster to iterate with FastWARC's ArchiveIterator and WarcRecordType and print record URLs plus the first few hundred characters. Note the handout's warning: this is unfiltered web, and you will hit content you'd rather not read.

2.2 · HTML to text conversion

Text extraction is the first place where an ostensibly boring engineering choice moves downstream loss. Any extractor has to decide what counts as "the content" of a page, and visible text is a bad proxy: menus, cookie banners, related-article rails, and footers are all visible. The course's reason for making you do this yourself rather than take WET is the DCLM ablation covered in L13, where swapping extractors measurably changed downstream benchmark scores. Two sub-problems hide here: pulling main content out of a DOM, and figuring out what bytes you even have — Resiliparse's encoding detection exists because a non-trivial minority of the web is not UTF-8 and a naive bytes.decode() will throw or, worse, silently produce mojibake that then poisons your language classifier.

Problem (extract_text): 3 points

Deliverable: (a) a function from HTML bytes to extracted text, wired to the adapter and passing its test; (b) a 2–3 sentence comparison of your extraction against the WET text for the same WARC file.

Use resiliparse.extract.html2text.extract_plain_text, which takes a str, so the real work is the decode step: try UTF-8, and on failure fall back to resiliparse.parse.encoding.detect_encoding() and decode with that. The single test is unusually strict — it asserts exact string equality against a checked-in fixture, so if you post-process the extractor's output at all (stripping blank lines, normalizing whitespace, collapsing newlines) the test fails. Do the normalization later in your pipeline, not inside the adapter. The adapter's return type is str | None, which is the handout's hint that returning None for undecodable input is acceptable. For part (b), the honest comparison is not "mine is better" but a specific observation: note whether your extractor keeps or drops navigation blocks, list markup, and repeated headers relative to WET, because those are exactly what §3.1's line dedup will later remove for free.

2.3 · Language identification

Almost every public pretraining corpus is language-filtered, not because multilingual data is bad but because at a fixed token budget, spreading capacity across a thousand languages buys you a worse model in each of them. The standard tool is fastText's lid.176 model — a linear classifier over character n-grams, small and fast enough to run over a whole crawl, which is the only property that matters at this scale. It returns a label like __label__en and a probability; filters keep documents above a confidence threshold, and choosing that threshold is a real decision, not a detail.

Problem (language_identification): 6 points

Deliverable: (a) a function returning a (language identifier, confidence in [0,1]) pair, wired to the adapter and passing both tests; (b) a 2–5 sentence discussion of what goes wrong downstream when language ID misfires and how you'd mitigate it in a deployed product; (c) a 2–5 sentence report on hand-labelling 20 extracted documents, the English fraction you observed, and the threshold you'd pick.

The mechanical trap is label formatting. fastText returns __label__en; the tests assert the string "en" and "zh" exactly, and the handout tells you to do any remapping inside the adapter. It also asserts isinstance(score, float) — fastText hands back a NumPy array, so cast it, or you will fail a test that has nothing to do with your classifier. The conceptual half is more interesting: lid.176 is trained largely on Wikipedia and news, so it degrades exactly where the web is weirdest — very short documents, code, heavily templated pages, and code-switched text. Part (b) wants you to reason about that asymmetrically: a low threshold admits garbage and non-target languages into training; a high threshold silently deletes dialects, minority varieties, and non-Latin scripts, which is a fairness problem you cannot see in aggregate loss. The usual mitigations are per-language thresholds, a human-audited sample, and reporting language distribution as a monitored pipeline metric rather than a one-off check.

2.4 · Personal identifiable information

Models memorize, and anything a model memorizes it can emit. Masking contact details in the training set is the cheapest available mitigation — cheap enough that it is standard, crude enough that nobody claims it is sufficient. You implement three regex-based maskers for emails, US phone numbers, and IPv4 addresses, each returning the rewritten text plus a count of substitutions made.

Problem (mask_pii): 3 points

Deliverable: three masking functions (email, phone, IPv4) wired to three adapters and passing their tests, plus a 2–5 sentence discussion of downstream problems from naive masking, plus a 2–5 sentence report of false positives and false negatives found in 20 real masked examples.

The replacement tokens are fixed strings and must match exactly: |||EMAIL_ADDRESS|||, |||PHONE_NUMBER|||, |||IP_ADDRESS|||. Two details in the tests catch people. First, test_mask_emails_existing_string feeds text that already contains the sentinel |||EMAIL_ADDRESS||| and expects a count of 2, not 3 — your regex must not match its own output, which it won't if you anchor on a real address pattern but will if you get creative. Second, the phone test iterates four formats in one loop — 2831823829, (283)-182-3829, (283) 182 3829, 283-182-3829 — and each must produce exactly one replacement, so a pattern with optional separators that can also match a bare 10-digit run is what you want; over-greedy patterns that swallow the trailing period in "…3829." fail on string equality. The write-up half is the part with actual content: naive PII masking is why models emit the literal string |||EMAIL_ADDRESS||| at inference, why IPv4 patterns eat version numbers and dates, and why documents about networking or contact directories can be degraded into near-noise. Reasonable mitigations are context-aware masking, replacing with plausible surrogates rather than sentinels, and dropping documents whose masked fraction crosses a threshold instead of keeping the wreckage.

2.5 · Harmful content

Two more fastText classifiers, both from the Dolma project and both trained on the Jigsaw toxic comment corpus of labelled Wikipedia comments: one for NSFW content, one for hate/toxic speech. You do not train these — you load the provided binaries and wrap them. The interesting content is entirely in what the classifiers are and are not: they were fit on short user comments from one site, and you are about to run them over arbitrary long web documents, which is a domain shift large enough that the score distribution you get is not the one Jigsaw's validation set implies.

Problem (harmful_content): 6 points

Deliverable: an NSFW classifier function and a toxic-speech classifier function, each returning a (label, confidence) pair and passing its sanity test; a 2–5 sentence discussion of downstream effects; and a 2–5 sentence report comparing classifier verdicts to your own on 20 real documents, with the harmful fraction and the thresholds you'd use.

Labels are asserted as exact strings: "nsfw" / "non-nsfw" and "toxic" / "non-toxic", with a positive float score. The models' native labels differ, so remap in the adapter, same as language ID. Both tests are two examples each, drawn from Jigsaw's own training set — the test comments say so — meaning they are essentially guaranteed to pass with the right model loaded and tell you nothing about accuracy. The handout is explicit that validating accuracy is your job, and part 4 is where you do it. What you will find in practice is that the classifiers fire on documents merely discussing sexuality, medicine, or violence, and on regional dialects and reclaimed slurs — the well-documented sociolect bias of toxicity classifiers. That is the substance of part 3: aggressive harmful-content filtering systematically deletes health information, LGBTQ content, and African-American English from the corpus, which produces a model that is worse at those topics and dialects rather than safer. Sensible answers involve high thresholds for deletion, a middle band that gets downweighted rather than dropped, and reporting what was removed by domain.

2.6 · Quality rules

Before any learned quality signal, the cheap win: syntactic rules that catch pages which are obviously not prose. Gopher's Appendix A is the canonical list, and the assignment asks for a named subset of it. These rules are worth understanding as a class rather than as four thresholds: each one targets a specific web pathology. Word-count bounds kill stubs and scraped databases. Mean word length outside 3–10 characters catches both tokenizer wreckage and concatenated-URL soup. Lines ending in ellipsis are truncated article teasers on index pages. And the alphabetic-character rule catches pages that are mostly numbers, prices, or symbols — a product listing that survived extraction as a wall of SKUs.

RuleKeep the document whenWhat it removes
Word countat least 50 and at most 100,000 wordsstubs, error pages; scraped dumps
Mean word lengthbetween 3 and 10 characters inclusivesymbol soup; concatenated tokens
Ellipsis linesat most 30% of lines end with "..."index and teaser pages
Alphabetic wordsat least 80% of words contain ≥1 alphabetic characterprice tables, numeric listings

Problem (gopher_quality_filters): 3 points

Deliverable: (a) a boolean filter implementing at least the four rules above, wired to the adapter and passing the test_gopher tests; (b) a 2–5 sentence comparison of the filter's verdicts to your own on 20 real documents.

The adapter returns a plain bool — True means the document passes and is kept. Read the seven test_gopher_* cases before writing the rules, because they pin the boundary semantics your prose reading might get backwards: each test supplies one string that must fail and, usually, a near-neighbour that must pass. "the be " * 100 must fail on mean word length while "the with " * 100 must pass, which tells you the mean is over whitespace-ish word tokens including short stopwords and that 3.0 is on the passing side of the boundary. The ellipsis test uses 70/100 failing and 30/260 passing lines. The word-count test's 100,000-word case is built from 50,000 repetitions of a nine-word sentence, so a slow tokenizer here costs you real seconds every test run — NLTK's word_tokenize is suggested but not required, and a plain str.split() is both faster and, for these tests, sufficient. Where you should expect disagreement with your own judgment in part (b) is on tables, poetry, code, and transcripts: all are legitimate text that these rules were never designed for.

2.7 · Quality classifier

The heaviest problem in the assignment at 15 points, and the one with the most design freedom. The idea, traced through the handout's citations, is a chain: PageRank observed that good pages link to good pages; GPT-2's WebText operationalized that by scraping outbound links from Reddit posts above a karma threshold; LLaMA used Wikipedia references as the trusted-link source instead. All of these produce a corpus that is high quality but far too small, so the modern move is to use it as supervision rather than as data: treat trusted pages as positives, random Common Crawl as negatives, train a cheap classifier, and use its score to rank the entire crawl. The staff hand you the trusted-link list — 43.5 M external URLs extracted from English Wikipedia — and the rest is yours.

The pipeline you have to build for this problem is longer than it looks: subsample the URL list, fetch those pages with wget --warc-file into your own WARC, run them through your own §2.2 extractor, apply your own language and Gopher filters to the positives (the handout points out the positives are not automatically clean), sample negatives from CC, format both as fastText training lines, train, and pick a threshold. Every earlier primitive gets reused here, which is the point.

Problem (quality_classifier): 15 points

Deliverable: (a) a trained quality classifier that maps text to a numeric quality score; (b) a function returning a (label, confidence) pair, wired to the adapter and correctly classifying the two provided sanity examples.

The test asserts the labels "wiki" and "cc" exactly, on two fixture files — high_quality_wiki_reference.txt and low_quality_cc.txt — so name your fastText labels to match rather than inventing "high"/"low" and remapping badly. Where people get stuck is the scraping: fetching even a modest subsample of 43.5 M URLs is slow, many are dead, and a meaningful fraction of what returns is not the page the Wikipedia editor cited. Budget for a few tens of thousands of positives, parallelize the fetch, and accept the loss rate. The deeper trap is negative-set construction. If you sample negatives from raw CC while your positives have been through your language and quality filters, your classifier learns to detect "went through the pipeline", not "is good writing" — it will score fluent English near 1.0 regardless of substance and give you nothing on top of §2.3 and §2.6. Apply the same preprocessing to both sides. Finally, note that a fastText classifier gives you a continuous score, and the useful output is not the label but the score's distribution over the crawl: the threshold you pick is a token-yield knob, and §4 is where that trade-off gets scored.

3.1 · Exact line deduplication

Deduplication is the step with the best measured return per line of code in the whole pipeline, and line-level exact dedup is its simplest form. The observation is that page chrome — nav bars, footers, cookie notices, "Sign in", license boilerplate — appears verbatim across thousands of pages, so a line that occurs more than once in the corpus is almost certainly not content. Two passes: count line occurrences across all input files, then rewrite each file keeping only its unique lines. Hashing the line instead of storing it keeps the counter table at fixed key size, which is the difference between fitting in memory and not.

Problem (exact_deduplication): 3 points

Deliverable: a function taking a list of input file paths and an output directory, which rewrites each input to the output directory under the same basename with every non-unique line removed.

"if the input paths are a/1.txt and a/2.txt, and the output directory is b/, your function should write the files b/1.txt and b/2.txt."Handout §3.1 — cs336_spring2025_assignment4_data.pdf

Read the semantics carefully, because they are stricter than "deduplicate": a line appearing twice is removed from both places, not kept once. The test loads five fixture documents from documents_with_line_duplicates/, runs your function, and asserts the five outputs match the five files in documents_line_deduplicated/ as a multiset of exact strings — so trailing-newline handling matters, and you must write all five files even if one ends up empty. It opens outputs with xopen, so gzip output is fine. The two-pass structure is not optional advice: you cannot decide whether to keep line 1 of file 1 until you have seen every file. In production this is a MapReduce, which is exactly how L14 presents it.

3.2 · MinHash + LSH document deduplication

Exact matching cannot see that two MIT licenses differing only in a name and a year are the same document. Fuzzy dedup fixes that, and the standard construction has three layers you should keep separate in your head. Jaccard similarity over word n-gram sets is the notion of "same" — |S ∩ T| / |S ∪ T| — but computing it for all pairs is quadratic and storing n-gram sets is expensive. MinHash replaces each document with a length-k signature: for each of k hash functions, the minimum hash value over the document's n-grams. The useful fact is that the probability two documents share a given minhash equals their Jaccard similarity, so the fraction of matching signature positions estimates it. LSH then avoids the all-pairs comparison: split the k-element signature into b bands of r values each (k = br), hash each band, and call two documents candidates if any band collides. Increasing b at fixed k raises recall and lowers precision — that is the only knob-direction fact you need, and it falls out of the fact that a pair must match all r values in at least one of b bands.

Problem (minhash_deduplication): 8 points

Deliverable: a function taking input file paths, a number of hashes, a number of bands, an n-gram length in words, a Jaccard threshold, and an output directory; it computes signatures, finds LSH candidates, verifies true n-gram Jaccard against the threshold, clusters the confirmed duplicates transitively, and writes out each input file keeping one randomly chosen survivor per cluster.

"normalize the text before computing minhash signatures and/or comparing Jaccard similarity by lowercasing, removing punctuation, normalizing whitespaces, and removing accents, and applying NFD unicode normalization."Handout §3.2, following Penedo et al., 2023 — handout PDF

Four things reliably go wrong. (1) Normalization is graded implicitly: the fuzzy test's two MIT licenses differ in whitespace and attribution, and without the normalization above their n-gram Jaccard sits below 0.8 and nothing gets removed. (2) Clustering must be transitive — if A~B in one band and B~C in another, all three are one cluster; a pairwise removal loop will over-delete. Union-find is the natural structure. (3) The candidate step is a filter, not a decision: you still compute exact n-gram Jaccard on candidate pairs and compare against the threshold, otherwise your b/r choice silently becomes your similarity threshold. (4) k must be divisible by b — the handout says you may assume it, so don't over-engineer, but do assert it. Use mmh3 with distinct seeds for the k hash functions; that is what "same family, different seeds" in the handout footnote means.

Testnum_hashesnum_bandsngramsjaccard_thresholdFiles in → out
exact duplicates1001050.85 → 4 (doc1 ≡ doc2)
fuzzy duplicates5005050.83 → 2 (rails ≈ react MIT license)

4 · Leaderboard: filter data for language modeling

Everything above was a primitive; this section is the assignment. You point your pipeline at 5,000 WET files and produce a training corpus, tokenize it, and train the staff's model on it. The scored quantity is validation loss on the C4-100-domains split of Paloma — the hundred most common domains in C4. Read that target carefully: you are not optimizing "good data" in the abstract, you are optimizing a fit to a specific, known, English web-text distribution, and the filters that win are the ones that shape your corpus toward it. L12 explains why the course scores it this way — perplexity on a genuinely disjoint split is smooth enough to rank near-identical submissions and hard to game.

The permitted-use rule is the sharpest line in the handout and worth quoting exactly, because it is easy to cross accidentally by, say, seeding your quality classifier's positives from the validation text.

"you are allowed to make use of the Paloma validation data in constructing filters or classifiers to process the CC WET files, but are not allowed to literally copy any of the validation data into your training data. The language model should never see any data from the validation set."Handout §4 — cs336_spring2025_assignment4_data.pdf

Problem (filter_data): 6 points

Deliverable: (a) a script or scripts that filter the 5,000 WET files in parallel into language modeling data, reporting how many examples each filter step kept, plus a written breakdown of what proportion of discards each step is responsible for; (b) the measured runtime over 5,000 files and an extrapolation to the full 100,000-WET crawl.

No test grades this. What grades it is whether your script instruments itself — the per-filter keep/discard counts are an explicit deliverable, and they are also the only way you will debug an over-aggressive pipeline. Order your filters by cost: language ID and Gopher rules are cheap and cut hard, so run them before the fastText quality classifier and long before any dedup. The handout supplies working concurrent.futures and submitit skeletons; the shape is one process per WET file, which parallelizes trivially. The genuine engineering problems appear at the seams: deduplication is global, so it cannot live inside the per-file map — you need a shuffle or a second pass. Note also that this stage uses WET files, i.e. Common Crawl's own extraction, not your §2.2 extractor (the handout switched from WARC to WET here in v1.0.1); your extractor's value in this section is the comparison you did in §2.2, not a re-extraction of 375 GB. tldextract is suggested because domain-level filtering — blocklists, per-domain caps to stop one forum dominating your corpus — is one of the highest-leverage filters available and is not covered by any earlier problem.

Problem (inspect_filtered_data): 4 points

Deliverable: (a) five random examples from your final filtered data with a 1–2 sentence quality judgment on each; (b) five documents your pipeline removed or modified, with which step did it and whether that was justified; (c) a description of any pipeline changes these observations motivated.

This is look_at_cc again, on the other end of the pipe, and it is the problem most likely to change your leaderboard number. The failure mode it is designed to catch is a pipeline that is quietly deleting most of what you want: an over-tight quality threshold that keeps only Wikipedia-like prose, a PII masker that shredded a whole class of documents, a Gopher word-count bound removing every short but good page. Part (c) explicitly invites you to iterate before training — take it, because a training run is expensive enough that you want to spend your reads before it, not after.

Problem (tokenize_data): 2 points

Deliverable: a script that tokenizes your filtered data with the GPT-2 tokenizer and serializes it, plus the token count of your final dataset.

"Make sure to serialize following the example code above, with ids_array.tofile(output_path), where ids_array is a np.uint16 numpy array of integer IDs. This ensures compatibility with the provided training script."Handout §4, problem tokenize_data — handout PDF

Two points, and mostly a compatibility contract: the training script np.memmaps your file as uint16, so any other dtype produces a corpus that loads without error and trains to nonsense. Append the GPT-2 end-of-sequence token <|endoftext|> after every document — the handout's starter code does it per line — or documents bleed into each other across the context window. The token count you report is a real signal about your pipeline: it is your token yield, the other half of the precision/recall trade-off you set with your quality threshold, and it is the number to quote when explaining a leaderboard result. This is a genuinely slow step over hundreds of gigabytes, hence the multiprocessing.Pool.imap pattern in the handout's starter code.

Problem (train_model): 2 points

Deliverable: the best validation loss recorded on C4-100-domains, the learning curve, and a description of what you did — submitted both in the write-up and as a PR to the leaderboard.

Set paths.train_bin, training.wandb_entity and training.wandb_project in your_data.yaml and launch with uv run torchrun --standalone --nproc_per_node=2 scripts/train.py --config-name=experiment/your_data. Leave paths.valid_bin alone.

"do not modify the training config (other than the path and wandb attributes mentioned above) or the training script."Handout §4 — handout PDF

Note one discrepancy between the handout prose and the code at this ref, because it changes your time budget by a factor of two: the PDF says you train for 200K iterations and reports a ~7-hour staff run, while train_config.py at spring2025 sets train_steps = 100_000 — CHANGELOG 1.0.4 (2025-05-19) is "Halve training tokens for the leaderboard run". Trust the config, not the prose; the leaderboard is scored on runs made with it.

Config keyValue at spring2025Note
vocab_size / context_length50257 / 512GPT-2 vocabulary and a 512-token window
d_model / d_ff / num_layers / num_heads768 / 2048 / 12 / 12GPT-2 small shape; d_ff = floor(d_model·8/3/64)·64
rope_theta10000.0RoPE, not GPT-2's learned positions
train_steps100,000handout prose says 200K; halved in v1.0.4
train_batch_size128 per device2 GPUs, DDP → 256 sequences per step
tokens per step (derived)131,072128 × 2 × 512
lr / warmup_ratio / schedule1e-3 / 0.01 / cosinecosine cycle spans train_steps
weight_decay / betas / eps0.1 / (0.9, 0.98) / 1e-9AdamW
max_grad_norm / dtype / compile1.0 / bfloat16 / true
eval_interval / eval_iterations2,000 / 1,000validation loss logged to wandb every 2k steps

For reference on what a number means: the leaderboard's own naive baseline row sits at 4.00 validation loss, the Spring 2025 top three landed between 3.19 and 3.23, and roughly half the class finished under 3.5. Set +training.save_checkpoints=True on the command line if you want to sample from the model with generate_with_gpt2_tok.py — reading a few samples is a faster sanity check on your corpus than staring at the loss curve.

What you hand in

Two artifacts to Gradescope and one pull request. writeup.pdf — typeset answers to every written question: look_at_cc (a)–(d) including the 25 annotations, the extraction comparison, the three filter-audit reports (language, PII, harmful, quality rules — each "look at 20 examples and report"), the filter-step discard breakdown and runtime extrapolation, the five kept plus five discarded examples, your token count, and your best validation loss with its learning curve. code.zip — produced by running ./test_and_make_submission.sh, which runs uv run pytest -v ./tests --junitxml=test_results.xml and then zips the tree; the graders unzip it and re-run the same script, so it must work from a clean checkout of your code.

The leaderboard PR goes to assignment4-data-leaderboard, adding a row to the Markdown table sorted by increasing loss, with your final validation loss, a link to a wallclock-time learning curve (a public wandb report or an image committed to ./images), and a description of what you did. One rule from the leaderboard README is worth knowing before you start rather than after:

"we will reorder the top 5 scoring students based on our reproduced training runs. Make sure you save a snapshot of your best data pipeline so it can be reproduced by us!"Leaderboard README — assignment4-data-leaderboard

Since the submission zip excludes *.bin, that snapshot has to be your code plus whatever makes your classifier and your filtered corpus reconstructible — seeds, thresholds, the URL subsample, the exact filter order. Write that down while you are running, not afterwards.

Materials, verified

Next: A5 Alignment and reasoning RL · Back to the assignments.