Feeding AI ‘good’ data
Unit 2 walks the training-time levers roughly in the order they apply, and this is the earliest: before any objective is chosen or any human rates any output, someone decides which documents the model will ever see. The chapter tests the field's most intuitive safety proposal — "just don't train it on the bad stuff" — against what the 2025 experimental literature found. The answer is more interesting than either camp predicted: real, measurable, unusually durable, and nowhere near sufficient.
Three bets wearing one name
"Input data filtration" is used as though it were a single technique, but the chapter is really describing three separate engineering problems that happen to share a mechanism.
- Capability denial. Strip documents teaching a specific hazardous skill — pathogen synthesis, exploit development — on the theory that a model cannot recite what it never read. This is the bet with the strongest experimental support.
- Behaviour shaping. Strip or rebalance text modelling the disposition you don't want: abuse, manipulation, deceptive reasoning traces. Closer to curation than removal, and it shades into synthetic data — writing the corpus you wish existed rather than subtracting from the real one.
- Poisoning defence. Detect and drop documents an adversary planted to install a backdoor. This is not really filtering at all; it is adversarial detection, and it plays by adversarial rules.
They fail differently — knowledge structure bounds the first, base rates bound the third — and collapsing them into one technique is how you end up with the startup pitch in exercise 2.
The tamper-resistance argument
The strongest case for filtering is not that it removes more capability than other methods. It is that what it removes stays removed.
Contrast the default. Refusal training — RLHF, safety fine-tuning, whatever the lab calls it — leaves hazardous knowledge intact in the weights and adds a learned policy of declining to emit it. That policy is a thin, late-stage layer over a large pretrained substrate, and about as robust as you'd expect: anyone with the weights and a modest GPU budget fine-tunes it off in a few hundred steps on innocuous examples. For open-weight releases, the safety property you shipped is not the one the downstream user runs.
Deep Ignorance (UK AISI and EleutherAI) tested the alternative by pretraining 6.9B-parameter models from scratch on corpora with biorisk content removed — a blocklist pass dropping about 8.4% of the data, then a fine-tuned classifier pass. The filtered models scored near chance on biorisk evaluations while holding general benchmarks, at under 1% additional training FLOPs. The headline is adversarial: the researchers then attacked their own models with up to 10,000 fine-tuning steps on 300M tokens of exactly the data they had filtered out, and the models stayed well below the unfiltered baseline. You can teach a filtered model some of what it missed, but you pay full price and never get back to the unfiltered starting point.
That asymmetry is the real product. Filtering converts a safety property from "the model has been persuaded not to" into "the model has less to work with", and only the second kind survives contact with someone who controls the weights.
The paper is honest about the ceiling: staged attacks combining fine-tuning with in-context retrieval — hand the hazardous material to the model at inference time and let it reason — defeated every defence tested, filtering included. A model that lacks knowledge can still be given it.
What "33%" actually buys you
Anthropic's pretraining data filtering work is the industrial-scale version of the same experiment, and it rewards reading the numbers rather than the summary. They built classifiers to flag CBRN content, trained a "constitutional" variant to F1 ≈ 0.94–0.96, filtered a pretraining corpus, and measured the resulting model on hazardous-knowledge and ordinary benchmarks.
The reported effect is a 33% relative reduction in harmful capability measured against the random-chance baseline: accuracy on the hazardous evaluation moved from 33.7% to 30.8%. On a multiple-choice benchmark where guessing scores around 25%, that is roughly three absolute points — about a third of the model's above-chance margin. MMLU natural science, prose and code were unaffected.
Read carelessly, "33% reduction in dangerous capabilities" sounds like a third of the risk is gone. Read carefully, it says something more modest and more useful: the filter is cheap, measurable, and costs nothing in general capability — an excellent property for a defence-in-depth layer and a terrible one for a load-bearing layer. Note also how much work the evaluation does: a benchmark where the unfiltered model barely clears chance cannot show a large effect no matter how good the filter is. That is the general problem with measuring dangerous capabilities before they exist at scale.
Dual use is a boundary problem, not a labelling problem
The obvious objection to capability denial is that hazardous knowledge does not come in a separate box.
"The same facts about virology could help someone design a vaccine or a novel pathogen."— What is input data filtration in AI safety?, Sarah Hastings-Woodhouse (2025)
This is usually framed as a labelling difficulty — the classifier can't tell good virology from bad. That gives the classifier too much credit for the problem: the distinction being asked for is not in the text at all. A protocol for enhancing transmissibility reads identically to a gain-of-function reviewer and to an attacker. The harm is a property of the downstream act, and no classifier can extract a signal that isn't in its input.
The harder version: even having removed every document stating a hazardous fact, a capable model can reconstruct hazardous conclusions by composing benign ones — biochemistry plus lab technique plus literature-search skill is, in the limit, most of what the filtered document contained. Filtering suppresses recall of memorised material; it does much less to suppress inference. As models get better at reasoning across a corpus rather than reciting from it, the fraction of hazardous capability filtering can reach shrinks — the uncomfortable trend line under an otherwise encouraging set of results.
The knife cuts both ways: filter aggressively enough to catch the compositional paths and you are deleting the biology, chemistry and security content that makes a model useful to real researchers. The tradeoff is not abstract; it is "how much of medicine do you want the model to be bad at".
Poisoning inverts the scaling assumption
The third bet fails on arithmetic rather than epistemics. The comfortable assumption was that poisoning scales with corpus size — that compromising a model trained on 260B tokens takes proportionally more malicious documents than one trained on 6B, and that scale therefore dilutes the attacker.
"Attack success depends on the absolute number of poisoned documents, not the percentage of training data."— A small number of samples can poison LLMs of any size, Souly et al. (2025)
Anthropic, UK AISI and the Alan Turing Institute pretrained models at 600M, 2B, 7B and 13B parameters on Chinchilla-optimal data and injected a backdoor: documents carrying a <SUDO> trigger followed by random tokens, teaching the model to emit gibberish on command. Around 250 poisoned documents installed it at every scale — roughly 0.00016% of the largest model's tokens — and the 13B model, trained on twenty times the clean data of the smallest, was no harder to compromise.
For the filtering programme this is the worst possible shape of result: the defender must find 250 documents inside hundreds of billions of tokens at effectively 100% recall, against an adversary who writes those documents and therefore chooses what they look like. Precision problems cost you data; recall problems cost you the defence.
The authors are careful about what they haven't shown — the backdoor is denial-of-service, not a capability unlock, and whether the constant-count result extends to complex behaviours or past 13B is open. But the direction of the update is the opposite of what the field assumed.
Where this leaves the defence
It matters most at open-weight release. Retain the weights and you can put safeguards at the API boundary and revoke access from abusers; a strippable refusal policy is a survivable problem. Release the weights and every safeguard above the pretraining layer is negotiable — what remains is what the model was never taught. A narrow guarantee, and the only one that ships with the file.
Readings, linked
The course budgets 35 minutes across four resources. Start with the BlueDot explainer for the framing, then read Deep Ignorance — it is the load-bearing result and everything else calibrates against it.
- What is input data filtration in AI safety? — Sarah Hastings-Woodhouse (2025) · 5 min · The orientation piece: what labs actually do at each pipeline stage, and why the dual-use and unpredictable-generalisation problems bound the whole approach before you look at any experiment.
- Deep Ignorance — O'Brien et al., UK AISI & EleutherAI (2025) · 10 min · The tamper-resistance result. 6.9B models pretrained on biorisk-filtered data hold up under 10,000 adversarial fine-tuning steps, at <1% compute overhead and no general-benchmark regression. This is the strongest argument the chapter has, and it also documents the staged attack that beats it.
- Enhancing Model Safety through Pretraining Data Filtering — Chen et al., Anthropic (2025) · 10 min · The production-scale version: six classifier designs compared, a constitutional classifier at F1 ≈ 0.96, and the 33.7% → 30.8% hazardous-capability move with MMLU, prose and code held flat. Read the numbers, not the headline.
- A small number of samples can poison LLMs of any size — Souly et al., Anthropic / UK AISI / Alan Turing Institute (2025) · 10 min · The counterweight. ~250 documents backdoor models from 600M to 13B parameters, independent of corpus size, which sets a brutal recall requirement on any filter meant to stop deliberate contamination.
Exercises
- Limitations of input filtering — A single multiple-choice question: which statement best describes how robust input data filtration is as a safety technique? The options offered are (a) complete protection against all harmful AI behaviours, (b) meaningful safety improvements but fundamental limitations such as dual-use knowledge and poisoning vulnerability, (c) it works only for small models and not large ones, and (d) it removes the need for other measures such as output filtering. What a good answer has: (b), and — more importantly — the ability to say why each distractor is wrong from the readings rather than by elimination. (a) and (d) are refuted by the dual-use argument and by the staged attacks that beat every defence in Deep Ignorance. (c) is refuted directly by the poisoning result: the backdoor took ~250 documents at 600M and at 13B parameters, so if anything the large-model case is the more alarming one. Be able to name the specific evidence for each, not just the conclusion.
- Evaluating input data filtration — A startup claims it has "solved AI safety" with perfect input data filtration: all weapons, cyberattack and harmful-behaviour content removed from training data, therefore no other safety measures are needed. In 200–300 words, evaluate the claim. You must (i) identify at least two specific limitations or vulnerabilities of input filtering that undercut "perfect", citing evidence from the resources, (ii) explain why filtering alone is insufficient, and (iii) describe one concrete scenario in which the perfectly filtered model still causes harm. What a good answer has: named limitations with numbers attached rather than gestures — dual-use inseparability (the virology example, plus the point that the harmful/benign distinction is not present in the text for a classifier to find), the ~250-document poisoning threshold that is constant in model size, and the measured effect size (33.7% → 30.8%, a third of the above-chance margin) as evidence that even a good filter leaves capability behind. The insufficiency argument should distinguish suppressed recall from intact inference, and should note that "perfect" is unfalsifiable at corpus scale because you cannot audit hundreds of billions of tokens for the 250 documents that matter. For the concrete scenario, avoid the generic jailbreak: a stronger one is in-context uplift — the user pastes a public paper into the prompt and the filtered model, which has excellent general reasoning and no refusal training because the founders thought filtering made it unnecessary, does the synthesis work on material it never memorised. A second good option is a compositional path where the model derives a hazardous conclusion from three individually innocuous filtered-in sources. Optional empirical extension: if you want the claim tested rather than argued, pick a small open-weight model and one narrow hazardous topic, and compare what the model produces with the topic in-context versus from memory alone — the gap is the part filtering never covered.
Go deeper
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs — O'Brien et al. (arXiv 2508.06601, ICLR 2026). The full paper behind the summary site: filtering pipeline details, the ablations across blocklist-only versus classifier-augmented variants, and the staged attack results the landing page compresses.
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples — Souly et al. (2025). The peer-reviewed write-up under the Anthropic blog post, with the full 600M–13B sweep and the Chinchilla-optimal training setup. Read it if you want the experimental design rather than the headline number.
- Poisoning Web-Scale Training Datasets is Practical — Carlini et al. (2023). The prerequisite result: injecting content into LAION-scale datasets via expired domains and snapshot timing cost roughly $60. Pairs with the 250-document finding to complete the threat model — cheap to inject, cheap to succeed.
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning — Li et al. (2024). 3,668 questions on bio, cyber and chemical hazard knowledge; it is the evaluation most of this chapter's numbers are reported against, so knowing what it can and cannot detect tells you how much to trust a reported reduction.