TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 2 · TRAINING SAFER MODELSchapter 2 · 50 min

Feeding AI ‘good’ data

BlueDot Impact · Technical AI Safety · unit 2, chapter 2
TL;DR — The cheapest safety intervention available is refusing to show the model something in the first place, and it works better than its reputation: filtering bioweapons-adjacent text out of pretraining produces safeguards that survive 10,000 steps of adversarial fine-tuning, at under 1% extra compute and no hit to general benchmarks — more than can be said for refusal training, which an attacker strips in an afternoon. But it is a layer, not a solution. It cannot separate dual-use knowledge from harmful application, because the harm lives in the user's intent rather than in the document; and ~250 malicious documents backdoor a model regardless of size, so a filter needs near-perfect recall against an adversary who need only slip a rounding error past it. Filtering constrains what a model has memorised, not what it can infer, and not what someone deliberately planted.

Unit 2 walks the training-time levers roughly in the order they apply, and this is the earliest: before any objective is chosen or any human rates any output, someone decides which documents the model will ever see. The chapter tests the field's most intuitive safety proposal — "just don't train it on the bad stuff" — against what the 2025 experimental literature found. The answer is more interesting than either camp predicted: real, measurable, unusually durable, and nowhere near sufficient.

Three bets wearing one name

"Input data filtration" is used as though it were a single technique, but the chapter is really describing three separate engineering problems that happen to share a mechanism.

They fail differently — knowledge structure bounds the first, base rates bound the third — and collapsing them into one technique is how you end up with the startup pitch in exercise 2.

The tamper-resistance argument

The strongest case for filtering is not that it removes more capability than other methods. It is that what it removes stays removed.

Contrast the default. Refusal training — RLHF, safety fine-tuning, whatever the lab calls it — leaves hazardous knowledge intact in the weights and adds a learned policy of declining to emit it. That policy is a thin, late-stage layer over a large pretrained substrate, and about as robust as you'd expect: anyone with the weights and a modest GPU budget fine-tunes it off in a few hundred steps on innocuous examples. For open-weight releases, the safety property you shipped is not the one the downstream user runs.

Deep Ignorance (UK AISI and EleutherAI) tested the alternative by pretraining 6.9B-parameter models from scratch on corpora with biorisk content removed — a blocklist pass dropping about 8.4% of the data, then a fine-tuned classifier pass. The filtered models scored near chance on biorisk evaluations while holding general benchmarks, at under 1% additional training FLOPs. The headline is adversarial: the researchers then attacked their own models with up to 10,000 fine-tuning steps on 300M tokens of exactly the data they had filtered out, and the models stayed well below the unfiltered baseline. You can teach a filtered model some of what it missed, but you pay full price and never get back to the unfiltered starting point.

That asymmetry is the real product. Filtering converts a safety property from "the model has been persuaded not to" into "the model has less to work with", and only the second kind survives contact with someone who controls the weights.

The paper is honest about the ceiling: staged attacks combining fine-tuning with in-context retrieval — hand the hazardous material to the model at inference time and let it reason — defeated every defence tested, filtering included. A model that lacks knowledge can still be given it.

What "33%" actually buys you

Anthropic's pretraining data filtering work is the industrial-scale version of the same experiment, and it rewards reading the numbers rather than the summary. They built classifiers to flag CBRN content, trained a "constitutional" variant to F1 ≈ 0.94–0.96, filtered a pretraining corpus, and measured the resulting model on hazardous-knowledge and ordinary benchmarks.

The reported effect is a 33% relative reduction in harmful capability measured against the random-chance baseline: accuracy on the hazardous evaluation moved from 33.7% to 30.8%. On a multiple-choice benchmark where guessing scores around 25%, that is roughly three absolute points — about a third of the model's above-chance margin. MMLU natural science, prose and code were unaffected.

Read carelessly, "33% reduction in dangerous capabilities" sounds like a third of the risk is gone. Read carefully, it says something more modest and more useful: the filter is cheap, measurable, and costs nothing in general capability — an excellent property for a defence-in-depth layer and a terrible one for a load-bearing layer. Note also how much work the evaluation does: a benchmark where the unfiltered model barely clears chance cannot show a large effect no matter how good the filter is. That is the general problem with measuring dangerous capabilities before they exist at scale.

Dual use is a boundary problem, not a labelling problem

The obvious objection to capability denial is that hazardous knowledge does not come in a separate box.

"The same facts about virology could help someone design a vaccine or a novel pathogen."— What is input data filtration in AI safety?, Sarah Hastings-Woodhouse (2025)

This is usually framed as a labelling difficulty — the classifier can't tell good virology from bad. That gives the classifier too much credit for the problem: the distinction being asked for is not in the text at all. A protocol for enhancing transmissibility reads identically to a gain-of-function reviewer and to an attacker. The harm is a property of the downstream act, and no classifier can extract a signal that isn't in its input.

The harder version: even having removed every document stating a hazardous fact, a capable model can reconstruct hazardous conclusions by composing benign ones — biochemistry plus lab technique plus literature-search skill is, in the limit, most of what the filtered document contained. Filtering suppresses recall of memorised material; it does much less to suppress inference. As models get better at reasoning across a corpus rather than reciting from it, the fraction of hazardous capability filtering can reach shrinks — the uncomfortable trend line under an otherwise encouraging set of results.

The knife cuts both ways: filter aggressively enough to catch the compositional paths and you are deleting the biology, chemistry and security content that makes a model useful to real researchers. The tradeoff is not abstract; it is "how much of medicine do you want the model to be bad at".

Poisoning inverts the scaling assumption

The third bet fails on arithmetic rather than epistemics. The comfortable assumption was that poisoning scales with corpus size — that compromising a model trained on 260B tokens takes proportionally more malicious documents than one trained on 6B, and that scale therefore dilutes the attacker.

"Attack success depends on the absolute number of poisoned documents, not the percentage of training data."— A small number of samples can poison LLMs of any size, Souly et al. (2025)

Anthropic, UK AISI and the Alan Turing Institute pretrained models at 600M, 2B, 7B and 13B parameters on Chinchilla-optimal data and injected a backdoor: documents carrying a <SUDO> trigger followed by random tokens, teaching the model to emit gibberish on command. Around 250 poisoned documents installed it at every scale — roughly 0.00016% of the largest model's tokens — and the 13B model, trained on twenty times the clean data of the smallest, was no harder to compromise.

For the filtering programme this is the worst possible shape of result: the defender must find 250 documents inside hundreds of billions of tokens at effectively 100% recall, against an adversary who writes those documents and therefore chooses what they look like. Precision problems cost you data; recall problems cost you the defence.

The authors are careful about what they haven't shown — the backdoor is denial-of-service, not a capability unlock, and whether the constant-count result extends to complex behaviours or past 13B is open. But the direction of the update is the opposite of what the field assumed.

Carry this away: pretraining filtration is the only training-time intervention an adversary with the weights cannot simply remove — which is why it earns its place even for a three-point benchmark move. Use it as the tamper-resistant floor under refusal training, output filtering and monitoring, scoped to narrow high-consequence domains where the dual-use cost is tolerable. Anyone calling it sufficient has confused a floor with a building.

Where this leaves the defence

It matters most at open-weight release. Retain the weights and you can put safeguards at the API boundary and revoke access from abusers; a strippable refusal policy is a survivable problem. Release the weights and every safeguard above the pretraining layer is negotiable — what remains is what the model was never taught. A narrow guarantee, and the only one that ships with the file.

Readings, linked

The course budgets 35 minutes across four resources. Start with the BlueDot explainer for the framing, then read Deep Ignorance — it is the load-bearing result and everything else calibrates against it.

Exercises

  1. Limitations of input filtering — A single multiple-choice question: which statement best describes how robust input data filtration is as a safety technique? The options offered are (a) complete protection against all harmful AI behaviours, (b) meaningful safety improvements but fundamental limitations such as dual-use knowledge and poisoning vulnerability, (c) it works only for small models and not large ones, and (d) it removes the need for other measures such as output filtering. What a good answer has: (b), and — more importantly — the ability to say why each distractor is wrong from the readings rather than by elimination. (a) and (d) are refuted by the dual-use argument and by the staged attacks that beat every defence in Deep Ignorance. (c) is refuted directly by the poisoning result: the backdoor took ~250 documents at 600M and at 13B parameters, so if anything the large-model case is the more alarming one. Be able to name the specific evidence for each, not just the conclusion.
  2. Evaluating input data filtration — A startup claims it has "solved AI safety" with perfect input data filtration: all weapons, cyberattack and harmful-behaviour content removed from training data, therefore no other safety measures are needed. In 200–300 words, evaluate the claim. You must (i) identify at least two specific limitations or vulnerabilities of input filtering that undercut "perfect", citing evidence from the resources, (ii) explain why filtering alone is insufficient, and (iii) describe one concrete scenario in which the perfectly filtered model still causes harm. What a good answer has: named limitations with numbers attached rather than gestures — dual-use inseparability (the virology example, plus the point that the harmful/benign distinction is not present in the text for a classifier to find), the ~250-document poisoning threshold that is constant in model size, and the measured effect size (33.7% → 30.8%, a third of the above-chance margin) as evidence that even a good filter leaves capability behind. The insufficiency argument should distinguish suppressed recall from intact inference, and should note that "perfect" is unfalsifiable at corpus scale because you cannot audit hundreds of billions of tokens for the 250 documents that matter. For the concrete scenario, avoid the generic jailbreak: a stronger one is in-context uplift — the user pastes a public paper into the prompt and the filtered model, which has excellent general reasoning and no refusal training because the founders thought filtering made it unnecessary, does the synthesis work on material it never memorised. A second good option is a compositional path where the model derives a hazardous conclusion from three individually innocuous filtered-in sources. Optional empirical extension: if you want the claim tested rather than argued, pick a small open-weight model and one narrow hazardous topic, and compare what the model produces with the topic in-context versus from memory alone — the gap is the part filtering never covered.

Go deeper

Next: Teaching AI right from wrong · Back to the map.