Can we train AI to be safe?
This is the unit's one-screen framing chapter — no readings, no exercises. It earns its place by naming the taxonomy the next three chapters hang off, so it is worth being precise about what that taxonomy is cutting.
Safety here is a training problem, not a programming problem
The intuition most engineers arrive with is wrong in a specific, load-bearing way. In ordinary software the gap between spec and implementation is a bug you can localise: some line does the wrong thing, you find it, you change it. Neural network behaviour has no such line. The weights are the output of a search, and the only handles on that search are the corpus, the loss, and the optimiser's budget. Nobody wrote "refuse to help synthesise a nerve agent" anywhere; a refusal is a behaviour that fell out of gradient descent on a lot of examples, and it can fall back out again under a prompt nobody anticipated.
The practical consequence: every safety property you get from training is statistical. You are not adding a guarantee, you are shifting a distribution. Hence "training safer models", and hence chapters 2–4 spending as much time on failure modes as on techniques.
Three levers, in the order the training run touches them
Input data filtering. The earliest and bluntest intervention: decide what the model is allowed to learn from in the first place. If a capability is never in the corpus, there is nothing for fine-tuning to elicit later. The strongest recent result here is Deep Ignorance, which filtered dual-use biology content out of pretraining and found the resulting open-weight models substantially resisted adversarial fine-tuning — the safety property survived people trying to train it away, which is exactly what post-hoc refusal training does not do. The catch is also instructive:
"…can still leverage such information when it is provided in context."— Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs, O'Brien et al. (2025)
Filtering removes knowledge, not the capacity to use it. Paste the paper into the prompt and the ignorance evaporates — a real defence with a precisely known hole.
Human feedback. Once a base model exists, you shape it toward the behaviour you want by having people rank outputs and training on those preferences — RLHF and its descendants. It is stunningly sample-efficient at behaviour change: InstructGPT reported that "outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters" (Ouyang et al., 2022). But note what the objective literally is: what a rater clicked approve on. Persuasive-and-wrong beats correct-and-dull under that objective, and sycophancy is not a bug in the implementation — it is the optimum of the stated loss.
Scalable oversight. The previous lever assumes the rater can tell which output is better. Push capability far enough and that breaks: on a 3,000-line diff or a novel proof, the rater is no longer the more competent party. Scalable oversight is the family of techniques for extracting a reliable signal anyway — decomposing the task, having models critique each other, or replacing human judgment with a written rule set the model applies to itself, as in Constitutional AI. It is the lever that has to work for anything beyond current models, and the least settled of the three.
What this framing leaves out
All three levers act during training. They exclude inference-time defences — classifiers and monitors sitting in front of and behind the model — and they say nothing about whether you can tell a model is unsafe, which is unit 3's evaluations and unit 4's interpretability. Read this unit as "how far does shaping the training run get you", with the honest answer arriving at the end of chapter 4.
Readings, linked
This chapter assigns no readings — it is a framing screen for the unit. What follows is the unit's own embedded video plus the three chapters it introduces, with the course's time budget; start with chapter 2, which is the earliest lever and the shortest chapter.
- Technical AI Safety Unit 2: Training safer models — BlueDot Impact (video, embedded in the chapter) · short · The course's own spoken version of the three-lever framing on this page.
- Feeding AI 'good' data — BlueDot Impact · 50 min · Lever one: curation and filtering of the pretraining corpus, and what filtering can and cannot remove.
- Teaching AI right from wrong — BlueDot Impact · 1 h 10 min · Lever two: RLHF, preference models, and the failure modes of optimising rater approval.
- More safety techniques — BlueDot Impact · 1 h 25 min · Lever three and the leftovers: scalable oversight, constitutional methods, and where the whole stack still leaks.
Exercises
This chapter has no exercises in the course — it is a unit overview. Two field-map extras below, both doable in an afternoon, that make the three-lever taxonomy concrete before you meet it chapter by chapter.
- Attribute the refusal code (field map extra) — Take a small open-weight instruct model and its matching base model (e.g. Qwen2.5-1.5B and Qwen2.5-1.5B-Instruct; both fit in a free Colab T4 in bf16). Write 20 prompts spanning clearly-fine, borderline, and clearly-harmful, run both models, and classify each response as complies / refuses / incoherent. The question you are answering: which of this unit's three levers is doing the work? A base model that never produces the content at all points at data; a base model that happily produces it while the instruct model refuses points at feedback. What a good answer has: a 20×2 table, an explicit count of the four cells (base-complies/instruct-refuses being the interesting one), and at least one case where you were surprised, with a hypothesis. Start here: (1) pip install transformers accelerate torch; (2) load both models with AutoModelForCausalLM, greedy decoding, max_new_tokens=200; (3) prompt the base model in completion form and the instruct model through its chat template — this asymmetry matters and is half the lesson; (4) score by hand, 40 outputs is fast; (5) re-run your three most interesting harmful prompts with the harmful content pasted into the prompt as context, and note whether refusal survives.
- Write the constitution you would actually ship (field map extra) — Pick a narrow deployed assistant you know well (a coding agent, a customer-support bot, a medical-triage front end) and draft 8–12 numbered principles a model could apply to critique and revise its own drafts, in the style of Constitutional AI. Then adversarially review your own list: for each principle, write the request that satisfies its letter and violates its intent. What a good answer has: principles that are checkable by reading a single response (not "be helpful"), at least two that visibly conflict with each other plus a stated tie-break rule, and an honest note on which principles your own attacks broke — the conflicts and the breakages are the actual output of this exercise, not the list.
Go deeper
- Training language models to follow instructions with human feedback — Ouyang et al., OpenAI (2022). The InstructGPT paper; the canonical description of lever two and the reason every deployed chat model is post-trained this way.
- Constitutional AI: Harmlessness from AI Feedback — Bai et al., Anthropic (2022). Replaces the human labeller with a written rule set and a critique loop — the most deployed instance of lever three.
- Deep Ignorance — O'Brien et al. (2025). The strongest evidence that lever one buys something the other two do not: tamper-resistance against adversarial fine-tuning of open weights.
- Measuring Progress on Scalable Oversight for Large Language Models — Bowman et al., Anthropic (2022). Turns "can a human supervise a stronger model" into an experiment with a number, which is what makes lever three tractable at all.
- Weak-to-Strong Generalization — Burns et al., OpenAI Superalignment (2023). The empirical analogue of the superhuman-supervision problem: fine-tune a strong model on a weak model's labels and measure how much of the strong model's own competence you recover.