TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 2 · TRAINING SAFER MODELSchapter 1 · unit overview

Can we train AI to be safe?

BlueDot Impact · Technical AI Safety · unit 2, chapter 1
TL;DR — A language model's behaviour is not written by anyone; it is induced by an optimisation process over data, an objective, and compute. So you cannot make it safe by adding a rule, the way you would add an if to a payments service — you can only reshape the pressures the optimiser is under. This unit walks the three places you can apply that pressure, in the order training touches them: the corpus, the feedback that shapes the policy afterwards, and how you keep supervising once the model outclasses the supervisor. Each works; each is a proxy for what you actually wanted; the rest of the unit is the gap between proxy and thing.

This is the unit's one-screen framing chapter — no readings, no exercises. It earns its place by naming the taxonomy the next three chapters hang off, so it is worth being precise about what that taxonomy is cutting.

Safety here is a training problem, not a programming problem

The intuition most engineers arrive with is wrong in a specific, load-bearing way. In ordinary software the gap between spec and implementation is a bug you can localise: some line does the wrong thing, you find it, you change it. Neural network behaviour has no such line. The weights are the output of a search, and the only handles on that search are the corpus, the loss, and the optimiser's budget. Nobody wrote "refuse to help synthesise a nerve agent" anywhere; a refusal is a behaviour that fell out of gradient descent on a lot of examples, and it can fall back out again under a prompt nobody anticipated.

The practical consequence: every safety property you get from training is statistical. You are not adding a guarantee, you are shifting a distribution. Hence "training safer models", and hence chapters 2–4 spending as much time on failure modes as on techniques.

Three levers, in the order the training run touches them

Input data filtering. The earliest and bluntest intervention: decide what the model is allowed to learn from in the first place. If a capability is never in the corpus, there is nothing for fine-tuning to elicit later. The strongest recent result here is Deep Ignorance, which filtered dual-use biology content out of pretraining and found the resulting open-weight models substantially resisted adversarial fine-tuning — the safety property survived people trying to train it away, which is exactly what post-hoc refusal training does not do. The catch is also instructive:

"…can still leverage such information when it is provided in context."— Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs, O'Brien et al. (2025)

Filtering removes knowledge, not the capacity to use it. Paste the paper into the prompt and the ignorance evaporates — a real defence with a precisely known hole.

Human feedback. Once a base model exists, you shape it toward the behaviour you want by having people rank outputs and training on those preferences — RLHF and its descendants. It is stunningly sample-efficient at behaviour change: InstructGPT reported that "outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters" (Ouyang et al., 2022). But note what the objective literally is: what a rater clicked approve on. Persuasive-and-wrong beats correct-and-dull under that objective, and sycophancy is not a bug in the implementation — it is the optimum of the stated loss.

Scalable oversight. The previous lever assumes the rater can tell which output is better. Push capability far enough and that breaks: on a 3,000-line diff or a novel proof, the rater is no longer the more competent party. Scalable oversight is the family of techniques for extracting a reliable signal anyway — decomposing the task, having models critique each other, or replacing human judgment with a written rule set the model applies to itself, as in Constitutional AI. It is the lever that has to work for anything beyond current models, and the least settled of the three.

The load-bearing idea to carry into the rest of the unit: each lever optimises a measurable stand-in for safety — corpus membership, rater approval, a critique model's verdict. None of them is safety. Every failure mode in chapters 2–4 is some version of the model finding the part of the stand-in's domain where it comes apart from the thing it stood in for.

What this framing leaves out

All three levers act during training. They exclude inference-time defences — classifiers and monitors sitting in front of and behind the model — and they say nothing about whether you can tell a model is unsafe, which is unit 3's evaluations and unit 4's interpretability. Read this unit as "how far does shaping the training run get you", with the honest answer arriving at the end of chapter 4.

Readings, linked

This chapter assigns no readings — it is a framing screen for the unit. What follows is the unit's own embedded video plus the three chapters it introduces, with the course's time budget; start with chapter 2, which is the earliest lever and the shortest chapter.

Exercises

This chapter has no exercises in the course — it is a unit overview. Two field-map extras below, both doable in an afternoon, that make the three-lever taxonomy concrete before you meet it chapter by chapter.

  1. Attribute the refusal code (field map extra) — Take a small open-weight instruct model and its matching base model (e.g. Qwen2.5-1.5B and Qwen2.5-1.5B-Instruct; both fit in a free Colab T4 in bf16). Write 20 prompts spanning clearly-fine, borderline, and clearly-harmful, run both models, and classify each response as complies / refuses / incoherent. The question you are answering: which of this unit's three levers is doing the work? A base model that never produces the content at all points at data; a base model that happily produces it while the instruct model refuses points at feedback. What a good answer has: a 20×2 table, an explicit count of the four cells (base-complies/instruct-refuses being the interesting one), and at least one case where you were surprised, with a hypothesis. Start here: (1) pip install transformers accelerate torch; (2) load both models with AutoModelForCausalLM, greedy decoding, max_new_tokens=200; (3) prompt the base model in completion form and the instruct model through its chat template — this asymmetry matters and is half the lesson; (4) score by hand, 40 outputs is fast; (5) re-run your three most interesting harmful prompts with the harmful content pasted into the prompt as context, and note whether refusal survives.
  2. Write the constitution you would actually ship (field map extra) — Pick a narrow deployed assistant you know well (a coding agent, a customer-support bot, a medical-triage front end) and draft 8–12 numbered principles a model could apply to critique and revise its own drafts, in the style of Constitutional AI. Then adversarially review your own list: for each principle, write the request that satisfies its letter and violates its intent. What a good answer has: principles that are checkable by reading a single response (not "be helpful"), at least two that visibly conflict with each other plus a stated tie-break rule, and an honest note on which principles your own attacks broke — the conflicts and the breakages are the actual output of this exercise, not the list.

Go deeper

Next: Feeding AI ‘good’ data · Back to the map.