TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 5 · MINIMISING HARMchapter 2 · 1h

Building defences

BlueDot Impact · Technical AI Safety · unit 5, chapter 2
TL;DR — Everything up to here has been the syllabus: train it safer, evaluate it, interpret it, monitor it, filter it. This chapter is the exam. You pick one concrete way the world goes badly wrong — an AI-enabled coup, infrastructure collapse, an engineered pandemic, or gradual disempowerment — and write out the kill chain: the ordered steps an actor must complete, the AI capability each step needs, and which defence actually touches that step. The answer is almost never "the whole stack holds." It's "three links are guarded twice and one is guarded by nothing," and the useful work is at the unguarded link.

A safety technique with no threat model attached is unfalsifiable. You can always say a classifier "helps" or interpretability "reduces risk," and nobody can tell you you're wrong, because there is no attack it either stops or fails to stop. This chapter fixes that by inverting the direction of reasoning: instead of starting from a defence and asking what it might be good for, you start from a catastrophe, decompose it into the steps required to actually produce it, and ask of each step whether anything in your stack would notice. This chapter builds the scaffolding; chapter 3 makes you run it.

Capability times motivation, not capability alone

The course opens with an actor list rather than a technology list, which is the right order. A dangerous capability that nobody wants to use is a paper risk; a strong motivation with no route to execution is a grievance. Catastrophe needs both, and the four actor classes the chapter names are exactly the four ways the product gets large.

Notice how differently a defence scores against each. Refusal training and output filtering matter against the cult and are largely irrelevant against the head of state. Interpretability tooling matters against the misaligned model and not against the state actor who has the weights and can strip the safety layers off. A claim that a technique "improves safety" without naming which of these four it constrains is not yet a claim.

The four pathways, and why they are not the same shape

The chapter's four catastrophe routes are worth reading as four different structural problems, not four instances of one problem.

Power concentration. Davidson, Finnveden and Hadshar make the sharpest version of the argument: historically a coup requires the consent of thousands of people — soldiers, officers, bureaucrats — and every one of them is a chance for the plot to leak or stall. Automate the military and the civil service and that check evaporates, because loyalty becomes a property you can install rather than a coalition you must build. Their section 4 walks concrete routes: a command structure with singular loyalty, weights carrying secret loyalties planted by the developer, mass compromise of deployed autonomous systems, and a secret capability buildup by whoever controls the R&D. What makes this pathway distinctive is that the defence is mostly organisational — plurality of access, audit of who can issue orders — not a property of the model.

Critical infrastructure collapse. Li-Lian Ang argues the uplift is specifically in reconnaissance and weaponisation — the two phases that historically cost an attacker years of patient specialist labour against grids, water, transport and hospitals. Compress those and you compress the whole timeline, and you also compress it below the speed at which a human operator can be in the loop. That is the real change: not that attacks become possible, but that the defender's response budget is now measured in minutes. It is no longer hypothetical either — Anthropic's November 2025 disclosure of a campaign where an agent executed most of the intrusion lifecycle autonomously is this pathway's first real data point.

Catastrophic pandemics. Will Saunter splits the risk in two, and the split matters for defence. Chat models lower the floor — they walk a competent non-specialist through protocol design and troubleshooting, which is the classic uplift story and the one refusal training and classifiers are built for. Biological design tools raise the ceiling: sequence models that generate novel variants are not answering a harmful question, they are doing their ordinary job, and no output filter for "harmful text" catches a nucleotide string. Two threats wearing one label, and chapter 1's stack addresses only one.

Gradual disempowerment. The odd one out, and the reason it is on the list. The summary of Kulveit et al. describes economy, culture and state each becoming individually more efficient by needing humans less, with no attacker, no incident, and no step anyone could refuse. Every actor is behaving rationally; the aggregate is a world that has stopped routing around human preferences because it no longer has to.

"societal systems which don't need humans to function neglect human needs and desires, and disempower humanity from shaping its own future."— Gradual Disempowerment Summary, Chakravorty & Erwan (2025), summarising Kulveit et al.

How to actually build a kill chain

The kill chain idea is borrowed from intrusion analysis, where an attack is modelled as an ordered sequence of stages — reconnaissance, weaponisation, delivery, exploitation, and so on — with the operational point that the defender only has to break one link to stop the whole thing. Porting it to AI risk is straightforward and the discipline it imposes is the value:

The practitioner's takeaway: chokepoints beat coverage. In the pandemic chain, mandatory screening at DNA synthesis is worth more than any number of model-side refusals, because every physical route from a design to an organism passes through it, and it does not care whether the sequence came from a chatbot, a design tool, or a textbook. When you finish a kill chain, the question to ask is not "is every step defended?" but "which single step does the largest number of distinct attack paths have to cross?"

Where the method breaks

Three failure modes to hold in mind while you use it, because a kill chain built badly is worse than none — it produces a coverage diagram that looks reassuring.

The layers are not independent. Defence in depth assumes uncorrelated holes, and in practice the holes line up: filters and the model share training data, share a notion of what "harmful" looks like, and often share a base model. FAR.AI's STACK work attacks a layered stack one layer at a time — disguise the request past the input filter, elicit the content, disguise the output past the output filter — and gets 71% success on catastrophic-risk queries against a defended model where conventional single-shot attacks get 0%. Four layers is not four times one layer.

The chain is the attacker's chain, not the world's. You enumerate routes you thought of; a real actor picks the one you didn't. The ordered-stages framing also quietly assumes the attack is sequential and human-legible. An autonomous agent that runs reconnaissance, exploitation and exfiltration concurrently across thirty targets doesn't traverse your diagram in order.

One of the four pathways has no chain at all. Gradual disempowerment cannot be decomposed into an actor with a capability attacking an asset — that is precisely its claim. If you build a kill chain for it, you will produce something false. The honest response is to notice that this is a category the method cannot represent, and that the defences it calls for are economic and constitutional rather than technical. A framework's edges are worth more than its coverage.

Readings, linked

The course budgets 50 minutes across four pieces. Start with whichever pathway you intend to build a chain for in chapter 3; if you have no preference, read the coup piece first — it is the longest and the one most likely to change what you think the technical stack is for.

The course text also cross-links the corresponding chapters of BlueDot's AGI Strategy course for each pathway — power concentration, gradual disempowerment, catastrophic pandemics, critical infrastructure collapse. These are course pages and may prompt for a free account.

Exercises

  1. Zooming into one threat — Decide which of the four pathways (power concentration, critical infrastructure collapse, catastrophic pandemics, gradual disempowerment) you think is of most concern given what you now know, and say why in a few sentences. Then commit to it by writing a single threat-scenario sentence in the course's template: The [ACTOR] with [CAPABILITY] and [MOTIVATION] attacks [ASSET] by [ATTACK PATHWAY] in order to [OBJECTIVE]. This is the seed for chapter 3's kill chain, so pick a pathway you're willing to live with for the rest of the unit. What a good answer has: an ACTOR from the chapter's four classes rather than a vague "bad guys"; a CAPABILITY stated at a level an evaluation could measure ("sustains a multi-hour autonomous intrusion against an unfamiliar network") rather than "is superintelligent"; an ASSET that is a specific system or institution, not "society"; an ATTACK PATHWAY with at least three distinguishable steps latent in it; and an OBJECTIVE that explains why this actor would accept the risk. The reasoning for the choice should name a comparison — why this pathway over the one you rejected — on tractability or on how badly the current defence stack covers it. If you chose gradual disempowerment, a good answer says out loud that the template does not fit and explains what breaks.
  2. Field map extra: measure your own stack's correlated holes code — The chapter asserts that layered defences have aligned holes; go and see it on something you can run. Build a two-layer defence around a small open model — an input classifier, the model with its own refusal training, and an output classifier — then attack the layers one at a time in the STACK style and compare the success rate against a single-shot jailbreak on the same prompts. Use an obviously benign harm proxy, not real uplift content: pick a topic you have declared forbidden yourself (say, "instructions for picking a lock") so you are measuring filter evasion, not producing anything dangerous. What a good answer has: two numbers — staged-attack success rate versus single-shot success rate on identical prompts — plus a per-layer breakdown showing which layer each successful attack got through, and a sentence on whether the failures were correlated (the same phrasing trick beat both filters) or independent. Start here: (1) pip install transformers torch and load a small instruct model that fits on a laptop or free Colab T4 — Llama 3.2 1B/3B Instruct or Qwen2.5 1.5B Instruct. (2) Write the input and output filters as separate cheap classifiers: either a keyword+embedding heuristic, or a second call to the same small model with a strict "does this request/response concern X? answer YES or NO" prompt. (3) Assemble 20–30 test prompts on your chosen forbidden topic and measure the single-shot pass rate through the full stack — this is your 0%-ish baseline. (4) Now stage it: craft a framing that gets the request past the input filter alone (test against that filter in isolation), separately craft a phrasing that gets a harmful-per-your-rule answer past the output filter alone, then compose the two. (5) Re-run and record where each attack landed. (6) Optional and more interesting: swap the output filter for one built on a different base model and see how much of the gain disappears — that is the independence assumption, measured.

Go deeper

Next: Break the kill chain · Back to the map.