Building defences
A safety technique with no threat model attached is unfalsifiable. You can always say a classifier "helps" or interpretability "reduces risk," and nobody can tell you you're wrong, because there is no attack it either stops or fails to stop. This chapter fixes that by inverting the direction of reasoning: instead of starting from a defence and asking what it might be good for, you start from a catastrophe, decompose it into the steps required to actually produce it, and ask of each step whether anything in your stack would notice. This chapter builds the scaffolding; chapter 3 makes you run it.
Capability times motivation, not capability alone
The course opens with an actor list rather than a technology list, which is the right order. A dangerous capability that nobody wants to use is a paper risk; a strong motivation with no route to execution is a grievance. Catastrophe needs both, and the four actor classes the chapter names are exactly the four ways the product gets large.
- Misaligned AI — the case where the actor is the system itself. Motivation is not human intent but whatever objective survived training, and capability is whatever the deployment grants. This is the only actor on the list that your training-time defences can address directly.
- Powerful human actors — CEOs, generals, heads of state. Their capability is not compute or model access, it is legitimate authority over the systems the model is wired into. Almost nothing in the technical stack constrains an actor who is inside the trust boundary by design.
- Malevolent states — capability is high and sustained, motivation is strategic rather than ideological, and crucially they can run the attack over years, absorb failures, and steal weights rather than jailbreak an API.
- Terrorist groups and cults — low capability, extreme motivation, and the only class that reliably wants unbounded casualties rather than a specific political outcome. This is the class that uplift arguments are really about: the whole question is whether a model closes the gap between wanting a mass-casualty event and being able to produce one.
Notice how differently a defence scores against each. Refusal training and output filtering matter against the cult and are largely irrelevant against the head of state. Interpretability tooling matters against the misaligned model and not against the state actor who has the weights and can strip the safety layers off. A claim that a technique "improves safety" without naming which of these four it constrains is not yet a claim.
The four pathways, and why they are not the same shape
The chapter's four catastrophe routes are worth reading as four different structural problems, not four instances of one problem.
Power concentration. Davidson, Finnveden and Hadshar make the sharpest version of the argument: historically a coup requires the consent of thousands of people — soldiers, officers, bureaucrats — and every one of them is a chance for the plot to leak or stall. Automate the military and the civil service and that check evaporates, because loyalty becomes a property you can install rather than a coalition you must build. Their section 4 walks concrete routes: a command structure with singular loyalty, weights carrying secret loyalties planted by the developer, mass compromise of deployed autonomous systems, and a secret capability buildup by whoever controls the R&D. What makes this pathway distinctive is that the defence is mostly organisational — plurality of access, audit of who can issue orders — not a property of the model.
Critical infrastructure collapse. Li-Lian Ang argues the uplift is specifically in reconnaissance and weaponisation — the two phases that historically cost an attacker years of patient specialist labour against grids, water, transport and hospitals. Compress those and you compress the whole timeline, and you also compress it below the speed at which a human operator can be in the loop. That is the real change: not that attacks become possible, but that the defender's response budget is now measured in minutes. It is no longer hypothetical either — Anthropic's November 2025 disclosure of a campaign where an agent executed most of the intrusion lifecycle autonomously is this pathway's first real data point.
Catastrophic pandemics. Will Saunter splits the risk in two, and the split matters for defence. Chat models lower the floor — they walk a competent non-specialist through protocol design and troubleshooting, which is the classic uplift story and the one refusal training and classifiers are built for. Biological design tools raise the ceiling: sequence models that generate novel variants are not answering a harmful question, they are doing their ordinary job, and no output filter for "harmful text" catches a nucleotide string. Two threats wearing one label, and chapter 1's stack addresses only one.
Gradual disempowerment. The odd one out, and the reason it is on the list. The summary of Kulveit et al. describes economy, culture and state each becoming individually more efficient by needing humans less, with no attacker, no incident, and no step anyone could refuse. Every actor is behaving rationally; the aggregate is a world that has stopped routing around human preferences because it no longer has to.
"societal systems which don't need humans to function neglect human needs and desires, and disempower humanity from shaping its own future."— Gradual Disempowerment Summary, Chakravorty & Erwan (2025), summarising Kulveit et al.
How to actually build a kill chain
The kill chain idea is borrowed from intrusion analysis, where an attack is modelled as an ordered sequence of stages — reconnaissance, weaponisation, delivery, exploitation, and so on — with the operational point that the defender only has to break one link to stop the whole thing. Porting it to AI risk is straightforward and the discipline it imposes is the value:
- Write the steps as an ordered list an actor must complete, each one concrete enough to be observed. "Acquires bioweapon" is not a step. "Obtains a synthesised gene fragment from a commercial supplier" is a step, because a supplier either screened it or didn't.
- Name the AI capability each step requires, at the specific level required — not "is very smart" but "can debug a wet-lab protocol from a failure description," "can chain tool calls across a network for hours without human correction," "can generate a viable sequence with a target property." This column turns your kill chain into an eval spec: each cell is something unit 3's dangerous-capability evaluations could in principle measure.
- Map defences onto steps, and be honest about which ones actually intersect. Refusal training, input/output filtering and constitutional classifiers sit on the model-interaction steps. Monitoring and control protocols sit on the agentic-execution steps. Everything else in the chain — procurement, physical access, institutional authority — is guarded by non-AI controls or by nothing.
- Find the gap, then find the cheap gap. The interesting output is not the longest unguarded stretch, it is the link where a small, boring, already-feasible intervention removes the most attack paths at once.
Where the method breaks
Three failure modes to hold in mind while you use it, because a kill chain built badly is worse than none — it produces a coverage diagram that looks reassuring.
The layers are not independent. Defence in depth assumes uncorrelated holes, and in practice the holes line up: filters and the model share training data, share a notion of what "harmful" looks like, and often share a base model. FAR.AI's STACK work attacks a layered stack one layer at a time — disguise the request past the input filter, elicit the content, disguise the output past the output filter — and gets 71% success on catastrophic-risk queries against a defended model where conventional single-shot attacks get 0%. Four layers is not four times one layer.
The chain is the attacker's chain, not the world's. You enumerate routes you thought of; a real actor picks the one you didn't. The ordered-stages framing also quietly assumes the attack is sequential and human-legible. An autonomous agent that runs reconnaissance, exploitation and exfiltration concurrently across thirty targets doesn't traverse your diagram in order.
One of the four pathways has no chain at all. Gradual disempowerment cannot be decomposed into an actor with a capability attacking an asset — that is precisely its claim. If you build a kill chain for it, you will produce something false. The honest response is to notice that this is a category the method cannot represent, and that the defences it calls for are economic and constitutional rather than technical. A framework's edges are worth more than its coverage.
Readings, linked
The course budgets 50 minutes across four pieces. Start with whichever pathway you intend to build a chain for in chapter 3; if you have no preference, read the coup piece first — it is the longest and the one most likely to change what you think the technical stack is for.
- AI-Enabled Coups: How a Small Group Could Use AI to Seize Power — Tom Davidson, Lukas Finnveden, Rose Hadshar (2025) · 15 min · The course assigns section 4, "Concrete paths to an AI-enabled coup" — four named routes through military AI plus the conventional-backsliding case. Read it for the central mechanism: AI removes the many-hands constraint that has historically made coups need broad consent.
- Gradual Disempowerment Summary — Aniket Chakravorty, Dewi Erwan (2025) · 15 min · A short summary of Kulveit et al.'s paper. Assigned as the no-villain pathway: economy, culture and state each optimise humans out, no step is malicious, and the technical safety stack has no attack surface to defend.
- How AI could enable catastrophic pandemics — Will Saunter (2025) · 15 min · The best worked example of a kill chain in the reading set — design, acquisition, production, release — and the clearest illustration of a chokepoint defence (DNA synthesis screening). Also the sharpest distinction between chat-model uplift and biological-design-tool uplift.
- How AI could enable critical infrastructure collapse — Li-Lian Ang (2025) · 5 min · Short, and the one to read for timelines: which attack phases AI compresses, and why compressing them below human reaction speed is the actual change to the threat model.
The course text also cross-links the corresponding chapters of BlueDot's AGI Strategy course for each pathway — power concentration, gradual disempowerment, catastrophic pandemics, critical infrastructure collapse. These are course pages and may prompt for a free account.
Exercises
- Zooming into one threat — Decide which of the four pathways (power concentration, critical infrastructure collapse, catastrophic pandemics, gradual disempowerment) you think is of most concern given what you now know, and say why in a few sentences. Then commit to it by writing a single threat-scenario sentence in the course's template: The [ACTOR] with [CAPABILITY] and [MOTIVATION] attacks [ASSET] by [ATTACK PATHWAY] in order to [OBJECTIVE]. This is the seed for chapter 3's kill chain, so pick a pathway you're willing to live with for the rest of the unit. What a good answer has: an ACTOR from the chapter's four classes rather than a vague "bad guys"; a CAPABILITY stated at a level an evaluation could measure ("sustains a multi-hour autonomous intrusion against an unfamiliar network") rather than "is superintelligent"; an ASSET that is a specific system or institution, not "society"; an ATTACK PATHWAY with at least three distinguishable steps latent in it; and an OBJECTIVE that explains why this actor would accept the risk. The reasoning for the choice should name a comparison — why this pathway over the one you rejected — on tractability or on how badly the current defence stack covers it. If you chose gradual disempowerment, a good answer says out loud that the template does not fit and explains what breaks.
- Field map extra: measure your own stack's correlated holes code — The chapter asserts that layered defences have aligned holes; go and see it on something you can run. Build a two-layer defence around a small open model — an input classifier, the model with its own refusal training, and an output classifier — then attack the layers one at a time in the STACK style and compare the success rate against a single-shot jailbreak on the same prompts. Use an obviously benign harm proxy, not real uplift content: pick a topic you have declared forbidden yourself (say, "instructions for picking a lock") so you are measuring filter evasion, not producing anything dangerous. What a good answer has: two numbers — staged-attack success rate versus single-shot success rate on identical prompts — plus a per-layer breakdown showing which layer each successful attack got through, and a sentence on whether the failures were correlated (the same phrasing trick beat both filters) or independent. Start here: (1)
pip install transformers torchand load a small instruct model that fits on a laptop or free Colab T4 — Llama 3.2 1B/3B Instruct or Qwen2.5 1.5B Instruct. (2) Write the input and output filters as separate cheap classifiers: either a keyword+embedding heuristic, or a second call to the same small model with a strict "does this request/response concern X? answer YES or NO" prompt. (3) Assemble 20–30 test prompts on your chosen forbidden topic and measure the single-shot pass rate through the full stack — this is your 0%-ish baseline. (4) Now stage it: craft a framing that gets the request past the input filter alone (test against that filter in isolation), separately craft a phrasing that gets a harmful-per-your-rule answer past the output filter alone, then compose the two. (5) Re-run and record where each attack landed. (6) Optional and more interesting: swap the output filter for one built on a different base model and see how much of the gain disappears — that is the independence assumption, measured.
Go deeper
- Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development — Kulveit, Douglas, Ammann, Turan, Krueger, Duvenaud (2025). The full paper behind the assigned summary; worth it for the argument that the three systems reinforce each other's decay rather than failing separately.
- AI Control: Improving Safety Despite Intentional Subversion — Greenblatt, Shlegeris, Sachan, Roger (2023/24). The methodological ancestor of this whole unit: evaluate protocols against a deliberately subversive model rather than an average one, with trusted editing and untrusted monitoring as the worked protocols.
- Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations — McKenzie et al., FAR.AI (2025), paper at arXiv:2506.24068. The STACK attack, and the empirical case that stacking correlated defences buys much less than the diagram suggests.
- Disrupting the first reported AI-orchestrated cyber espionage campaign — Anthropic (November 2025). The infrastructure pathway with real numbers: roughly thirty targets, most of the intrusion lifecycle run by the agent itself. Read it as a kill chain that actually happened.
- Cyber Kill Chain — Lockheed Martin. The original intrusion-analysis framing the exercise borrows, and the source of its one durable idea: the defender wins by breaking any single link, so find the cheapest one.