TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 5 · MINIMISING HARMchapter 3 · 50 min · exercises

Break the kill chain

BlueDot Impact · Technical AI Safety · unit 5, chapter 3
TL;DR — A threat scenario is a sentence; a kill chain is the ordered set of things that all have to go right for that sentence to come true. Staging it converts a scary story into an engineering problem, because every stage is a place a defender can stand. The payoff is finding the load-bearing capability — the one link with no cheap substitute — and aiming the course's three lever positions at it: prevent it in training, detect it with evals and interpretability, constrain it at inference with control and filtering. The method's limit is that it assumes a discrete adversary taking discrete steps, which is exactly the shape gradual-disempowerment and inside-the-defence failures do not have.

This chapter is the unit's assembly step, and it assigns no new readings. Units 2 through 5 handed you a toolbox — safer training data, RLHF and Constitutional AI, dangerous-capability evals, interpretability probes, monitoring, filtering, AI control — and chapter 2 handed you a threat scenario you wrote yourself. The question nobody answers by accident is: against my specific threat, which tool goes where, and does the stack hold? The kill chain is the data structure that makes that answerable instead of rhetorical.

What a kill chain actually is

The idea is borrowed from network defence. Lockheed Martin's Cyber Kill Chain answered a specific institutional failure: security teams were writing incident reports about individual malware samples and learning nothing transferable. It reframed an intrusion as a sequence — reconnaissance, weaponisation, delivery, exploitation, installation, command-and-control, actions on objectives — making the sequence the unit of analysis rather than the payload. That lets you say something a per-incident report never does: the attacker had to complete all seven stages, and I only had to break one.

That asymmetry inverts the usual attacker-advantage framing. Under "the attacker only has to be right once", defence is hopeless; under the kill chain, the attacker has to be right at every stage in order and the defender picks whichever stage is cheapest to instrument. Neither is wrong — they describe different games. The kill chain applies when the harm requires a chain of successes rather than one lucky shot, which is true of every catastrophic pathway the course names: a coup, an engineered pandemic, an infrastructure takedown. None of those is one prompt. MITRE later generalised the same move into ATT&CK and, for machine-learning systems specifically, ATLAS; borrow their stage names rather than inventing your own.

Scenario versus chain: the difference that does the work

At the end of chapter 2 you wrote a scenario in the form The [ACTOR] with [CAPABILITY] and [MOTIVATION] attacks [ASSET] by [ATTACK PATHWAY] in order to [OBJECTIVE]. That sentence is a claim about the world: falsifiable in principle, useless in practice, because it has no internal joints — nowhere to put a defence, and no way to argue about whether the defence would work.

Staging it gives it joints. Say a well-resourced group wants a frontier model to shorten the path from published virology literature to a working pathogen. As a chain that becomes: acquire model access that is not rate-limited or attributable; elicit synthesis knowledge the model was trained to refuse; convert it into an executable protocol; acquire the physical materials past screening; do the wet-lab work with troubleshooting; release. Now there are six arguments instead of one, with different experts and different owners. Access control is a deployment question, elicitation a jailbreak-robustness question, and materials screening not an AI question at all — which is itself a finding, because some of your best choke points sit entirely outside the model.

The discipline that makes staging honest is writing, per stage, three things: what the attacker does, what has to be true for it to succeed, and what a defender could observe if it were happening. The third column is the one people skip and the one that decides whether the chain is actionable. A stage nobody can observe is not a choke point; it is a hope.

Finding the load-bearing link

The chapter's sharpest prompt is whether removing one capability collapses the entire threat. Read literally that is usually too strong — real chains have redundancy — but as a search heuristic it is excellent, because it forces you to distinguish two things that look alike on a diagram. A stage is conjunctive if the attacker needs it and has no substitute: an AND gate. It is disjunctive if there are three other ways to get the same effect: an OR gate. Defences on OR-gate stages buy almost nothing, because the attacker reroutes for the cost of an afternoon; defences on AND-gate stages buy the whole chain. So the question is not "which capability sounds scariest" but "which has no cheap substitute" — and the test is a counterfactual you run adversarially against yourself: if I successfully deny this, what does a competent attacker do instead, and what does it cost them? If the answer is "use an open-weights model with the safeguards stripped", your choke point is a speed bump, and it is much better to learn that on paper.

The second filter is where the choke point is cheapest to hold. Anthropic's Constitutional Classifiers work reports that arithmetic honestly: on 10,000 synthetic jailbreak prompts the unguarded success rate was 86%, the guarded rate 4.4%, at the price of a 0.38% increase in refusals and a 23.7% increase in compute. Those four numbers are what a choke point looks like once you cost it out — effectiveness, false-positive burden, unit economics. A defence you cannot afford to run on every request is not deployed, and a defence that refuses 5% of legitimate traffic gets turned off by whoever owns the revenue.

"Historically, jailbreaks have proved difficult to detect and block: these kinds of attacks were described over 10 years ago, yet to our knowledge there are still no fully robust deep-learning models in production."— Constitutional Classifiers: Defending against universal jailbreaks, Anthropic (2025)

Three lever positions: prevent, detect, constrain

Once you have the load-bearing capability, the chapter asks you to attack it from three directions — which are the course's earlier units re-sorted by when they act.

Prevent acts before the capability exists: data filtering, unlearning and refusal training from unit 2. It is the only lever that removes risk rather than managing it, and the most brittle, because capabilities are entangled — the virology you would like the model not to know is the virology that makes it useful to vaccine researchers. Prevention also fails silently: you cannot easily prove absence, and a capability that was suppressed rather than removed can be elicited by a better prompt or restored by a few hundred fine-tuning steps.

Detect acts on the trained artefact: dangerous-capability evaluations from unit 3, interpretability probes from unit 4, runtime monitoring from unit 5.1. Detection is where most of the field's effort sits, and it has a structural weakness the exercise will surface if you are honest: base rates. A monitor with 99% accuracy on a stream where 1 in 100,000 actions is an attack raises roughly a thousand false alarms for every true one. At production volume that is not a monitor, it is a denial-of-service on your review team. This is why the useful published work measures catch rate at a fixed alarm budget — see Anthropic's work on cheap monitors — and why a tiered design (cheap model screens everything, expensive model adjudicates the top 1%) is usually the only affordable shape.

Constrain acts at the moment of use and assumes the other two levers already failed: permissions, sandboxing, human approval for irreversible actions, task decomposition so no single instance holds the whole plan, resampling suspicious trajectories. This is AI control; its defining move is dropping the assumption that the model is on your side.

"Researchers have not evaluated whether such techniques still ensure safety if the model is itself intentionally trying to subvert them."— AI Control: Improving Safety Despite Intentional Subversion, Greenblatt, Shlegeris, Sachan & Roger (2023)

That distinction changes how you evaluate. Prevention and detection can be scored against average behaviour; a constraint has to be scored against a red team that knows your protocol — which is why the control literature reports a game between a blue-team protocol and a best-response attack policy, not a benchmark number.

The one thing to carry away: write the counterfactual next to every defence you propose. "Add a classifier at stage 3" is not a finding. "Add a classifier at stage 3; the attacker's cheapest reroute is an open-weights model, costing them a month and X capability" is a finding — and it is the sentence that tells you whether to build the classifier or go work on something else.

Where the method breaks

Three failure modes are worth naming before you trust your own chain.

It assumes an adversary with stages. The most-discussed structural risk in this unit's readings, gradual disempowerment, has no attacker, no reconnaissance and no moment of exploitation — it is the aggregate of many locally rational decisions to hand more of the economy and the state to systems not accountable to humans. There is no link to break because there is no chain, only a gradient. Pick that pathway and you will find the template fighting you; the right response is to say so rather than force stages onto it.

The defender may be inside the chain. When the threat actor is a misaligned model rather than a human using one, several of your defences are implemented by models drawn from the same distribution as the threat — the monitor, the classifier, the summariser feeding your dashboard. Correlated failure is the default assumption, not the pessimistic one. The standard mitigation is a trust asymmetry: a weaker but trusted model watches a stronger untrusted one.

Capability lists drift into wishlists. "Long-term planning" is not a capability, it is a genre. A capability you can defend against has an operationalisation — a task, a threshold, a measurement procedure. If yours cannot in principle become an eval with a pass/fail line, you cannot detect it, and two of your three levers are gone before you start.

Readings, linked

The course assigns no new readings for this chapter — the full 50 minutes is exercise time. What follows is the material the exercise actually draws on, all of it from earlier in unit 5; if you are picking one thing up first, make it the AI control overview, since it supplies the prevent / detect / constrain vocabulary the third exercise is built on.

Exercises

Three exercises are served by the course; all three build on the threat-scenario sentence you wrote at the end of chapter 2. A fourth, clearly labelled as a field-map extra, turns the third one's hand-waving about detection into a number. You can do all four without a BlueDot login; the course's only role there is saving your answers.

  1. Step-by-step breakdown code — Take your threat scenario and break it into the stages an attacker would actually execute, in order: reconnaissance, delivery, exploitation, persistence, action on objectives (rename or add stages if your pathway needs them — an engineered-pandemic chain wants an "acquire physical materials" stage that has no cyber analogue). The point of the transformation is that your scenario says what might happen while the chain says how it unfolds, and only the second form exposes choke points. Use the course's Notion template, or reproduce it as a five-column table. What a good answer has: one row per stage, and for each stage three filled cells — the attacker's action, the precondition that must hold for it to succeed, and the observable a defender could see if it were happening right now. Stages are in causal order and each one's precondition is genuinely produced by the stage above it. At least one stage is marked as sitting outside the AI system entirely; at least one is marked disjunctive (the attacker has substitutes) versus conjunctive (they do not). If your chosen pathway is gradual disempowerment, the correct answer is an argument for why the template does not apply, plus whatever partial structure you can salvage. Start here: (1) make it a machine-readable artefact rather than prose — a YAML or CSV file with fields stage, action, precondition, observable, gate: and|or, defence, counterfactual_cost; (2) validate it in ten lines of Python that assert every row has all seven fields and that the file has ≥4 stages, so an incomplete chain fails loudly; (3) render it with a tiny script into a Mermaid flowchart (graph TD, one node per stage, AND-gates as solid arrows and OR-gates as dashed) so the shape of the chain is visible at a glance; (4) keep the file — exercises 2 to 4 all write into it.
  2. Capabilities required for harm — List three to five specific technical capabilities or behaviours your threat requires the AI to have. Be concrete about what the system would need to be able to do, not about how bad the outcome is. The course's own examples set the register: self-replication, goal persistence across instances, long-horizon planning and coordination, resource acquisition (compute, money), bioweapon knowledge and synthesis support. What a good answer has: every capability is attached to at least one stage from exercise 1, so you can see which stage dies if the capability is absent. Every capability is written so that an evaluation could in principle return pass or fail on it — "can reconstruct a synthesis protocol from public literature to the point where a technician with a BSL-2 lab succeeds" rather than "knows biology". None of them is a restatement of the objective (if one of your capabilities is "cause a pandemic", it is a goal, not a capability). Note which capabilities today's deployed models plausibly already have, which are contested, and which are clearly absent — that split is the difference between a mitigation you need this year and one you are forecasting.
  3. Building defences — Pick the single most important capability from your list and interrogate it: if we prevented just this one, would the whole threat collapse? Then, drawing on everything the course has covered, work out three defences against it — how to stop the model acquiring it during training, how to detect it once it exists, and how to constrain the model so it cannot use the capability to take dangerous actions. What a good answer has: an explicit verdict on whether the threat really collapses, and if it only degrades, an estimate of by how much and what the attacker's reroute is. Then one concrete technique per lever, each named from the course rather than invented — for prevent, data filtering or unlearning or refusal training from unit 2; for detect, a dangerous-capability eval from unit 3, an interpretability probe from unit 4, or an input/output classifier from 5.1; for constrain, a control protocol from 5.1 (trusted monitoring, human approval on irreversible actions, task decomposition, resampling). Each defence carries its failure mode: what an adversary who knows the defence exists does about it. A strong answer ends by naming the residual gap — the part of the chain that all three levers miss — because that gap is the actual research agenda, and it is the thing unit 6 is going to ask you to go work on.
  4. Field map extra: cost your monitor code — Exercise 3 lets you write "detect it with a monitor" and move on. This makes you check whether that monitor survives contact with production volume. Pick the one detection defence you proposed, build a small version of it, and report its catch rate at a fixed false-positive budget rather than its accuracy. What a good answer has: an explicit deployment volume (say 10 million requests per day), an explicit assumed base rate for the attack stage you are monitoring, and the resulting confusion matrix in absolute counts — how many true positives, how many false alarms per day, and how many human review-hours that implies at a minute apiece. Then a sentence on whether the defence is deployable, and if not, what tiering or threshold change would make it so. Start here: (1) build a 200-item labelled set — 100 benign requests that superficially resemble your attack stage (a virologist asking legitimate questions, an admin doing legitimate privileged operations) and 100 positives you write by hand or generate; the hard negatives are the whole experiment, since separating attacks from obviously-unrelated traffic is trivial and proves nothing. (2) Use a small instruct model as the trusted monitor — Qwen2.5-1.5B-Instruct or Llama-3.2-1B-Instruct via HuggingFace transformers runs on a free Colab T4 or a laptop CPU at this scale — prompted to return a 0–100 suspicion score; a cheap API model works too if you prefer. (3) Score all 200 items, then use sklearn.metrics.roc_curve to get TPR at FPR = 0.01 and 0.001 rather than picking a threshold by eye. (4) Multiply out to your deployment volume and write down the daily false-alarm count. (5) Optional but where the insight is: add a second tier — send only the top 1% by suspicion to a larger adjudicator model — and recompute both the catch rate and the compute cost, which is the same tiered-defence trade-off Anthropic reports in the cheap monitors work. (6) Write the number into the defence column of your exercise-1 file.

Go deeper

Next: Your next steps · Back to the map.