What might success look like?
Chapter 1 argued that advanced AI is a technical problem worth taking seriously. That leaves a hole: worth taking seriously toward what? Every research agenda later in this course — interpretability, evaluations, scalable oversight, control — is implicitly a bet on a story about how this ends well. This chapter makes those stories explicit, because an agenda that is excellent under one is close to useless under another.
The question is strategic, not technical
Imagine a lever with real influence over how frontier AI gets built for the next decade — a research budget, a government, or just your own career. What do you do with it? The technical literature answers narrower questions ("can we tell whether this model is deceiving us"), not this one. Two variables generate the field's three answers. Coordination feasibility: can enough of the world's compute, talent and money be brought under common rules for long enough to matter? Alignment tractability: is making a much-smarter-than-human system reliably do what we intended a problem that yields to normal scientific effort, or something we cannot currently even state? Yes to both gives camp 1. No to coordination, yes to tractability gives camp 2. No to tractability leaves coordination as the only move — camp 3.
Camp 1 — build it slowly and safely
This camp takes one asymmetry seriously: advanced AI plausibly cures diseases, compresses decades of science and lifts people out of poverty, so walking away is itself a moral cost someone pays. The answer is not "don't" but "not yet, and not like this" — build it under the regime we apply to nuclear power and pharmaceuticals, safety demonstrated before deployment, an authority that can say no, pace set by the safety case rather than the product roadmap.
The interesting consequence is technical, not political. If your plan is "go slowly enough to be sure", you need techniques that can make you sure. A behavioural fix — an output filter, a round of fine-tuning that removes the bad completions you found — tells you the failure is no longer where you looked, not that it is gone. Camp 1 therefore weights interpretability and formal verification heavily, because those could in principle produce a claim about what a system is rather than what it did on your test set.
Where it breaks: the plan is only as good as its enforcement, and enforcing a slow-and-safe regime is exactly the coordination problem the other camps call unsolvable. The concrete proposals all try to make coordination cheaper — pooling frontier development into one international lab (Brundage's "CERN for AI"), or building alignment in careful increments where each generation helps verify the next. Raemon notes that even a lab fully intending this finds the hard part organisational: keeping safety discipline under commercial pressure, retaining people who would rather move faster elsewhere, and actually stopping when its own tripwire fires. High-reliability cultures in nuclear plants and hospitals took decades to build; labs are trying to have one on arrival.
Camp 2 — accept the race, buy safety at the margin
This camp grants camp 1's goal and denies its premise. Nobody is going to stop: too many actors, incentives too strong, and a rule binding only those who agree to it hands the frontier to whoever did not. Given that, the highest-value thing a safety-minded person can do is make the systems that will be built as safe as possible for as long as possible.
It splits in two. Pragmatic safety: ship many imperfect, cheap, robust techniques and get them adopted widely, since a method used by everyone beats a better one used by nobody. Win to control: get a decisive capability lead, then spend it — automating alignment research faster than anyone automates capabilities, or denying dangerous capability to others. Aschenbrenner states why the automation step is load-bearing:
"there's no way we'll manage to solve alignment for true superintelligence directly; covering that vast of an intelligence gap seems extremely challenging."— IIIc. Superalignment, Leopold Aschenbrenner (2024)
The move is a bootstrap: align systems only somewhat smarter than us with ordinary empirical methods, then spend millions of copies of them on the alignment research for the generation after. Carlsmith formalises this as two competing feedback loops — capability improving capability, safety improving safety — where the game is keeping the second in control of the first. def/acc generalises it beyond AI: don't slow technology down, differentially accelerate its defensive parts.
Where it breaks: a lead is a wasting asset. Espionage, independent rediscovery, open weights and your own publications all compress it — the research you release to make everyone safer is the research that shortens your lead. There is also a structural problem the camp rarely faces squarely: "let the good actors win" needs an answer to who certifies goodness, and a strategy whose success condition is one organisation holding decisive power over everyone else has a governance failure mode resembling the misuse risk it was meant to prevent. Racing costs something even when it works — it normalises the shortcut, exactly the habit you need absent when the tripwire fires.
Camp 3 — don't build it
This camp denies alignment tractability outright. If a sufficiently capable system cannot be controlled in principle, every technique in camps 1 and 2 is a way of feeling safer while walking toward the same cliff, and the only intervention that changes the outcome is not arriving. The Statement on Superintelligence is the mainstream articulation — note how carefully conditional it is:
"We call for a prohibition on the development of superintelligence, not lifted before there is 1. broad scientific consensus that it will be done safely and controllably, and 2. strong public buy-in."— Statement on Superintelligence, Future of Life Institute (2025)
The engineering objection is that "stop" is not a null action — it is a policy with a specification problem as hard as anything in alignment. Capability comes from compute, data and algorithms, and only one is physically countable. Freeze chip production tomorrow and progress continues through better recipes and better use of installed hardware. A real halt must say exactly what is prohibited (training runs above a threshold? deployment? publication?), define a measurable trigger, and be enforceable in every jurisdiction, indefinitely. Aguirre's "Keep the Future Human" is the most concrete attempt: it hangs everything on compute, the one auditable input, proposing compute accounting for training and inference, hard caps enforced in the silicon, strict liability for systems that are simultaneously highly autonomous, highly general and superhuman, and lighter tiers for systems that are only one or two of those. Steal that decomposition even if you reject the conclusion — the danger is in the conjunction, so forbidding the conjunction while permitting each property alone costs far less than a blanket ban. Positions inside the camp run from that kind of targeted autonomy limit to indefinite moratorium.
Why this decides what you work on
Here is the part that matters for a technical career. Each camp answers "is this technique worth the time" differently, and the split is roughly guarantees versus mitigations. Camps 1 and 3 want work that could license a decision — an interpretability result, a verification argument, an eval strong enough that a lab or regulator would actually halt on it. A mitigation that reduces observed bad behaviour without explaining it is, to them, a way of hiding the evidence you needed. Camp 2 inverts the ranking: cheap, robust and adoptable this quarter beats better-but-late. Both are defensible. What is not is holding a camp implicitly and being surprised your portfolio doesn't match it.
Stages, not teams
The course closes with a caveat worth taking literally: most careful people treat these as a sequence, not a tribe. Pause now to buy time for safety techniques, then proceed under camp-1 conditions. Or build a lead precisely to gain leverage to impose a camp-1 regime on everyone else. Anthropic's core-views post makes the same move: it declines to bet on one world and picks research whose value holds across optimistic, intermediate and pessimistic cases — including the case where the safety work's main output is the evidence that convinces everyone to stop. Apply that as a portfolio test to your own project. Name the scenario in which your work is decisive, then check what it is worth in the other two.
Readings, linked
The chapter itself is a 10-minute read with no separate resources block; these are the eight proposals it links inline, in the order the course names them. If you read one, read Aguirre's executive summary — it is the only piece here that turns a strategic position into a mechanism you could actually audit.
- My recent lecture at Berkeley and a vision for a "CERN for AI" — Miles Brundage (2024) · ~15 min · The canonical camp-1 institutional proposal: pool frontier development internationally, secure the chips and datacenters first, agree a joint scaling plan before scaling. Argues consolidating capability need not mean consolidating power.
- "Carefully Bootstrapped Alignment" is organizationally hard — Raemon (2023) · ~20 min · The best objection to camp 1 from inside camp 1. Even granting the technical plan works, executing it needs a high-reliability safety culture, and those take decades to build and one quarter of commercial pressure to lose.
- IIIc. Superalignment (Situational Awareness) — Leopold Aschenbrenner (2024) · ~35 min · The automate-alignment-research argument in full: RLHF will not scale past human evaluability, so align the somewhat-superhuman generation and spend it on the rest. Note: the course links this through an
arc.netquote-shortener that may not resolve for you; this is the canonical source it points at. - AI for AI Safety — Joe Carlsmith (2025) · ~40 min · Frames the whole strategy space as two feedback loops — capability improving capability, safety improving safety — and asks what it takes for the second to stay in control of the first. Part 5 of his alignment-problem series.
- What is def/acc anyway? — Matt Clifford (2024) · ~8 min · The differential-acceleration position: don't brake technology, steer it, by funding the defensive and agency-preserving parts of the frontier. Assigned as the "race but choose your direction" flavour of camp 2.
- Keep the Future Human — executive summary — Anthony Aguirre (2025) · ~20 min · The most mechanism-level camp-3 proposal: compute accounting, hardware-enforced caps on training and inference, strict liability for the autonomy × generality × superhuman-intelligence conjunction, tiered rules below it.
- Statement on Superintelligence — Future of Life Institute (2025) · ~2 min · A single sentence and a signatory list spanning Bengio, Hinton, five Nobel laureates and a politically improbable spread of public figures. Read it for the precise conditionality of the ask, not the length.
- If Anyone Builds It, Everyone Dies — Eliezer Yudkowsky & Nate Soares (2025) · book · The strongest form of the intractability claim, and the reason camp 3 exists at all: if superhuman systems cannot be controlled even in principle, no amount of careful building changes the outcome.
Exercises
This chapter ships no exercises of its own — it is a framing reading between chapter 1's problem statement and chapter 3's difficulty argument. Two extras below, built to make the framing operational.
- Position statement with a falsifier (field map extra) — In under 200 words, state which camp you are currently in and why, then do the harder half: name the single concrete observation that would move you to each of the other two camps. Then go find the strongest published argument against your camp (Raemon against camp 1, the lead-durability critique against camp 2, the specification-problem critique against camp 3) and write two sentences on why it does not move you yet. What a good answer has: a falsifier that is an observation rather than an argument — "an international compute-accounting regime signed by the US and China holds for 18 months", "a frontier lab's own eval triggers and it actually halts a run", "an interpretability result predicts a behaviour nobody had observed" — plus honesty about which of the two underlying variables (coordination feasibility, alignment tractability) you are actually uncertain about. A common failure is a falsifier so demanding that nothing could ever satisfy it.
- Camp-tag the safety literature code (field map extra) — Build a small classifier that reads a safety agenda or lab policy document and predicts which camp it implies, then measure whether an LLM's labels agree with yours. This turns the chapter's taxonomy from a vibe into something with an error rate, and the disagreements are where the real learning is. What a good answer has: at least 20 documents spanning all three camps and some deliberately ambiguous ones (Anthropic's RSP, OpenAI's Preparedness Framework, a METR eval paper, an interpretability agenda, the FLI statement); your own labels recorded before you see the model's; a reported agreement statistic rather than an eyeballed impression; and a written analysis of the three or four documents where you and the model disagreed most, since those are usually documents that genuinely straddle camps rather than model errors. Start here: (1) collect the docs as plain text —
trafilaturaorreadability-lxmlfor the HTML, one file per doc, and keep alabels.csvof your own judgements with a confidence column; (2) write a prompt that gives the model the three camp definitions from this chapter and asks for a JSON{camp, confidence, evidence_quote}— demanding a supporting quote is what makes the errors diagnosable; (3) run it with theanthropicSDK (a cheap fast model is fine — this is classification, not reasoning-heavy) or, entirely locally,transformerswithQwen2.5-7B-Instructin 4-bit on a free Colab T4; (4) score agreement withsklearn.metrics.cohen_kappa_score— below about 0.4 usually means your camp definitions are underspecified, not that the model is bad; (5) sharpen the prompt with the two disambiguating questions from this page's grid ("does this assume coordination is feasible?", "does this assume alignment is tractable?") and re-run to see whether kappa moves; (6) inspect the residual disagreements by hand and write up which documents refuse to fit, and why.
Go deeper
- The Checklist: What Succeeding at AI Safety Will Involve — Sam Bowman (2024). The single most direct answer to this chapter's title question from inside a frontier lab: what has to be true, phase by phase, for this to have gone well. Notably argues you do not need perfect alignment if safeguards and monitoring are strong enough — a camp-2 position argued carefully.
- Core Views on AI Safety — Anthropic (2023). The portfolio framing: optimistic, intermediate and pessimistic worlds, and research chosen to be worth doing in all three — including the pessimistic case where the output is the evidence that justifies stopping.
- My Techno-Optimism — Vitalik Buterin (2023). The original d/acc essay that def/acc descends from, and clearer on the mechanism: prefer technologies whose defence-offence ratio favours defence, across bio, cyber, info and governance, not just AI.
- FLI press release on the superintelligence statement — Future of Life Institute (2025). Useful context for the statement itself: who signed, the accompanying polling on public appetite, and how it differs from the 2023 six-month-pause letter (conditional prohibition rather than a timed pause).