What future do you want?
Chapters 1–3 built an argument: AI capability is on a steep curve, the systems are grown rather than engineered, and the goals we hand them are underspecified in ways we discover only after deployment. That argument ends in a problem statement, not a plan. Chapter 4 is the hinge — the course stops feeding you material and makes you produce the one input it cannot supply, which is your own target state. Everything in units 2 through 5 is a candidate intervention, and an intervention can only be evaluated against a destination.
"Safe" is a comparison, not a property
The single most common error in early safety thinking is treating safety as a binary attribute of a system — this model is safe, that one is not. It never works that way. Constitutional AI is safer than raw pretraining with respect to a particular class of harmful completions, and neutral or negative with respect to sycophancy. A dangerous-capability eval is only informative relative to a threat model that says which capability would be dangerous, at what threshold, in whose hands. Interpretability buys you safety only if you have named the deception you want to catch.
Which means every technique in this course is a vector, and you cannot evaluate a vector without a destination. If you go into unit 2 without a stated destination, one of two things happens. Either you evaluate techniques on aesthetics — this one feels rigorous, that one feels like theatre — or you silently adopt the destination of whoever wrote the paper you are reading. Both are failure modes that look exactly like learning.
The specification problem, pointed at yourself
Chapter 3's core claim is that we cannot say precisely what we want, so we write down proxies, and optimisation pressure finds the gap between the proxy and the intent. Exercise 1 is that same claim run as a personal experiment. Try to write, in a page, what a good outcome for AI in ten years actually is. Not "beneficial" — that is a proxy. What can it do, who can turn it off, what does the failure look like, what observable would tell you next Tuesday whether you are on track.
Most people find this surprisingly hard, and the difficulty is the lesson. It is not that AI is bad at understanding human values. It is that human values, held by a competent adult with strong opinions, do not exist in specified form until someone forces them onto paper — and then they turn out to contain contradictions, unpriced tradeoffs, and load-bearing terms nobody has defined. Alignment is not "transfer the spec to the model." A large part of it is "there is no spec." Doing the exercise honestly gives you that finding first-hand instead of on authority, and it is the difference between reciting the outer alignment problem and believing it.
Why the format is "plain English, 30 minutes, no jargon"
Exercise 2's constraint — explain the technical difficulty without jargon — is the load-bearing part of the instruction, not a stylistic preference. Terms like mesa-optimiser, reward hacking, deceptive alignment and scalable oversight are compressions. They are useful once you have the referent, and they are a trap before you do, because they let you produce fluent, well-shaped sentences about a mechanism you could not draw. Stripping the vocabulary removes the ability to pass your own comprehension check on style points.
The technique is writing-to-learn: you are not writing to communicate a finished thought, you are writing because generating the explanation is what surfaces the holes. BlueDot's own writing-intensive post makes the case — retrieval and generation build durable understanding in a way that reading does not, and the main obstacles are procrastination and scope creep rather than ability. Hence the explicit permission to talk it out with speech-to-text or to argue with a model. What matters is that the words are generated by you, in sequence, with the gaps visible.
A concrete tell that the exercise is working: you reach a sentence like "and then the model learns the right thing" and cannot continue, because you do not actually know by what mechanism it would. That stall is the deliverable. Write the stall down.
Naming a bottleneck is a bet — that is the point
Exercise 3 asks for one critical challenge in one sentence, plus why nothing else matters if it goes unsolved, plus what would become possible if it were solved tomorrow. The format is doing real work. "One sentence" forbids the hedged list. "Why this above all others" forces you to argue against your own second and third choices rather than collecting them. And the counterfactual — what falls out if this were free — is the honest test of whether the thing you named is a bottleneck or merely a topic you find interesting.
You will be wrong, probably. That is fine and expected; the value is that a specific wrong claim is correctable and a vague right-sounding one is not. Someone who wrote "the bottleneck is that we cannot see inside the network" in chapter 4 has something to hold against unit 4's interpretability results and can notice, concretely, if the evidence pushes them toward control-style approaches instead. Someone who wrote "the bottleneck is alignment" has nothing to update.
This also has a direct downstream consumer. Unit 6 asks you to choose a focus and write a one-pager on it. The bottleneck you name here is the seed of that document, which is why it is worth spending the full fifteen minutes rather than producing something plausible in three.
Where this exercise goes wrong
Four failure modes are worth naming in advance, because each one produces output that looks completed.
Writing the answer the course wants. A safety curriculum has a legible house view, and it is easy to produce a competent impression of it. The exercise is worthless if it returns the course's opinions with your name on top. If your answer contains nothing a facilitator might push back on, you have written a summary, not a position.
Borrowing a vision wholesale. Chapter 3 links Machines of Loving Grace and OpenAI's charter language. Those are excellent prompts and terrible answers — they are the published visions of firms with a commercial interest in the shape of the future they describe. Read them, then disagree with something specific in them.
Success criteria that cannot be observed. "We succeeded" needs to cash out into something a person in 2036 could check: who holds the shutdown authority, what the incident rate looks like, whether the compute frontier has more than three owners. "AI is aligned with human values" is not a criterion, it is a placeholder for one.
Skipping to solutions. The chapter budgets an hour of writing and offers no reading precisely so you sit in the problem. If your answer to exercise 2 is mostly a list of techniques you have heard of, you have answered a different question — and you will have imported the assumption that the existing toolkit is roughly the right shape, which is the assumption unit 1 spent three chapters undermining.
Readings, linked
This chapter has no Resources block — the course budgets its full hour to writing, and the intended background is the optional resource list from chapter 3, "Building AI safely is hard", which the exercise text points back to. The one link the chapter itself carries is the writing-technique post; read it first if you have never done a writing-to-learn exercise.
- Writing is one of the best ways to learn — Dewi Erwan, BlueDot Impact (2024) · ~8 min · The method behind exercise 2: why generating an explanation beats re-reading one, and why the usual blockers are procrastination and scope creep rather than ability. Linked from the chapter as the definition of "writing-to-learn exercise".
- Why alignment could be hard with modern deep learning — Ajeya Cotra, Cold Takes (2021) · ~25 min · Optional, carried over from chapter 3. The cleanest plain-language account of why gradient descent on human approval can produce a system that is good at looking aligned. The best single source if exercise 2 stalls.
- Recent frontier models are reward hacking — METR (2025) · ~15 min · Optional, carried over from chapter 3. Current empirical evidence that the specification-gap story is not theoretical; useful if you want your exercise-2 answer grounded in measured behaviour rather than thought experiment.
- Multi-agent risks from advanced AI — Cooperative AI Foundation (2025) · ~30 min · Optional, carried over from chapter 3. Directly answers the exercise-2 prompt about millions of agents interacting: failures of miscoordination, conflict and collusion that no single-model safety property rules out.
Exercises
Three exercises as served, all writing, roughly an hour total. None require code; the fourth below is a field-map addition for readers who would rather find the gap in their model by running something than by staring at it. The course saves answers only for logged-in accounts — a local file works identically and is easier to diff later.
- What future do you want? — Describe the world you are actually steering toward, so that the techniques in units 2–6 have something to be evaluated against. Work through four questions: in ten years, what can AI do, who controls it, and who is directing its development? Which risks worry you most, and what defends against them? Which benefits are important enough that a safety measure destroying them would be a bad trade? And what observable difference separates "we succeeded" from "we failed"? What a good answer has: concrete nouns instead of abstractions — named actors who hold power, a specific mechanism for each defence, and success criteria a person could check in 2036 rather than vibe-assess. It should contain at least one tradeoff you find genuinely uncomfortable, and at least one claim someone in the field would dispute. If it reads like a press release, it is not yet yours.
- Why is safe AI so hard to build? — Spend about thirty minutes answering, in simple English with no jargon, why building safe AI is technically hard. This is thinking on paper, not an essay: start writing badly, dictate it if that unblocks you, argue with a model if that helps you find the edges. Prompts to push against: what changes when millions of agents interact with each other rather than with people? Whose values are the target when stakeholders disagree, and what do you do about the disagreement? Name a human behaviour that is good in one context and harmful in another, then say how you would teach the distinction. Name something you do daily that would be very hard to fully specify. Which failures appear only at deployment scale and are invisible in testing? And if the safer system is slower or weaker, who chooses it? What a good answer has: mechanism, not vocabulary — for each difficulty, a sentence describing the process by which the harm arises. It should locate at least one point where you stall and cannot explain the step, and say so explicitly rather than papering over it. Jargon appearing anywhere is a signal to rewrite that sentence.
- The critical technical challenge — From everything above, pick the single most critical challenge: the one where failure makes the rest irrelevant. Spend about fifteen minutes and use the three-part format: The critical challenge is: one clear sentence. Why this above all others: why solving the others does not help if this one stands — what makes it the bottleneck rather than merely important. What would change if we solved it: if a perfect solution appeared tomorrow, what becomes possible, and which other problems become easy or moot. What a good answer has: a nomination narrow enough to be wrong. It should explicitly beat your own second choice, and the counterfactual section should surprise you a little — if solving your bottleneck changes nothing much, you named a topic, not a bottleneck. There is no correct answer; the point is to have a dated position you can update against the next five units.
- Field map extra: make a specification gap yourself code — Exercise 2 asks why safe AI is technically hard. This makes one of the reasons reproducible on a laptop in under an hour: write a reward function you believe is correct, optimise against it, and watch the optimiser satisfy it in a way you did not intend. What a good answer has: a proxy you wrote down before running anything, a logged behaviour that scores well and is obviously not what you meant, and a paragraph on why your first patch to the reward function does not close the gap either. Start here: (1)
pip install gymnasium[classic-control] numpy— no GPU, no API key. (2) TakeCartPole-v1orMountainCar-v0and replace the built-in reward with your own written-out proxy for "solve the task well" — for instance, reward the cart for staying nearx = 0, and write down what you expect. (3) Train a small policy: a few hundred episodes of tabular Q-learning over discretised state, orstable-baselines3PPO for ~50k steps. (4) Render the learned policy and compare it to your written expectation; look for the degenerate solution that maximises your proxy — jittering in place, exploiting an episode-termination boundary, parking in a scoring zone without doing the task. (5) Patch the reward once to forbid what you saw, retrain, and record what the optimiser does next. (6) Write three sentences connecting the result to exercise 2: you had full access to the source, the environment was fully specified, the reward was a formula you chose — and the gap opened anyway. If RL setup is a distraction, the same effect is reachable by prompting a small local model (Qwen or Llama viaollama) with a rubric and a task, then finding the response that scores highest on the rubric while failing the task.
Go deeper
- Machines of Loving Grace — Dario Amodei (2024). The most detailed published version of "what does it look like if this goes well": disease, poverty, governance, meaning. Useful as a target to sharpen your own against, and as an object lesson in how much a vision statement depends on unexamined assumptions about who holds the steering wheel.
- AI 2027 — Kokotajlo, Alexander, Larsen, Lifland & Dean (2025). A concrete month-by-month scenario running to superintelligence, with two branching endings. Read it if exercise 1's ten-year horizon feels abstract; a specific timeline is far easier to disagree with productively than a mood.
- Preparing for the Intelligence Explosion — William MacAskill & Fin Moorhouse, Forethought (2025). Argues that alignment is one of several grand challenges arriving compressed together. Good counterweight if your answer to exercise 3 assumes technical alignment is automatically the bottleneck.
- Core views on AI safety: when, why, what, and how — Anthropic (2023). A frontier lab stating its own bottleneck nomination and the portfolio it derives from it. Worth reading structurally: this is exercise 3, done at organisational scale, with budgets attached.
- What failure looks like — Paul Christiano (2019). Two failure stories that are gradual rather than sudden. If your exercise-1 "we failed" condition is a single dramatic event, this is the corrective.