Making AI go well
Strategy documents about advanced AI share a soft spot. They describe a target state — "stays under human control", "refuses to help with a bioweapon", "we can tell whether it is scheming" — and move on, as if that were a configuration setting rather than an open research problem. This chapter announces the course is going to look under that assumption. It teaches nothing technical yet; it sets the terms of a four-unit argument about what today's techniques can and cannot deliver.
Two questions that sound like one
"How do we make AI go well?" is what the AGI strategy course answers, in the currency of actors and incentives: who holds the compute, which threat pathways are live, which plan you are betting on. That yields claims like "build defences and diffuse AI" or "hand control to an aligned superintelligence." Both are coherent. Neither tells you how to build a defence or verify an alignment.
The technical question is narrower and far less forgiving: given a transformer trained on internet-scale data by gradient descent and then shaped by feedback, which safety properties can you actually cause to hold, how confidently, and how does that confidence decay as capability rises? That is what separates a plan from a proposal — and where most of the bad news lives, which is why units 2–5 are about techniques and their gaps rather than ambitions.
Failures at the technical layer are silent at the strategy layer
A regime requiring labs to test for dangerous capabilities is only as strong as the evals: if an eval measures "can the model do X when asked plainly" and the threat is "will it do X when it has reason to hide", the regime reports green while risk is unchanged. A plan resting on human oversight is only as strong as the oversight — which gets harder exactly as the overseen system outgrows the overseer. Strategy inherits every weakness of the mechanisms it delegates to, and cannot detect them in its own vocabulary.
"Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings."— Unsolved Problems in ML Safety, Hendrycks, Carlini, Schulman & Steinhardt (2021)
Five years old and now an understatement — the same argument in a line: capability, novelty and stakes rise together, each eroding a different assumption safety rests on.
The spine: defence in depth
The structure carried over from strategy is layered — stop the behaviour being learned, constrain the capability if it is learned, survive the action if it is taken anyway. The course walks those layers asking what tooling exists at each: training-time intervention (unit 2), evaluation (unit 3), reading the internals (unit 4), containment (unit 5, which assumes the earlier layers leaked).
Read that as a chain, not a menu. No layer is sufficient alone; the real question is whether the failure modes are independent. If a capable model defeats your training, your evals and your monitor for the same underlying reason — all three graded by systems weaker than it is — stacking them buys much less than the diagram suggests. Unit 5's kill-chain exercise makes you confront that on your own example.
The four exclusions
The chapter is explicit about what it will not cover. Each cut is defensible; three leak:
- Policy detail — deferred to a separate course, on the theory that the substrate comes first. But a threshold in a frontier safety framework is a technical claim in legal clothes.
- Compute governance — the cleanest cut: a supply-chain and mechanism-design problem, not a model-behaviour one.
- AI security — weight theft and self-exfiltration. Load-bearing: every deployment-time mitigation in unit 5 assumes the weights stayed where you put them.
- ML fundamentals — assumed, not taught. If backprop, tokenisation, pretraining versus post-training and RL "policy" are not comfortable, do BlueDot's AI foundations modules first; unit 4 is unreadable without them.
Readings, linked
This chapter has no Resources block and no assigned reading time — it is a two-screen orientation page plus a short video. What it does carry is a set of back-links into the strategy course and one prerequisite; start with the video, and only detour into the strategy chapters if the "defence in depth" framing is new to you.
- Technical AI Safety Unit 1: The technical challenge with AI — BlueDot Impact (2025) · 1 min 33 · The embedded intro video. Sets the unit's framing in less time than reading about it takes; verified as public on YouTube, no login needed.
- AGI Strategy (course) — BlueDot Impact · ~25h · The prior course this chapter hands off from. Free and pay-what-you-want; the chapter assumes its conclusions, not its detail.
- Drivers of AI progress — BlueDot, AGI Strategy unit 2 ch 1 · 30 min · Compute, data and algorithms as the three inputs. Assigned so that "the models keep getting better" is a mechanism you can reason about rather than a mood.
- Pathways to harm — BlueDot, AGI Strategy unit 3 ch 1 · ~55 min · The threat-model catalogue: power concentration, gradual disempowerment, pandemics, infrastructure collapse. Supplies the "harm" that safety techniques are measured against.
- AGI Strategy unit 4 ch 1 — “What might success look like?” — BlueDot · 5 min · Linked from this chapter as "plans for making AI go well"; the live page now opens unit 4 (Defence in Depth) with a success-picture chapter of the same name as TAS 1.2. Read it as the plans-level companion to the next chapter here.
- Building defences — BlueDot, AGI Strategy unit 4 ch 2 · The layered-defence chapter this course is structurally built on: prevent → constrain → withstand. The single most useful back-link on the page.
- AI Foundations — BlueDot Impact (Notion) · self-paced · The stated prerequisite for anyone without working ML background. Public Notion page, no login; it renders only with JavaScript on.
Exercises
This chapter ships no exercises — it is pure orientation, and the course's first real work lands in chapter 1.4. Two field-map extras below, because an orientation chapter is exactly where a cheap pre-commitment pays off later.
- Write your prior before the course argues with it (field map extra) — In under 300 words, answer three questions in writing and date the file: (a) which of the strategy course's plans do you currently believe is most likely to actually work — government control, aligned superintelligence, or defend-and-diffuse; (b) which technical layer you expect to be strongest today — training, evals, interpretability, or containment; (c) what is the largest capability jump you believe current techniques would survive without a qualitative change in approach. What a good answer has: a specific commitment on each, each with a falsifier — the observation that would move you. Re-open the file after unit 5 and diff it; the value is entirely in the diff, and a prior you cannot be wrong about scores zero.
- Sanity-check the prerequisite (field map extra) code — Before unit 2, confirm the ML baseline the course assumes rather than discovering the gap in unit 4. Load a small open model and reproduce four things end to end: tokenisation, a forward pass, next-token logits, and a single gradient step. What a good answer has: a notebook where you can point at the tensor that is a logit distribution and say what its shape means, plus one sentence distinguishing pretraining from post-training in terms of what the loss is computed against. Start here: (1) free Colab or any laptop with 8 GB RAM; pip install transformers torch. (2) Use a genuinely small model — gpt2 or Qwen/Qwen2.5-0.5B-Instruct; nothing here needs a frontier model. (3) Tokenise a sentence, print the token ids and the decoded pieces, and find a word that splits into three tokens. (4) Run the model, take outputs.logits[0, -1], softmax it, and print the top-10 next tokens with probabilities. (5) Compute the loss against the true next token and call loss.backward(); print the norm of the gradient on one weight matrix. (6) Write the one-sentence pretraining-vs-post-training answer at the top of the notebook. If any step needed a tutorial, do the AI Foundations modules before unit 2.
Go deeper
- An Approach to Technical AGI Safety and Security — Shah, Irpan, Turner, Wang, Conmy, Lindner et al., Google DeepMind (2025). The closest thing to this course's argument written as a single research agenda: misuse and misalignment split apart, then model-level and system-level mitigations stacked. Read the introduction now, the rest after unit 5.
- Unsolved Problems in ML Safety — Hendrycks, Carlini, Schulman & Steinhardt (2021). The four-bucket taxonomy — robustness, monitoring, alignment, systemic safety — that most later field maps, including this course's unit structure, are variations on.
- Concrete Problems in AI Safety — Amodei, Olah, Steinhardt, Christiano, Schulman & Mané (2016). The paper that made "specification, robustness, oversight" concrete enough to fund. Dated in its examples and startlingly current in its failure modes; useful for seeing which of the 2016 problems are still open.
- Core Views on AI Safety — Anthropic (2023). A lab stating in public which safety bets it is making and under which assumptions about difficulty. Worth reading now as the genre unit 3 will teach you to audit rather than take at face value.
- BlueDot Impact — all courses. Where the excluded scope actually lives: policy, compute governance and the foundations modules are separate offerings, all free.