TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 1 · THE TECHNICAL CHALLENGE WITH AIchapter 4 · 1h · exercises

What future do you want?

BlueDot Impact · Technical AI Safety · unit 1, chapter 4
TL;DR — Unit 1 ends with no readings and three writing prompts, and that is deliberate. The rest of the course is a parade of techniques — data curation, RLHF, evals, interpretability, control — and every one of them is only ever safer than what, measured against what target. So before you meet the techniques you write down the target: the world you are steering toward, why steering is technically hard in plain English, and which single failure would make all the other work moot. The artifact you produce here is a dated, falsifiable statement of your own position. Keep it; you will diff against it in unit 6.

Chapters 1–3 built an argument: AI capability is on a steep curve, the systems are grown rather than engineered, and the goals we hand them are underspecified in ways we discover only after deployment. That argument ends in a problem statement, not a plan. Chapter 4 is the hinge — the course stops feeding you material and makes you produce the one input it cannot supply, which is your own target state. Everything in units 2 through 5 is a candidate intervention, and an intervention can only be evaluated against a destination.

"Safe" is a comparison, not a property

The single most common error in early safety thinking is treating safety as a binary attribute of a system — this model is safe, that one is not. It never works that way. Constitutional AI is safer than raw pretraining with respect to a particular class of harmful completions, and neutral or negative with respect to sycophancy. A dangerous-capability eval is only informative relative to a threat model that says which capability would be dangerous, at what threshold, in whose hands. Interpretability buys you safety only if you have named the deception you want to catch.

Which means every technique in this course is a vector, and you cannot evaluate a vector without a destination. If you go into unit 2 without a stated destination, one of two things happens. Either you evaluate techniques on aesthetics — this one feels rigorous, that one feels like theatre — or you silently adopt the destination of whoever wrote the paper you are reading. Both are failure modes that look exactly like learning.

The specification problem, pointed at yourself

Chapter 3's core claim is that we cannot say precisely what we want, so we write down proxies, and optimisation pressure finds the gap between the proxy and the intent. Exercise 1 is that same claim run as a personal experiment. Try to write, in a page, what a good outcome for AI in ten years actually is. Not "beneficial" — that is a proxy. What can it do, who can turn it off, what does the failure look like, what observable would tell you next Tuesday whether you are on track.

Most people find this surprisingly hard, and the difficulty is the lesson. It is not that AI is bad at understanding human values. It is that human values, held by a competent adult with strong opinions, do not exist in specified form until someone forces them onto paper — and then they turn out to contain contradictions, unpriced tradeoffs, and load-bearing terms nobody has defined. Alignment is not "transfer the spec to the model." A large part of it is "there is no spec." Doing the exercise honestly gives you that finding first-hand instead of on authority, and it is the difference between reciting the outer alignment problem and believing it.

The practitioner's takeaway: write the target down, date it, and store it somewhere you will find it again. A safety agenda you cannot state in a paragraph is not an agenda, it is a mood — and an undated one cannot be checked for drift.

Why the format is "plain English, 30 minutes, no jargon"

Exercise 2's constraint — explain the technical difficulty without jargon — is the load-bearing part of the instruction, not a stylistic preference. Terms like mesa-optimiser, reward hacking, deceptive alignment and scalable oversight are compressions. They are useful once you have the referent, and they are a trap before you do, because they let you produce fluent, well-shaped sentences about a mechanism you could not draw. Stripping the vocabulary removes the ability to pass your own comprehension check on style points.

The technique is writing-to-learn: you are not writing to communicate a finished thought, you are writing because generating the explanation is what surfaces the holes. BlueDot's own writing-intensive post makes the case — retrieval and generation build durable understanding in a way that reading does not, and the main obstacles are procrastination and scope creep rather than ability. Hence the explicit permission to talk it out with speech-to-text or to argue with a model. What matters is that the words are generated by you, in sequence, with the gaps visible.

A concrete tell that the exercise is working: you reach a sentence like "and then the model learns the right thing" and cannot continue, because you do not actually know by what mechanism it would. That stall is the deliverable. Write the stall down.

Naming a bottleneck is a bet — that is the point

Exercise 3 asks for one critical challenge in one sentence, plus why nothing else matters if it goes unsolved, plus what would become possible if it were solved tomorrow. The format is doing real work. "One sentence" forbids the hedged list. "Why this above all others" forces you to argue against your own second and third choices rather than collecting them. And the counterfactual — what falls out if this were free — is the honest test of whether the thing you named is a bottleneck or merely a topic you find interesting.

You will be wrong, probably. That is fine and expected; the value is that a specific wrong claim is correctable and a vague right-sounding one is not. Someone who wrote "the bottleneck is that we cannot see inside the network" in chapter 4 has something to hold against unit 4's interpretability results and can notice, concretely, if the evidence pushes them toward control-style approaches instead. Someone who wrote "the bottleneck is alignment" has nothing to update.

This also has a direct downstream consumer. Unit 6 asks you to choose a focus and write a one-pager on it. The bottleneck you name here is the seed of that document, which is why it is worth spending the full fifteen minutes rather than producing something plausible in three.

Where this exercise goes wrong

Four failure modes are worth naming in advance, because each one produces output that looks completed.

Writing the answer the course wants. A safety curriculum has a legible house view, and it is easy to produce a competent impression of it. The exercise is worthless if it returns the course's opinions with your name on top. If your answer contains nothing a facilitator might push back on, you have written a summary, not a position.

Borrowing a vision wholesale. Chapter 3 links Machines of Loving Grace and OpenAI's charter language. Those are excellent prompts and terrible answers — they are the published visions of firms with a commercial interest in the shape of the future they describe. Read them, then disagree with something specific in them.

Success criteria that cannot be observed. "We succeeded" needs to cash out into something a person in 2036 could check: who holds the shutdown authority, what the incident rate looks like, whether the compute frontier has more than three owners. "AI is aligned with human values" is not a criterion, it is a placeholder for one.

Skipping to solutions. The chapter budgets an hour of writing and offers no reading precisely so you sit in the problem. If your answer to exercise 2 is mostly a list of techniques you have heard of, you have answered a different question — and you will have imported the assumption that the existing toolkit is roughly the right shape, which is the assumption unit 1 spent three chapters undermining.

Readings, linked

This chapter has no Resources block — the course budgets its full hour to writing, and the intended background is the optional resource list from chapter 3, "Building AI safely is hard", which the exercise text points back to. The one link the chapter itself carries is the writing-technique post; read it first if you have never done a writing-to-learn exercise.

Exercises

Three exercises as served, all writing, roughly an hour total. None require code; the fourth below is a field-map addition for readers who would rather find the gap in their model by running something than by staring at it. The course saves answers only for logged-in accounts — a local file works identically and is easier to diff later.

  1. What future do you want? — Describe the world you are actually steering toward, so that the techniques in units 2–6 have something to be evaluated against. Work through four questions: in ten years, what can AI do, who controls it, and who is directing its development? Which risks worry you most, and what defends against them? Which benefits are important enough that a safety measure destroying them would be a bad trade? And what observable difference separates "we succeeded" from "we failed"? What a good answer has: concrete nouns instead of abstractions — named actors who hold power, a specific mechanism for each defence, and success criteria a person could check in 2036 rather than vibe-assess. It should contain at least one tradeoff you find genuinely uncomfortable, and at least one claim someone in the field would dispute. If it reads like a press release, it is not yet yours.
  2. Why is safe AI so hard to build? — Spend about thirty minutes answering, in simple English with no jargon, why building safe AI is technically hard. This is thinking on paper, not an essay: start writing badly, dictate it if that unblocks you, argue with a model if that helps you find the edges. Prompts to push against: what changes when millions of agents interact with each other rather than with people? Whose values are the target when stakeholders disagree, and what do you do about the disagreement? Name a human behaviour that is good in one context and harmful in another, then say how you would teach the distinction. Name something you do daily that would be very hard to fully specify. Which failures appear only at deployment scale and are invisible in testing? And if the safer system is slower or weaker, who chooses it? What a good answer has: mechanism, not vocabulary — for each difficulty, a sentence describing the process by which the harm arises. It should locate at least one point where you stall and cannot explain the step, and say so explicitly rather than papering over it. Jargon appearing anywhere is a signal to rewrite that sentence.
  3. The critical technical challenge — From everything above, pick the single most critical challenge: the one where failure makes the rest irrelevant. Spend about fifteen minutes and use the three-part format: The critical challenge is: one clear sentence. Why this above all others: why solving the others does not help if this one stands — what makes it the bottleneck rather than merely important. What would change if we solved it: if a perfect solution appeared tomorrow, what becomes possible, and which other problems become easy or moot. What a good answer has: a nomination narrow enough to be wrong. It should explicitly beat your own second choice, and the counterfactual section should surprise you a little — if solving your bottleneck changes nothing much, you named a topic, not a bottleneck. There is no correct answer; the point is to have a dated position you can update against the next five units.
  4. Field map extra: make a specification gap yourself code — Exercise 2 asks why safe AI is technically hard. This makes one of the reasons reproducible on a laptop in under an hour: write a reward function you believe is correct, optimise against it, and watch the optimiser satisfy it in a way you did not intend. What a good answer has: a proxy you wrote down before running anything, a logged behaviour that scores well and is obviously not what you meant, and a paragraph on why your first patch to the reward function does not close the gap either. Start here: (1) pip install gymnasium[classic-control] numpy — no GPU, no API key. (2) Take CartPole-v1 or MountainCar-v0 and replace the built-in reward with your own written-out proxy for "solve the task well" — for instance, reward the cart for staying near x = 0, and write down what you expect. (3) Train a small policy: a few hundred episodes of tabular Q-learning over discretised state, or stable-baselines3 PPO for ~50k steps. (4) Render the learned policy and compare it to your written expectation; look for the degenerate solution that maximises your proxy — jittering in place, exploiting an episode-termination boundary, parking in a scoring zone without doing the task. (5) Patch the reward once to forbid what you saw, retrain, and record what the optimiser does next. (6) Write three sentences connecting the result to exercise 2: you had full access to the source, the environment was fully specified, the reward was a formula you chose — and the gap opened anyway. If RL setup is a distraction, the same effect is reachable by prompting a small local model (Qwen or Llama via ollama) with a rubric and a task, then finding the response that scores highest on the rubric while failing the task.

Go deeper

Next: Can we train AI to be safe? · Back to the map.