TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 1 · THE TECHNICAL CHALLENGE WITH AIchapter 1 · orientation, ~5 min

Making AI go well

BlueDot Impact · Technical AI Safety · unit 1, chapter 1
TL;DR — The hinge between BlueDot's strategy course and its technical one. Strategy asks which futures we want and which plans could get there; this course asks what those plans quietly assume — what are these systems made of, and which safety properties can we actually engineer into them? It fixes the scope (training, evaluation, interpretability, containment), names four exclusions, and previews the output: a kill chain and a one-page fundable plan. Remember: every safety plan bottoms out in a technique, and techniques have failure modes plans do not mention.

Strategy documents about advanced AI share a soft spot. They describe a target state — "stays under human control", "refuses to help with a bioweapon", "we can tell whether it is scheming" — and move on, as if that were a configuration setting rather than an open research problem. This chapter announces the course is going to look under that assumption. It teaches nothing technical yet; it sets the terms of a four-unit argument about what today's techniques can and cannot deliver.

Two questions that sound like one

"How do we make AI go well?" is what the AGI strategy course answers, in the currency of actors and incentives: who holds the compute, which threat pathways are live, which plan you are betting on. That yields claims like "build defences and diffuse AI" or "hand control to an aligned superintelligence." Both are coherent. Neither tells you how to build a defence or verify an alignment.

The technical question is narrower and far less forgiving: given a transformer trained on internet-scale data by gradient descent and then shaped by feedback, which safety properties can you actually cause to hold, how confidently, and how does that confidence decay as capability rises? That is what separates a plan from a proposal — and where most of the bad news lives, which is why units 2–5 are about techniques and their gaps rather than ambitions.

Failures at the technical layer are silent at the strategy layer

A regime requiring labs to test for dangerous capabilities is only as strong as the evals: if an eval measures "can the model do X when asked plainly" and the threat is "will it do X when it has reason to hide", the regime reports green while risk is unchanged. A plan resting on human oversight is only as strong as the oversight — which gets harder exactly as the overseen system outgrows the overseer. Strategy inherits every weakness of the mechanisms it delegates to, and cannot detect them in its own vocabulary.

"Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings."— Unsolved Problems in ML Safety, Hendrycks, Carlini, Schulman & Steinhardt (2021)

Five years old and now an understatement — the same argument in a line: capability, novelty and stakes rise together, each eroding a different assumption safety rests on.

The spine: defence in depth

The structure carried over from strategy is layered — stop the behaviour being learned, constrain the capability if it is learned, survive the action if it is taken anyway. The course walks those layers asking what tooling exists at each: training-time intervention (unit 2), evaluation (unit 3), reading the internals (unit 4), containment (unit 5, which assumes the earlier layers leaked).

Read that as a chain, not a menu. No layer is sufficient alone; the real question is whether the failure modes are independent. If a capable model defeats your training, your evals and your monitor for the same underlying reason — all three graded by systems weaker than it is — stacking them buys much less than the diagram suggests. Unit 5's kill-chain exercise makes you confront that on your own example.

The four exclusions

The chapter is explicit about what it will not cover. Each cut is defensible; three leak:

What to carry out of an orientation chapter: for every technique the course introduces, write the one sentence saying what would have to be true about the model for it to stop working. It is almost always "the model is capable enough to notice it is being trained, tested or watched." The techniques differ; the sentence recurs — noticing that is most of what unit 1 installs.

Readings, linked

This chapter has no Resources block and no assigned reading time — it is a two-screen orientation page plus a short video. What it does carry is a set of back-links into the strategy course and one prerequisite; start with the video, and only detour into the strategy chapters if the "defence in depth" framing is new to you.

Exercises

This chapter ships no exercises — it is pure orientation, and the course's first real work lands in chapter 1.4. Two field-map extras below, because an orientation chapter is exactly where a cheap pre-commitment pays off later.

  1. Write your prior before the course argues with it (field map extra) — In under 300 words, answer three questions in writing and date the file: (a) which of the strategy course's plans do you currently believe is most likely to actually work — government control, aligned superintelligence, or defend-and-diffuse; (b) which technical layer you expect to be strongest today — training, evals, interpretability, or containment; (c) what is the largest capability jump you believe current techniques would survive without a qualitative change in approach. What a good answer has: a specific commitment on each, each with a falsifier — the observation that would move you. Re-open the file after unit 5 and diff it; the value is entirely in the diff, and a prior you cannot be wrong about scores zero.
  2. Sanity-check the prerequisite (field map extra) code — Before unit 2, confirm the ML baseline the course assumes rather than discovering the gap in unit 4. Load a small open model and reproduce four things end to end: tokenisation, a forward pass, next-token logits, and a single gradient step. What a good answer has: a notebook where you can point at the tensor that is a logit distribution and say what its shape means, plus one sentence distinguishing pretraining from post-training in terms of what the loss is computed against. Start here: (1) free Colab or any laptop with 8 GB RAM; pip install transformers torch. (2) Use a genuinely small model — gpt2 or Qwen/Qwen2.5-0.5B-Instruct; nothing here needs a frontier model. (3) Tokenise a sentence, print the token ids and the decoded pieces, and find a word that splits into three tokens. (4) Run the model, take outputs.logits[0, -1], softmax it, and print the top-10 next tokens with probabilities. (5) Compute the loss against the true next token and call loss.backward(); print the norm of the gradient on one weight matrix. (6) Write the one-sentence pretraining-vs-post-training answer at the top of the notebook. If any step needed a tutorial, do the AI Foundations modules before unit 2.

Go deeper

Next: What might success look like? · Back to the map.