BlueDot Impact · Technical AI Safety · six units · twenty-seven chapters
The problem is technical. So is the way out.
BlueDot's Technical AI Safety course walks from why a capable model is hard to make safe to what a working engineer can build about it: training methods, evaluations, interpretability, defence in depth, and where to go work. This map re-explains each chapter in its own words, pins the course page and every reading at the top, and writes the exercises out so you can do them without the login.
$bluedot.org/courses/technical-ai-safety
▲ amber = understanding the problem — units 1, 2, 4
▼ cyan = doing something about it — units 3, 5, 6
00 the map
Six units, one argument
Frontier models are trained, not programmed, so their goals are learned and their behaviour is only ever sampled. Everything in this course is a response to that one fact.
UNIT 1 · 2h
What "going well" means, why building safe AI is hard even for people who want to, and what future you are actually aiming at.
UNIT 2 · 4h
The levers you have at training time: data, RLHF and its descendants, and the newer techniques that try to shape what a model wants.
UNIT 3 · 6h
Evaluations: capability versus propensity, and how Anthropic, OpenAI, Google DeepMind and Meta actually test before shipping.
UNIT 4 · 2.5h
Interpretability: what it means to say a model "thinks", and what the tools can and cannot see inside one.
UNIT 5 · 2.5h
Assume the model is unsafe and design around it: layered defences, monitoring, and breaking the kill chain.
UNIT 6 · 3h+
Choosing a focus, writing your one-pager, and the roles, fellowships and programs that take people in.
Read it in order once. After that the map is the index: amber pages explain, cyan pages equip.
01 unit 1 · amber
The technical challenge with AI
~2h · 4 chapters
The framing unit. It sets the bar ("AI going well" is more than "no catastrophe"), then shows why a well-intentioned lab still cannot simply build a safe model, and closes by asking what you would want the outcome to be.
1.1
What the course is for, who it is for, and the shape of the problem it is about to hand you.
1.2 · 10 min
Success is not the absence of disaster; it is a world where powerful AI is steerable, honest and broadly beneficial. A picture worth having before the failure modes.
1.3 · 1h 50
The core reading block: specification, generalisation and oversight fail in ways that grow with capability. The chapter every later unit answers.
1.4 · 1h · exercises
Three writing exercises. You state the future you are working toward before the course tells you how to work toward it.
02 unit 2 · amber
Training safer models
~4h · 4 chapters
If behaviour is learned, the first place to intervene is the learning. Data, then feedback, then the techniques that go beyond both.
2.1
The unit's question, and why "just train it to be good" is a research programme rather than an instruction.
2.2 · 50 min
Pretraining data as a safety lever: filtering, curation and synthetic data, and what data cannot fix.
2.3 · 1h 10
RLHF, constitutional AI and their relatives: how preference feedback shapes a model, and the ways it can be gamed.
2.4 · 1h 25
A tour of the newer training-time ideas, from deliberative alignment to unlearning, and what each buys you.
03 unit 3 · cyan
Detecting danger
~6h · 6 chapters
You cannot trust a model you cannot test. Evaluations split into "can it" and "will it", and the four frontier labs each run a different version of the exam.
3.1 · 1h 40
Capability evals versus propensity evals, dangerous-capability thresholds, and why a passed eval is weaker evidence than it looks.
3.2
The frontier safety frameworks as a genre: thresholds, mitigations, and the promise to pause. Pick one lab to study in depth.
3.3 · 1h
The Responsible Scaling Policy and ASL levels, read as an engineering spec.
3.4 · 1h
The Preparedness Framework: tracked categories, capability thresholds, and what triggers a hold.
3.5 · 1h
The Frontier Safety Framework: critical capability levels and the deployment and security mitigations tied to them.
3.6 · 1h
The Advanced AI Scaling Framework (v2, April 2026, formerly the Frontier AI Framework): how Meta reasons about risk when the weights leave the building, now with loss of control as a third domain.
04 unit 4 · amber
Understanding AI
~2.5h · 2 chapters
Interpretability is the bet that we can read a model's reasoning off its weights and activations rather than infer it from its outputs.
4.1 · 35 min
Features, circuits and superposition: the vocabulary for talking about what happens inside a transformer.
4.2 · 2h · 2 exercises · code
Sparse autoencoders, probes and attribution in use, and a hands-on exercise with real model internals.
05 unit 5 · cyan
Minimising harm
~2.5h · 3 chapters
Security thinking applied to models: assume compromise, layer the defences, and cut the chain between a bad intention and a bad outcome.
5.1 · 50 min
Why "the model might be trying to hurt you" is the right design assumption even when it probably is not.
5.2 · 1h
Pick a threat pathway and stress-test the stack against it: the actors, the four catastrophe routes, and how to build a kill chain.
5.3 · 50 min · exercises
Four exercises that make you map an attack path end to end and find the cheapest place to cut it.
06 unit 6 · cyan
Start contributing
~3h · 8 chapters
The course ends where careers start: pick a focus, write the one-pager, and walk the lists of roles, fellowships and programs with the links in hand.
6.1
The unit's overview and the biggest list of doors in the course.
6.2 · 1h 50
Research agendas laid side by side so you can pick one on evidence rather than vibes.
6.3 · 1h · exercise
The artefact the course wants you to leave with: a one-page plan, with the templates and examples it points to.
6.4
Where the jobs are and how the good applications read.
6.5
MATS, Astra, ARENA and the rest: what each selects for and when to apply.
6.6
The policy-side on-ramps for people with a technical background.
6.7
The fellowships that do not fit the two boxes above.
6.8
Courses, bootcamps and programs to keep the momentum after this one.
07 the exercises
Everything you have to build or write yourself
The course's exercises live behind a login. Each is written out on its chapter page, with the code ones flagged so you can do them in a notebook. This is the index.
State the future you are aiming at, the biggest risk to it, and the role you could play.
Reason about what data filtering can and cannot remove from a model.
Preference feedback, reward hacking and the limits of human raters.
Compare the newer techniques on what they assume and what they cost.
Design or critique an evaluation for a dangerous capability.
Read one frontier safety framework as a spec and say where it would fail.
Hands-on with model internals: a probe on GPT-2 small with controls, and the ARENA transformer-interp chapter.
Pick a threat pathway and write the one-sentence threat scenario; then build its kill chain.
Map an attack path and find the cheapest cut.
Rank agendas against your background and the field's needs.
The one-page plan the course wants you to leave with.
08 reading log
Tick them off
Twenty-seven chapters. The checks live in this browser only.
09 refs
Where this came from
The course is BlueDot Impact's; every chapter page here links back to the chapter it explains and to the readings the course assigns. The explanations are original. Read the course for the discussion, the cohort and the certificate.