TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 4 · UNDERSTANDING AIchapter 2 · 2 hours

Interpretability in practice

BlueDot Impact · Technical AI Safety · unit 4, chapter 2
TL;DR — Chapter 4.1 gave you the vocabulary — features, circuits, superposition. This chapter asks the harder question: what does any of it do for you on a Tuesday? The honest answer is that the ambitious version — reach inside a deployed model and switch off deception — is not shipping. What is shipping is interpretability as an input to other safety machinery: a linear probe that flags a hallucinated entity mid-sentence, a chain of thought you can read for intent, an SAE feature list that gives an auditor a lead, a deliberately-broken model you can test your detectors against. Carry this: the field's current product is better evals and better training data, not surgery, and a technique is only worth its cost if it beats a dumb baseline on a task you actually care about.

Unit 4 opens with the science of what is inside a transformer. This chapter is where the course cashes that science out. It is deliberately a tour of working tools rather than a theory — probes, chain-of-thought monitoring, model organisms, sparse autoencoders in an audit — and it is deliberately unsettled. The course is upfront that techniques here move in and out of fashion as models change. That churn is the real lesson: the useful skill is not memorising today's method but being able to look at a new interpretability result and ask whether it is load-bearing.

Two ambitions, one of which is real right now

There are two things you might want from understanding a model's internals. The first is direct intervention: find the circuit responsible for violence or deception or flattery, and clamp it. This is the version that gets drawn on whiteboards. It requires that the behaviour you care about corresponds to something localised and stable enough to grab, that grabbing it does not wreck everything else, and that the model has not simply routed around you. On today's frontier models, none of those three is reliably true.

The second is indirect application: use the understanding as evidence that feeds a technique which does not itself need to be interpretable. If you learn which pretraining data produces a bad behaviour, you filter that data. If you learn that a model represents "I am currently making this up" as a direction in activation space, you attach a classifier to that direction and route those outputs to a checker. If you learn that models sometimes reason one way and narrate another, you stop treating the narration as evidence and redesign the eval.

Almost every practical win in this chapter is the second kind. That is not a consolation prize. Interpretability's comparative advantage is that it sees things behavioural testing cannot: a behaviour that has not been triggered yet still leaves a representation you can look for. But the deliverable is a signal handed to some other system, not a scalpel.

Probes: the cheapest thing that works

A probing classifier is almost embarrassingly simple. Run text through the model, freeze it, grab the hidden activations at some layer, and train a small model — usually logistic regression on a single vector — to predict a property you labelled: is this sentence in French, is this chess position winning, is this entity fabricated. If a linear probe succeeds, the property is linearly represented at that layer: the model is not just capable of the distinction, it has already computed it and written it down somewhere you can read. BlueDot's own explainer on probing classifiers is the gentle version of this, and it is also the one that lists the traps.

The traps matter more than the method. A probe tells you information is present, not that the model uses it — the direction you found might be a side effect the rest of the network ignores, which is why serious probe work pairs the classifier with a causal intervention (add the direction back in, or ablate it, and see whether behaviour moves). A sufficiently expressive probe can also learn the task itself from generic features, so a strong result may be measuring your probe rather than the model; the standard control is to check that the same probe fails on a shuffled-label or random-network baseline. And a probe can only find what you can already name and label, which means it is structurally blind to anything the model knows that you do not.

The chapter's case study, Hallucination Probes (Obeso, Arditi, Ferrando, Freeman, Holmes and Nanda, 2025), is what a probe looks like when it clears those bars. Instead of a coarse "is this answer true" label, they annotate hallucinations at token level — which specific entity in a long generation is fabricated — and train lightweight linear probes on the hidden states to flag it as the text is being produced. Reported AUC is above 0.85 for linear probes and above 0.89 for LoRA variants in long-form settings, against uncertainty-based baselines that do not clear 0.76. Note the shape of the win: it is not that a probe exists, it is that the probe beats the obvious cheap alternative (the model's own token probabilities) on a realistic task, and does so fast enough to run during generation. That comparison is the thing to look for in every interpretability paper you read.

Chain of thought: an interpretability surface you can lose

Reasoning models externalise part of their computation as text. That is an extraordinary gift — a window into intermediate steps that costs nothing to read, needs no SAE, and generalises across labs. Chain of Thought Monitorability (Korbak et al., 2025, with forty-odd authors from most of the frontier labs and academia) is a position paper arguing that this window should be treated as safety infrastructure and actively defended.

"CoT monitorability may be fragile."— Chain of Thought Monitorability, Korbak et al. (2025)

Fragile in a specific, mechanical way. The window exists because some tasks genuinely require serialised working memory, and because nothing has trained the model to hide. Both conditions can erode: optimise directly against the chain of thought (penalise bad-looking thoughts) and you train a model that thinks bad thoughts quietly; move reasoning into continuous latent space for efficiency and there is no text left to read; scale capability far enough and the model may not need to write anything down. So the paper's ask is that labs measure monitorability, publish the number, and weigh it when making architecture and training decisions — a rare case of an alignment agenda whose main deliverable is a design constraint rather than a tool.

The complementary point is that a chain of thought is not a transcript of the computation. It is another output, generated by the same network, and it can be a rationalisation. The mechanistic evidence for this is direct: in On the Biology of a Large Language Model, Lindsey et al. trace cases where the model works backwards from an answer a human hinted at while narrating a forward derivation. Meanwhile Thought Anchors (Bogdan, Macar, Nanda, Conmy, 2025) shows the text is not uniformly important either: a small number of sentences — usually planning or uncertainty-management moves — dominate where the trace ends up, measured by resampling each sentence and watching the answer distribution shift. If you are building a CoT monitor, both results tell you where to spend attention.

Features and circuits, at production scale

Superposition (chapter 4.1) is why you cannot read neurons directly: the model packs more concepts than it has dimensions. Sparse autoencoders are the standard response — learn an overcomplete, sparsely-activating basis in which the packed directions come apart. Scaling Monosemanticity (Templeton, Conerly et al., 2024) established this works on a deployed model rather than a toy, extracting up to ~34M features from Claude 3 Sonnet's middle-layer residual stream, including features that fire on both concrete instances and abstract discussion of a concept, that transfer across languages and to images, and that steer behaviour when clamped. Safety-relevant abstractions — deception, sycophancy, unsafe code, self-reference — show up in the feature list.

Attribution graphs are the next rung: not "what concepts exist" but "which ones caused this specific output". Replace MLP layers with an interpretable transcoder, then trace the causal path from input tokens through features to the logit. That is the machinery behind the biology paper's case studies — the model representing Texas as an unstated intermediate on the way from Dallas to Austin, planning a rhyme word before writing the line that lands on it. Useful calibration: these graphs are built on a replacement model, cover a fraction of the real computation, and each one is hours of human work to read. Treat them as microscopy, not monitoring.

Model organisms, and the audit that ties it together

You cannot evaluate a detector without positives, and genuinely deceptive frontier models are (hopefully) not available. Model Organisms of Misalignment (Hubinger, Schiefer, Denison, Perez, 2023) proposes manufacturing them: deliberately train models that exhibit the failure — reward hacking, deceptive alignment, a hidden goal — under controlled conditions, then use them as a testbed for both mitigation and measurement.

"in vitro demonstrations of the kinds of failures that might pose existential threats"— Model Organisms of Misalignment, Hubinger et al. (2023)

The obvious objection is that a hand-built failure may be nothing like a naturally emergent one, and the agenda concedes this — the proposed arc runs from heavily-guided demonstrations toward increasingly naturalistic ones, with the early versions valued as existence proofs and calibration targets rather than forecasts.

Auditing language models for hidden objectives (Marks, Treutlein et al., 2025 — also on arXiv) is the chapter's payoff because it uses a model organism to score the tools. The team trained a model with a concealed objective — exploit known quirks of RLHF reward models, "RM sycophancy" — then ran a blind auditing game: four teams, no knowledge of the objective, asked to find it. Three succeeded. SAE features were part of how: an auditor browsing features active while the model wrote a recipe found reward-model-bias features sitting incongruously among cooking features, which is exactly the "surfaces what you did not know to look for" property probes lack.

Read the caveats as carefully as the result. Teams with training-data access also succeeded by plain semantic search, so the honest claim is that interpretability was one viable route, not a unique one. The objective was one the researchers installed, so difficulty was not adversarially set. And the auditors were skilled humans spending real hours. That is still a landmark: it is one of the first end-to-end evaluations of an alignment technique against ground truth, which is the format the field needs far more of.

The practitioner's filter, in one line: does the technique beat a dumb baseline, on a task you care about, with a causal check attached? Probes without an intervention test may be reading a direction the model ignores. SAE features without a steering result are a vocabulary, not a finding. A CoT monitor is worth exactly as much as the faithfulness you can demonstrate. Ask for the baseline column before you ask how the method works.

Readings, linked

The course budgets 1 hour for the five required readings (2 hours for the chapter with the exercise). Start with the BlueDot probing explainer — it is five minutes and it sets up the vocabulary the other four assume — then read Hallucination Probes immediately after as the worked example. The Marks et al. audit is the one to spend real time on.

Exercises

  1. Understanding an interpretability technique code — Pick one technique from the readings (required or optional): linear probing, chain-of-thought monitoring, sparse autoencoders, attribution graphs, model organisms, or thought-anchor resampling. Write it up in plain English, no jargon, under five headings: Goal — what is it trying to uncover? Mechanism — how does it work, step by step? Evidence — what concrete findings has it actually produced? Application — how are those findings being used to improve training or evaluation, if at all? Robustness — name one key limitation or failure mode. The course budgets 45 minutes reading and 15 minutes writing. What a good answer has: a mechanism section someone could re-derive from your description alone; specific findings with numbers or named results rather than "researchers found it useful"; an Application section honest enough to say "not yet" if that is the truth; and a Robustness section that names a failure mode intrinsic to the method (a probe reading a direction the model does not use; an SAE whose features are basis artefacts; a CoT monitor defeated by the training pressure it creates) rather than a generic "more research is needed". Writing it after running the technique once is what separates a summary from understanding — hence the code path below. Start here (probes on a laptop or free Colab, ~1 hour): (1) pip install transformer_lens torch scikit-learn, load gpt2-small — it is 124M parameters and runs on CPU, comfortably on a free Colab T4. (2) Build a labelled set of a few hundred short prompts with a binary property you care about; the classic easy start is present-vs-past tense or English-vs-French, then graduate to something safety-flavoured like "the assistant is about to refuse". (3) Run model.run_with_cache(prompts) and pull resid_post at the final token for every layer. (4) Fit logistic regression per layer, plot held-out accuracy against layer depth — you will see where the model computes the property. (5) The step most people skip: take the probe's weight vector, add it to the residual stream at that layer with a forward hook, and check whether generations actually change. That is the difference between "the information is there" and "the model uses it". (6) Run the control — refit on shuffled labels, and on a randomly-initialised model — and report those numbers next to your real one. For an SAE version instead, swap steps 3–5 for a Gemma Scope SAE on gemma-2-2b (pretrained, free weights, Colab tutorial in the model card) and inspect which features fire on your positive set via Neuronpedia; nnsight is the alternative to TransformerLens if you want the same hooks with remote execution on larger models.
  2. ARENA chapter 1: transformer interpretability code (the course's optional hands-on) — Work through the mechanistic interpretability notebooks: reverse-engineer induction heads in a two-layer attention-only model, replicate the indirect object identification circuit in GPT-2 small with activation patching, then the superposition and SAE material. The course flags this as difficult and expects at least a full day even for experienced ML engineers, and suggests finding collaborators in your cohort. What a good answer has: a working notebook where each claim about a head or a feature is backed by an intervention — patch it, ablate it, and show the metric move — not just an attention heatmap that looks suggestive. Start here: (1) Open learn.arena.education chapter 1 and use the Colab links rather than a local install; a free T4 is enough for everything up to the SAE sections. (2) If transformer_lens hooks feel opaque, do the "Transformer from scratch" section first — the rest assumes you know what resid_pre, attn_out and hook_z refer to. (3) Do "Intro to Mech Interp" (induction heads) before IOI; IOI is the same tools at ten times the complexity. (4) Budget the day the course warns about, and stop after IOI if you are short — the circuit-discovery loop is the transferable skill. (5) To connect it back to this chapter's readings, finish with circuit-tracer on Gemma-2 2B and reproduce a small attribution graph of your own.

Go deeper

Next: Assuming harm · Back to the map.