TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 4 · UNDERSTANDING AIchapter 1 · 35 min

How does AI think?

BlueDot Impact · Technical AI Safety · unit 4, chapter 1
TL;DR — Everything so far has judged models from the outside: train, sample, score. Interpretability judges them from the inside, with a three-word vocabulary. A feature is a direction in activation space that means something; a circuit is features wired together by weights to compute something; superposition is why features are not neurons — a layer packs far more features than it has dimensions by giving them almost-orthogonal directions, which is exactly why single neurons fire for unrelated things. Sparse autoencoders are the best current attempt to undo that packing. Remember: the vocabulary is load-bearing, and four of this chapter's nine readings argue it will not scale to the problem it was invented for.

Units 2 and 3 gave you two levers: change the training so the model behaves better, and build evaluations so you notice when it does not. Both operate entirely on inputs and outputs, and this unit is the course admitting that limit — if your only evidence is "we tried a lot of things and it was fine", you have evidence, not an argument. Chapter 1 hands you the vocabulary you need before any of chapter 2's tooling makes sense.

The ceiling on black-box evidence

Behavioural testing has a structural weakness no budget fixes: you can only test inputs you thought of. For most engineering that is fine — failures are roughly randomly distributed and sampling finds them. Safety-relevant failures are not. The scenario people actually worry about, a model that behaves one way when it thinks it is being evaluated and another when it thinks it is deployed, is defined by being invisible to sampling.

Interpretability's pitch is that the mechanism is a smaller object than the behaviour: inputs are effectively infinite, weights are finite, and reading the computation off the weights means reasoning about the thing itself rather than a sample of its outputs. Whether that survives contact with a frontier model is what this chapter is really about.

Features: the unit of explanation

Start with vision, where the answers are checkable by eye. The Distill circuits programme found units in an image classifier that respond to curves at a specific orientation, then made the claim falsifiable: dataset examples that fire the unit, a synthesised input that maximises it, a curve rotated through 360° with the response tracking the angle, and an ablation that degrades curve-dependent behaviour. That last step is what matters — "this neuron correlates with curves" is cheap; "intervening on it changes what the network does about curves" has teeth.

A feature generalises that: a property of the input the network has learned to represent internally — "curve at 30°", "this token is inside a Python string literal", "the text is Arabic script". The refinement that matters is that a feature is a direction in a layer's activation space, not a unit of it. You could conflate the two in 2015; you cannot now, and the reason is superposition.

Circuits: what the weights actually say

Once you have features, the weights between two layers stop being an opaque matrix and become a readable statement about how one feature is built from others. The canonical example: a car detector draws its strongest positive weight from a wheel detector at the bottom of its receptive field and a window detector at the top. That is not a correlation inferred from behaviour — it is a spatial rule written in the weight values, checkable by ablating either input.

A circuit is such a subgraph: features, the weights connecting them, and the computation that composition implements. In transformers the same move gives you induction heads — a pair of attention heads implementing "I saw A B earlier; I am now looking at A; predict B" — a mechanism you can identify, ablate, and watch in-context learning collapse. The programme also makes a bolder claim: universality, that analogous features keep reappearing across architectures trained on similar data. If it holds, findings transfer between models. If not, every model needs its own map.

Superposition: why the neuron is the wrong unit

Open a real network and most units are polysemantic: one fires for cat faces, car fronts, and the legs of a spider. The lazy read is that the network is messy. Toy Models of Superposition gives the better one: it is being efficient, in a way you would have designed yourself.

The argument runs in three steps. The world has vastly more features worth representing than any layer has dimensions. Features are sparse — "Golden Gate Bridge" is absent from essentially every input, and so is nearly every other specific concept. And given sparsity, you can pack many more than n features into n dimensions by assigning them nearly-but-not-quite orthogonal directions: each pair interferes a little, but two sparse features rarely fire together, so the cost is small and the benefit — thousands of things instead of hundreds — is large. It is compressed sensing, discovered by gradient descent. In the toy setting you can watch the network switch strategies as sparsity rises, from "represent a few features cleanly and discard the rest" to "represent all of them in overlapping directions".

The consequence is severe: activations live in a compressed, non-privileged basis, and the neuron basis is not the one the model uses. Reading a network neuron by neuron is like reading a zip file byte by byte and concluding the data is meaningless. This single fact explains most of what interpretability tooling does.

Sparse autoencoders: undoing the packing

If the model compressed its features into a smaller basis, learn the decompression. Train a small autoencoder on a layer's activations with a hidden layer much wider than the layer and a sparsity penalty forcing only a handful of hidden units active at once; each becomes a candidate feature direction. Towards Monosemanticity ran this on a one-layer transformer and recovered thousands of features from a few hundred neurons — Arabic script, DNA sequences, legal boilerplate — far cleaner than the neurons they came from. Scaling Monosemanticity then ran it on a production model and produced the strongest evidence this field has: clamp the Golden Gate Bridge feature high and the model starts insisting it is the bridge. Not "this direction correlates with X" but "intervene and behaviour moves exactly as the hypothesis predicts."

Three caveats travel with every SAE result. Dictionary size is a hyperparameter you chose, and widening it splits features into finer ones with no principled stopping point. Reconstruction is lossy, so what you read is not the whole computation. And "feature" stays an empirical, not a formal, notion: it is defined by the procedure that found it.

Two camps, and what each one owes you

The chapter splits the field into basic science — reverse-engineer the model completely, every layer, every parameter — and pragmatic — answer one question about one behaviour, the way a doctor diagnoses a symptom without first solving all of biology. They fail differently. Basic science can produce a decade of true, beautiful results that never cash out into a deployment decision. Pragmatic work can produce a story that is locally correct and globally wrong: you find a mechanism behind the bad output, patch it, and the model routes around it. These are bets, not just preferences — on whether superposition is a nuisance to be solved or a permanent fact about how large networks store things.

The case against — which this chapter assigns you

The reading list is not a sales pitch: four of the nine readings attack the paradigm, from three directions.

The reductionism objection. Hendrycks and Hiscott argue that a large network is a complex system whose behaviour comes from an enormous number of weak interactions, so no decomposition into named parts will be faithful enough to act on — the right level of analysis is higher, nearer representations than circuits. They cite specifics, not vibes: sparse autoencoders that "underperformed a simple baseline" at detecting harmful intent, and a major lab deprioritising SAE research over "disappointing results". Read it adversarially in both directions — the piece is polemical, and the underlying results are real.

The theory-of-impact objection. Segerie grants much of the science and attacks the chain from result to safety decision. If you find a circuit, what deployment call changes? For most goals interpretability names, a cheaper direct method — a behavioural eval, adversarial training, a governance intervention — already does the job, making the theory of change redundant rather than complementary. And the work is dual-use: understanding a model well enough to fix it is understanding it well enough to make it stronger.

The detection objection. The sharpest version comes from inside the field. Neel Nanda, one of the paradigm's most prominent practitioners, argues that catching a deceptive superintelligence by interpretability hits the same walls as black-box testing: proving a negative, over an adversarially-chosen space, in a representation you only partly understand.

"Neither interpretability nor black box methods offer a high reliability path to safeguards for superintelligence."— Interpretability Will Not Reliably Find Deceptive AI, Neel Nanda (2025)

His conclusion is not "abandon the field" — it is that interpretability belongs in a defence-in-depth stack as one imperfect instrument among several, which is exactly the argument unit 5 makes about everything else.

Carry this away: interpretability results sit on a ladder of evidential strength, and most public claims sit lower than their write-ups suggest. Rung 1: the unit correlates with a concept in dataset examples. Rung 2: ablating it degrades the related behaviour. Rung 3: steering it moves behaviour the predicted way and nothing else. Rung 4: the technique beats a simple non-interpretability baseline on a task someone cares about. Safety arguments need rungs 3 and 4 — name a paper's rung before deciding what it proves.

Readings, linked

The course budgets 35 minutes for the four core readings; the remaining five are optional and are where the real disagreement lives. Start with the Rational Animations video — it makes features and polysemanticity visual in a way no prose does — then read Hastings-Woodhouse for the transformer vocabulary. If you only add one optional piece, add Scott Alexander's explainer, because it walks the superposition argument slowly.

Exercises

This chapter ships no exercises of its own — it is a reading chapter, and the hands-on work is deferred to 4.2, Interpretability in practice. The four below are field map extras, ordered so the first two build the intuitions the chapter only describes.

  1. Field map extra — catch a polysemantic neuron in the act code — Pick a small open model, choose one MLP neuron in a middle layer, and find the inputs that make it fire hardest. Then argue from the evidence whether the neuron has one meaning or several. What a good answer has: at least 20 max-activating text snippets for your chosen neuron; an honest verdict on whether they share a concept (most will not); a comparison against a sparse-autoencoder feature read at the same layer, showing whether the SAE feature is cleaner; and one sentence on what you cannot conclude from max-activating examples alone. Start here: (1) install TransformerLens in a free Colab and load gpt2-small — it fits comfortably in a T4 or even on CPU; (2) stream a few thousand snippets from an open corpus such as OpenWebText or Wikipedia through the model with run_with_cache; (3) record the max activation of your chosen neuron per snippet and keep the top 20; (4) print them with the peak token highlighted; (5) look up the same layer on Neuronpedia and compare the neuron's dashboard to a nearby SAE feature's; (6) write the verdict.
  2. Field map extra — rebuild superposition from scratch code — Reproduce the core Toy Models result on your laptop: show that a network with fewer dimensions than features will represent all of them anyway once the features are sparse enough. What a good answer has: a plot of the learned feature-direction overlap matrix at three sparsity levels; a clear description of the transition from "represent a few features orthogonally and ignore the rest" to "represent everything in overlapping directions"; and one paragraph connecting what you saw to why reading a real model neuron-by-neuron fails. Start here: (1) generate synthetic data with 20 features, each active with probability p and uniform magnitude when active; (2) build a tiny PyTorch model that projects 20 → 5 dims with a weight matrix W, then reconstructs with Wᵀ plus a bias and a ReLU; (3) train on MSE with feature importances decaying geometrically; (4) sweep p from 1.0 down to 0.01, retraining each time; (5) heatmap WᵀW for each run — off-diagonal mass is interference, and its appearance is superposition; (6) check your reading against the original paper after you have formed your own. Runs on CPU in under two minutes per sweep point.
  3. Field map extra — steer a feature and report the failure — Use a hosted sparse-autoencoder interface to find a feature for a concept you choose, clamp it, and document both what worked and what broke. What a good answer has: the feature you picked and why you believe it means what you think it means; three generations at increasing steering strength; the strength at which output quality collapses; and an assessment of whether the feature is specific to your concept or is really a broader one that merely includes it. How: Neuronpedia hosts Gemma Scope SAE features with a steering interface in the browser — no install, no GPU. This is rung 3 of the evidence ladder above; notice how much harder it is to reach than rung 1.
  4. Field map extra — write the theory of impact, then attack it — Choose one concrete interpretability result from the readings and write the full causal chain from that result to a decision a real person would make differently: who, about what deployment, with what threshold. Then write the strongest objection Segerie or Hendrycks would make to your chain, and adjudicate it honestly. What a good answer has: a chain with no step that reads "and then it becomes safer"; a named decision-maker and a named decision; an explicit statement of which rung of the evidence ladder the result actually reaches; and a conclusion that is allowed to be "this chain does not close". If every chain you can write fails, say so — that is Segerie's whole argument and you should be able to state it in your own words before you decide whether you believe it.

Go deeper

Next: Interpretability in practice — the same ideas with your hands on them: sparse autoencoders, probes and attribution applied to a real model. Back to the map.