How does AI think?
Units 2 and 3 gave you two levers: change the training so the model behaves better, and build evaluations so you notice when it does not. Both operate entirely on inputs and outputs, and this unit is the course admitting that limit — if your only evidence is "we tried a lot of things and it was fine", you have evidence, not an argument. Chapter 1 hands you the vocabulary you need before any of chapter 2's tooling makes sense.
The ceiling on black-box evidence
Behavioural testing has a structural weakness no budget fixes: you can only test inputs you thought of. For most engineering that is fine — failures are roughly randomly distributed and sampling finds them. Safety-relevant failures are not. The scenario people actually worry about, a model that behaves one way when it thinks it is being evaluated and another when it thinks it is deployed, is defined by being invisible to sampling.
Interpretability's pitch is that the mechanism is a smaller object than the behaviour: inputs are effectively infinite, weights are finite, and reading the computation off the weights means reasoning about the thing itself rather than a sample of its outputs. Whether that survives contact with a frontier model is what this chapter is really about.
Features: the unit of explanation
Start with vision, where the answers are checkable by eye. The Distill circuits programme found units in an image classifier that respond to curves at a specific orientation, then made the claim falsifiable: dataset examples that fire the unit, a synthesised input that maximises it, a curve rotated through 360° with the response tracking the angle, and an ablation that degrades curve-dependent behaviour. That last step is what matters — "this neuron correlates with curves" is cheap; "intervening on it changes what the network does about curves" has teeth.
A feature generalises that: a property of the input the network has learned to represent internally — "curve at 30°", "this token is inside a Python string literal", "the text is Arabic script". The refinement that matters is that a feature is a direction in a layer's activation space, not a unit of it. You could conflate the two in 2015; you cannot now, and the reason is superposition.
Circuits: what the weights actually say
Once you have features, the weights between two layers stop being an opaque matrix and become a readable statement about how one feature is built from others. The canonical example: a car detector draws its strongest positive weight from a wheel detector at the bottom of its receptive field and a window detector at the top. That is not a correlation inferred from behaviour — it is a spatial rule written in the weight values, checkable by ablating either input.
A circuit is such a subgraph: features, the weights connecting them, and the computation that composition implements. In transformers the same move gives you induction heads — a pair of attention heads implementing "I saw A B earlier; I am now looking at A; predict B" — a mechanism you can identify, ablate, and watch in-context learning collapse. The programme also makes a bolder claim: universality, that analogous features keep reappearing across architectures trained on similar data. If it holds, findings transfer between models. If not, every model needs its own map.
Superposition: why the neuron is the wrong unit
Open a real network and most units are polysemantic: one fires for cat faces, car fronts, and the legs of a spider. The lazy read is that the network is messy. Toy Models of Superposition gives the better one: it is being efficient, in a way you would have designed yourself.
The argument runs in three steps. The world has vastly more features worth representing than any layer has dimensions. Features are sparse — "Golden Gate Bridge" is absent from essentially every input, and so is nearly every other specific concept. And given sparsity, you can pack many more than n features into n dimensions by assigning them nearly-but-not-quite orthogonal directions: each pair interferes a little, but two sparse features rarely fire together, so the cost is small and the benefit — thousands of things instead of hundreds — is large. It is compressed sensing, discovered by gradient descent. In the toy setting you can watch the network switch strategies as sparsity rises, from "represent a few features cleanly and discard the rest" to "represent all of them in overlapping directions".
The consequence is severe: activations live in a compressed, non-privileged basis, and the neuron basis is not the one the model uses. Reading a network neuron by neuron is like reading a zip file byte by byte and concluding the data is meaningless. This single fact explains most of what interpretability tooling does.
Sparse autoencoders: undoing the packing
If the model compressed its features into a smaller basis, learn the decompression. Train a small autoencoder on a layer's activations with a hidden layer much wider than the layer and a sparsity penalty forcing only a handful of hidden units active at once; each becomes a candidate feature direction. Towards Monosemanticity ran this on a one-layer transformer and recovered thousands of features from a few hundred neurons — Arabic script, DNA sequences, legal boilerplate — far cleaner than the neurons they came from. Scaling Monosemanticity then ran it on a production model and produced the strongest evidence this field has: clamp the Golden Gate Bridge feature high and the model starts insisting it is the bridge. Not "this direction correlates with X" but "intervene and behaviour moves exactly as the hypothesis predicts."
Three caveats travel with every SAE result. Dictionary size is a hyperparameter you chose, and widening it splits features into finer ones with no principled stopping point. Reconstruction is lossy, so what you read is not the whole computation. And "feature" stays an empirical, not a formal, notion: it is defined by the procedure that found it.
Two camps, and what each one owes you
The chapter splits the field into basic science — reverse-engineer the model completely, every layer, every parameter — and pragmatic — answer one question about one behaviour, the way a doctor diagnoses a symptom without first solving all of biology. They fail differently. Basic science can produce a decade of true, beautiful results that never cash out into a deployment decision. Pragmatic work can produce a story that is locally correct and globally wrong: you find a mechanism behind the bad output, patch it, and the model routes around it. These are bets, not just preferences — on whether superposition is a nuisance to be solved or a permanent fact about how large networks store things.
The case against — which this chapter assigns you
The reading list is not a sales pitch: four of the nine readings attack the paradigm, from three directions.
The reductionism objection. Hendrycks and Hiscott argue that a large network is a complex system whose behaviour comes from an enormous number of weak interactions, so no decomposition into named parts will be faithful enough to act on — the right level of analysis is higher, nearer representations than circuits. They cite specifics, not vibes: sparse autoencoders that "underperformed a simple baseline" at detecting harmful intent, and a major lab deprioritising SAE research over "disappointing results". Read it adversarially in both directions — the piece is polemical, and the underlying results are real.
The theory-of-impact objection. Segerie grants much of the science and attacks the chain from result to safety decision. If you find a circuit, what deployment call changes? For most goals interpretability names, a cheaper direct method — a behavioural eval, adversarial training, a governance intervention — already does the job, making the theory of change redundant rather than complementary. And the work is dual-use: understanding a model well enough to fix it is understanding it well enough to make it stronger.
The detection objection. The sharpest version comes from inside the field. Neel Nanda, one of the paradigm's most prominent practitioners, argues that catching a deceptive superintelligence by interpretability hits the same walls as black-box testing: proving a negative, over an adversarially-chosen space, in a representation you only partly understand.
"Neither interpretability nor black box methods offer a high reliability path to safeguards for superintelligence."— Interpretability Will Not Reliably Find Deceptive AI, Neel Nanda (2025)
His conclusion is not "abandon the field" — it is that interpretability belongs in a defence-in-depth stack as one imperfect instrument among several, which is exactly the argument unit 5 makes about everything else.
Readings, linked
The course budgets 35 minutes for the four core readings; the remaining five are optional and are where the real disagreement lives. Start with the Rational Animations video — it makes features and polysemanticity visual in a way no prose does — then read Hastings-Woodhouse for the transformer vocabulary. If you only add one optional piece, add Scott Alexander's explainer, because it walks the superposition argument slowly.
- What Do Neural Networks Really Learn? Exploring the Brain of an AI Model — Rational Animations (2024) · 15 min · The basic-science approach on image models, animated: features, circuits, and why one neuron represents several unrelated things. Assigned first because the visual case is the one where you can check the claims yourself.
- Introduction to Mechanistic Interpretability — Sarah Hastings-Woodhouse (2024) · 5 min · The course's own overview of the circuits perspective, plus sparse autoencoders and feature steering. The chapter also points at Zoom In: An Introduction to Circuits (Olah et al., Distill, 2020) for the technical version — that is the founding document of the programme and worth the extra half hour.
- Neel Nanda on the race to read AI minds — Robert Wiblin, 80,000 Hours (2025) · 5 min · The course asks only for "The interview in a nutshell" — a state-of-the-field snapshot from someone running interpretability at a frontier lab. For research directions, the chapter also links Nanda's "How to become a mechanistic interpretability researcher".
- The Misguided Quest for Mechanistic AI Interpretability — Dan Hendrycks and Laura Hiscott (2025) · 10 min · The reductionism objection: networks are complex systems, emergent behaviour does not decompose into simple mechanisms, and celebrated techniques including SAEs have underdelivered. The course's own link points at the old Substack domain, which now redirects here.
- MoSSAIC: AI Safety After Mechanism — Farr et al., ODYSSEY 2025 · optional · Names the "causal-mechanistic paradigm" explicitly and proposes a supplementary framework for when it fails, connecting obfuscation results to MIRI-style threat models. Note: OpenReview now sits behind a browser-verification challenge, so the PDF may not open on first click — the forum page above is the stable landing point.
- God Help Us, Let's Try To Understand The Paper On AI Monosemanticity — Scott Alexander, Astral Codex Ten (2023) · 25 min · optional · The best plain-English walkthrough of superposition and sparse autoencoders that exists. Pairs directly with the two primary sources it explains: Toy Models of Superposition (2022) and Towards Monosemanticity (2023).
- Against Almost Every Theory of Impact of Interpretability — Charbel-Raphaël Segerie (2023) · 20 min · optional · The theory-of-impact critique: deception detection is not what interpretability is good at, "feature" is fuzzy, cheaper direct methods usually dominate, and the dual-use risk is underweighted. The linked section is specifically on what the end state of interpretability would even look like.
- Interpretability Will Not Reliably Find Deceptive AI — Neel Nanda (2025) · 10 min · optional · The insider's limits case, and the reading that most changes how you should read every other result here. Ends in defence-in-depth, which is the hinge into unit 5.
- AGI Safety — Connor Leahy, FLI Interpretability Conference, MIT — Conjecture (2023) · 15 min · optional · Listed by the course as "Barriers to Mechanistic Interpretability for AGI Safety"; the actual video title is the one above. A talk on why the interpretability we have may not be the interpretability an AGI safety case would need.
Exercises
This chapter ships no exercises of its own — it is a reading chapter, and the hands-on work is deferred to 4.2, Interpretability in practice. The four below are field map extras, ordered so the first two build the intuitions the chapter only describes.
- Field map extra — catch a polysemantic neuron in the act code — Pick a small open model, choose one MLP neuron in a middle layer, and find the inputs that make it fire hardest. Then argue from the evidence whether the neuron has one meaning or several. What a good answer has: at least 20 max-activating text snippets for your chosen neuron; an honest verdict on whether they share a concept (most will not); a comparison against a sparse-autoencoder feature read at the same layer, showing whether the SAE feature is cleaner; and one sentence on what you cannot conclude from max-activating examples alone. Start here: (1) install TransformerLens in a free Colab and load
gpt2-small— it fits comfortably in a T4 or even on CPU; (2) stream a few thousand snippets from an open corpus such as OpenWebText or Wikipedia through the model withrun_with_cache; (3) record the max activation of your chosen neuron per snippet and keep the top 20; (4) print them with the peak token highlighted; (5) look up the same layer on Neuronpedia and compare the neuron's dashboard to a nearby SAE feature's; (6) write the verdict. - Field map extra — rebuild superposition from scratch code — Reproduce the core Toy Models result on your laptop: show that a network with fewer dimensions than features will represent all of them anyway once the features are sparse enough. What a good answer has: a plot of the learned feature-direction overlap matrix at three sparsity levels; a clear description of the transition from "represent a few features orthogonally and ignore the rest" to "represent everything in overlapping directions"; and one paragraph connecting what you saw to why reading a real model neuron-by-neuron fails. Start here: (1) generate synthetic data with 20 features, each active with probability p and uniform magnitude when active; (2) build a tiny PyTorch model that projects 20 → 5 dims with a weight matrix W, then reconstructs with Wᵀ plus a bias and a ReLU; (3) train on MSE with feature importances decaying geometrically; (4) sweep p from 1.0 down to 0.01, retraining each time; (5) heatmap WᵀW for each run — off-diagonal mass is interference, and its appearance is superposition; (6) check your reading against the original paper after you have formed your own. Runs on CPU in under two minutes per sweep point.
- Field map extra — steer a feature and report the failure — Use a hosted sparse-autoencoder interface to find a feature for a concept you choose, clamp it, and document both what worked and what broke. What a good answer has: the feature you picked and why you believe it means what you think it means; three generations at increasing steering strength; the strength at which output quality collapses; and an assessment of whether the feature is specific to your concept or is really a broader one that merely includes it. How: Neuronpedia hosts Gemma Scope SAE features with a steering interface in the browser — no install, no GPU. This is rung 3 of the evidence ladder above; notice how much harder it is to reach than rung 1.
- Field map extra — write the theory of impact, then attack it — Choose one concrete interpretability result from the readings and write the full causal chain from that result to a decision a real person would make differently: who, about what deployment, with what threshold. Then write the strongest objection Segerie or Hendrycks would make to your chain, and adjudicate it honestly. What a good answer has: a chain with no step that reads "and then it becomes safer"; a named decision-maker and a named decision; an explicit statement of which rung of the evidence ladder the result actually reaches; and a conclusion that is allowed to be "this chain does not close". If every chain you can write fails, say so — that is Segerie's whole argument and you should be able to state it in your own words before you decide whether you believe it.
Go deeper
- A Mathematical Framework for Transformer Circuits — Elhage et al. (2021). The transformer-native version of the circuits picture, and where induction heads are introduced. Denser than anything the course assigns, and the thing to read once the vision examples feel too easy.
- On the Biology of a Large Language Model — Anthropic (2025). Circuit tracing applied to a production model: multi-step reasoning, planning ahead in poetry, and cases where the model's stated reasoning does not match its computed reasoning. The strongest current answer to "does any of this work at scale?" — read alongside the plain-language companion post.
- Curve Detectors — Cammarata et al., Distill (2020). One feature, studied to exhaustion. The best available model of what "we understand this feature" is supposed to mean, and a useful calibration for how much work a single honest claim costs.
- Concrete Steps to Get Started in Mechanistic Interpretability — Neel Nanda. The standard on-ramp if the exercises above went well: a curriculum, paper list, and open problems, from the author of two of this chapter's readings.
- Technical AI Safety Unit 4: Understanding AI — BlueDot Impact. The short unit-intro video embedded on the chapter page itself, framing where interpretability sits relative to units 2 and 3.