TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 6 · START CONTRIBUTINGchapter 2 · 1h 50m

Choose your focus

BlueDot Impact · Technical AI Safety · unit 6, chapter 2
TL;DR — Five units of survey are worthless until you convert them into one bet. This chapter makes you pick one technical intervention, argue why it beats the alternatives against the threat model you built in unit 5, and then check whether the world agrees. The mechanism is four research agendas — a funder's, a government lab's, a frontier lab's, a small independent org's — read side by side, so overlap shows what the field considers load-bearing and disagreement shows where the open questions are. The failure it prevents is not choosing wrong; it is choosing nothing because you assume someone competent is already on it.

Every survey course has the same problem at the end: you know enough to have opinions and not enough to feel entitled to them, so the safe move is to keep reading. Unit 6 exists to break that loop, and chapter 2 is where the break happens — not by giving you a ranked list of what to work on, but by handing you four organisations' ranked lists, which visibly disagree, and making the disagreement your problem. Chapters 1 and 3 bracket it: chapter 1 asks what kind of contribution you are aiming at, chapter 3 makes you write the choice down as a one-pager. This is the part in the middle where the choice actually gets made.

The real bottleneck is agency, not information

The chapter opens with a piece that has nothing to do with AI — Neel Nanda on becoming someone who does things — and the placement is deliberate. By now the limiting factor has stopped being knowledge; you can name six threat models and a dozen techniques. What stops people is a pair of beliefs: I couldn't actually do that, and someone smarter is surely on it already. The second is wrong far more often than it feels. Whole categories of work — reproducing a published result on an open-weights model, maintaining an eval nobody owns, writing the tooling that would make someone else's experiment cheap — sit unclaimed not because they are hard but because they are unglamorous and unassigned.

"Being the kind of person who does things, an agent, is a skill, and I think it is a trainable skill."— Become a person who Actually Does Things, Neel Nanda (2021)

The practical form of that claim: the choice you are about to make is reversible and cheap. Not a PhD topic — just what you spend the next twenty to forty hours on, so you have something concrete to show and something concrete to be wrong about. Being wrong in public about a narrow technical question is one of the fastest ways into this field; being undecided in private is one of the slowest.

Three archetypes, three different first moves

The most useful lens the chapter supplies comes from MATS' interviews with 31 people who actually hire for these roles. They found that "AI safety researcher" is not one job. Iterators are empiricists who get a tight experiment loop spinning and turn the crank fast. Connectors work at the conceptual layer, generating and stress-testing the frames that empirical work later tests. Amplifiers multiply other people's output — infrastructure, project management, communication, the unsexy scaffolding without which a five-person research team ships nothing.

"Iterators are strong empiricists who build tight, efficient feedback loops for themselves and their collaborators."— Talent Needs of Technical AI Safety Teams, yams & Carson Jones et al. (2024)

This matters for choosing a focus because the archetypes have wildly different ramp times, and the agendas you are about to read are not equally accessible to each. The report's blunt finding is that Connectors take years of immersion and argument to mature — so "I'll start by developing a novel theory of agent foundations" is a poor opening move eight weeks into a course, however attractive it looks. Iterators can produce something a lab cares about almost immediately, because the bottleneck at the empirical layer is throughput. Amplifiers are the most systematically undervalued by newcomers and the most consistently in demand by the people doing the hiring.

Two caveats the course does not press hard enough. These are hiring lenses, not personality types — most working researchers do all three in different proportions, and the archetype to optimise for is the one describing your next six months, not your soul. And the report is from 2024: AI control was a niche agenda when those interviews happened and is now a standard line item at multiple labs. Read the archetypes as durable, the specific talent gaps as a snapshot.

Reading four agendas as a market, not a menu

The optional resources are the chapter's real payload, and they are worth more read against each other than in sequence. Four organisations with different incentives published what they want worked on:

Where they overlap, the field has converged. Evaluations, interpretability, monitoring of untrusted models, and chain-of-thought faithfulness appear in nearly all of them — your best signal that the problem is real and that work on it will be legible to reviewers, funders, and hiring managers. Where they diverge, the field is still arguing: AISI will spend years on theory yielding asymptotic guarantees, while Redwood's framing assumes powerful AI arrives before any proof does and optimises for protocols that survive a misaligned model today. Notice which side your instincts land on — it determines which agenda's open problems will feel like real problems to you.

One thing to carry: agendas are attributions of interest, not of authorship. A direction on a lab's list does not mean that lab invented it — weak-to-strong generalisation sits on Anthropic's recommended-directions list, but the technique comes from OpenAI's 2023 paper (Burns et al.). Before citing a technique as "X's approach," follow the link to the primary paper and check the byline. Getting this wrong in a one-pager is an unforced error a reviewer notices immediately.

Testing the choice: four questions that do real work

The chapter's second exercise gives you a set of questions that look like homework and are actually a filter. Run them honestly and most candidate focuses fail:

Where this chapter breaks

Three honest limitations. Staleness: these are 2024–2025 documents in a field with roughly annual turnover of fashionable problems, and one of the four organisations has since changed its name — check the org's current page before building a plan on it. Prestige capture: it is easy to pick the direction attached to the most impressive lab rather than the one where your marginal contribution is largest; the funder's list and the small-org backlog are better guides to marginal value. Over-commitment: "pick one" is a forcing function for a decision you can act on this month, not a vow — BlueDot's own guidance is blunt that a strong portfolio means skipping the project and just applying.

Readings, linked

The course budgets 35 minutes for the four core readings and about an hour for the second exercise, totalling 1h 50m with the five optional agenda documents. Start with the MATS talent-needs report — it reframes what you are choosing between before you look at any agenda.

Exercises

  1. Prioritise a single intervention — From everything the course has covered, commit to one technical AI safety technique: evals, interpretability, control protocols, scalable oversight, unlearning, adversarial robustness, monitoring, or another you can name precisely. The constraint that makes this exercise work is the link back to unit 5: choose the technique you believe is most effective against the specific threat model you developed there, not the one that is most interesting in the abstract. Use the four optional agendas as your candidate pool. What a good answer has: the technique named at a level of precision you could search for (not "interpretability" but "probing for deceptive-reasoning features in open-weight chat models"); an explicit statement of the unit-5 threat it addresses; one sentence on the causal path from the technique working to that threat being reduced; and at least one named alternative you rejected, with the reason. If your justification would survive having the technique swapped for a different one, it isn't a justification yet.
  2. Do your own research — Spend roughly an hour investigating the technique you chose and answer three questions. What does success look like? — describe the concrete state of the world in which this technique has worked, in terms someone could check. What is the current status? — are governments, frontier labs, or independent orgs already doing this, at what scale, and if the answer is "barely anyone," work out why: access constraints, missing prerequisites, or genuine neglect. Who could you join or contribute to? — name specific organisations, teams, or open repositories. What a good answer has: at least three cited primary sources published in the last 12–18 months (agenda documents, papers, org pages) rather than a summary of the course; a dated snapshot of who is working on it; an explicit "why not already solved" claim you would defend; and a named gap where your particular background gives you an edge. Write it in a form you can paste straight into chapter 3's one-pager.
  3. Agenda-overlap diff code — Field map extra, not part of the course. Turn the four agendas into a comparison table instead of reading them linearly, so that convergence and disagreement become visible rather than remembered. Build a small script that extracts each agenda's list of research areas, normalises the names to a shared vocabulary, and emits a matrix of area × organisation plus a ranked list of areas by how many organisations name them. What a good answer has: a table where at least one area is claimed by all four sources, at least one is claimed by exactly one, and a short written note on what each pattern implies for a newcomer's marginal value. Start here: (1) save the four agenda pages locally with curl — note that the Coefficient Giving page blocks automated fetches, so copy that one out of a browser by hand; (2) strip to text with trafilatura or readability-lxml; (3) pull candidate headings with a regex over <h2>/<h3> tags, or ask a small local model such as Qwen2.5-7B-Instruct via ollama to emit one JSON array of area names per document; (4) normalise synonyms by hand into a mapping dict — "mech interp" / "understanding model cognition" / "interpretability" are one row; (5) build the matrix with pandas and print it; (6) write three sentences on the areas with a count of 4 and the areas with a count of 1. Runs on a laptop in under two hours and produces an artifact you can put in the one-pager.

Go deeper

Next: Create your 1-pager · Back to the map.