Choose your focus
Every survey course has the same problem at the end: you know enough to have opinions and not enough to feel entitled to them, so the safe move is to keep reading. Unit 6 exists to break that loop, and chapter 2 is where the break happens — not by giving you a ranked list of what to work on, but by handing you four organisations' ranked lists, which visibly disagree, and making the disagreement your problem. Chapters 1 and 3 bracket it: chapter 1 asks what kind of contribution you are aiming at, chapter 3 makes you write the choice down as a one-pager. This is the part in the middle where the choice actually gets made.
The real bottleneck is agency, not information
The chapter opens with a piece that has nothing to do with AI — Neel Nanda on becoming someone who does things — and the placement is deliberate. By now the limiting factor has stopped being knowledge; you can name six threat models and a dozen techniques. What stops people is a pair of beliefs: I couldn't actually do that, and someone smarter is surely on it already. The second is wrong far more often than it feels. Whole categories of work — reproducing a published result on an open-weights model, maintaining an eval nobody owns, writing the tooling that would make someone else's experiment cheap — sit unclaimed not because they are hard but because they are unglamorous and unassigned.
"Being the kind of person who does things, an agent, is a skill, and I think it is a trainable skill."— Become a person who Actually Does Things, Neel Nanda (2021)
The practical form of that claim: the choice you are about to make is reversible and cheap. Not a PhD topic — just what you spend the next twenty to forty hours on, so you have something concrete to show and something concrete to be wrong about. Being wrong in public about a narrow technical question is one of the fastest ways into this field; being undecided in private is one of the slowest.
Three archetypes, three different first moves
The most useful lens the chapter supplies comes from MATS' interviews with 31 people who actually hire for these roles. They found that "AI safety researcher" is not one job. Iterators are empiricists who get a tight experiment loop spinning and turn the crank fast. Connectors work at the conceptual layer, generating and stress-testing the frames that empirical work later tests. Amplifiers multiply other people's output — infrastructure, project management, communication, the unsexy scaffolding without which a five-person research team ships nothing.
"Iterators are strong empiricists who build tight, efficient feedback loops for themselves and their collaborators."— Talent Needs of Technical AI Safety Teams, yams & Carson Jones et al. (2024)
This matters for choosing a focus because the archetypes have wildly different ramp times, and the agendas you are about to read are not equally accessible to each. The report's blunt finding is that Connectors take years of immersion and argument to mature — so "I'll start by developing a novel theory of agent foundations" is a poor opening move eight weeks into a course, however attractive it looks. Iterators can produce something a lab cares about almost immediately, because the bottleneck at the empirical layer is throughput. Amplifiers are the most systematically undervalued by newcomers and the most consistently in demand by the people doing the hiring.
Two caveats the course does not press hard enough. These are hiring lenses, not personality types — most working researchers do all three in different proportions, and the archetype to optimise for is the one describing your next six months, not your soul. And the report is from 2024: AI control was a niche agenda when those interviews happened and is now a standard line item at multiple labs. Read the archetypes as durable, the specific talent gaps as a snapshot.
Reading four agendas as a market, not a menu
The optional resources are the chapter's real payload, and they are worth more read against each other than in sequence. Four organisations with different incentives published what they want worked on:
- A funder's list. The technical AI safety RFP from Coefficient Giving (the organisation the course still calls Open Philanthropy — it rebranded in late 2025) enumerates roughly twenty-one areas it will write cheques for, at scales from API credits to seed funding for new orgs. A funder's list is deliberately broad: it tells you what is fundable, which is a different question from what is important, but a much more actionable one.
- A government lab's theory. The UK AISI Alignment Team agenda is the opposite shape: not a list of areas but a single argument. Decompose alignment proposals into safety-case sketches to expose the gaps; treat honesty as the necessary condition to attack first; pursue asymptotic guarantees with theory and empirics together; then aim at automating alignment research itself. The most theory-forward of the four, and the most explicit about what would count as sufficient evidence.
- A frontier lab's inventory. Anthropic's Alignment Science team's recommended directions runs to roughly seventeen concrete directions across oversight, model cognition, and robustness — capability and alignment evals, interpretability, chain-of-thought faithfulness, behavioural and activation monitoring, weak-to-strong and easy-to-hard generalisation, honesty, jailbreak benchmarks, unlearning. An inventory of problems a lab would hire you to solve tomorrow.
- An independent org's backlog. Redwood's project proposals are the least polished and the most immediately usable: control protocols and monitoring, training-time alignment and reward hacking, interpretability, plus alignment-faking extensions. Redwood published them so outsiders could pick one up — the shortest path from "I read an agenda" to "I am doing a project."
Where they overlap, the field has converged. Evaluations, interpretability, monitoring of untrusted models, and chain-of-thought faithfulness appear in nearly all of them — your best signal that the problem is real and that work on it will be legible to reviewers, funders, and hiring managers. Where they diverge, the field is still arguing: AISI will spend years on theory yielding asymptotic guarantees, while Redwood's framing assumes powerful AI arrives before any proof does and optimises for protocols that survive a misaligned model today. Notice which side your instincts land on — it determines which agenda's open problems will feel like real problems to you.
Testing the choice: four questions that do real work
The chapter's second exercise gives you a set of questions that look like homework and are actually a filter. Run them honestly and most candidate focuses fail:
- What does success look like? If you cannot describe the state of the world in which this technique has worked — a measurable property, a class of failure that no longer occurs, a safety case that now closes — you have chosen a topic, not an intervention.
- What is the current status, and if nobody is doing it, why not? The highest-yield question and the one most often skipped. An unworked problem usually has a reason: it needs frontier-model access, a data source nobody has, or a result that does not exist yet. Sometimes the reason is just that it is boring and unassigned — that is the case you want.
- Who could you join or contribute to? Naming three organisations, or three people whose recent output you have read, converts a topic into a path. If you can't name them, the topic may be real but your map of it is not.
- Does it bite on your threat model? The link back to unit 5 keeps this from being a popularity contest. Interpretability is excellent work; it is not obviously the fastest lever against a threat model dominated by, say, misuse of open-weight models by low-resource actors.
Where this chapter breaks
Three honest limitations. Staleness: these are 2024–2025 documents in a field with roughly annual turnover of fashionable problems, and one of the four organisations has since changed its name — check the org's current page before building a plan on it. Prestige capture: it is easy to pick the direction attached to the most impressive lab rather than the one where your marginal contribution is largest; the funder's list and the small-org backlog are better guides to marginal value. Over-commitment: "pick one" is a forcing function for a decision you can act on this month, not a vow — BlueDot's own guidance is blunt that a strong portfolio means skipping the project and just applying.
Readings, linked
The course budgets 35 minutes for the four core readings and about an hour for the second exercise, totalling 1h 50m with the five optional agenda documents. Start with the MATS talent-needs report — it reframes what you are choosing between before you look at any agenda.
- Become a person who Actually Does Things — Neel Nanda (2021; the course lists it as 2022) · 5 min · The agency argument: doing things is a trainable skill, and "someone else is probably on it" is usually false. Sets the emotional preconditions for making a choice at all. Also on LessWrong.
- Should you do an AI safety research / engineering project? — Li-Lian Ang (2025) · 5 min · A go/no-go filter: do a project if you can code and are stuck applying or still exploring; skip it if your portfolio is already strong and just apply. Prevents the project from becoming procrastination.
- The software engineer's guide to making your first AI safety contribution in <1 week — Li-Lian Ang (2025) · 10 min · The operational recipe once you have chosen: block 20–40 hours, pick one of three shapes (fix an open-source safety tool, replicate and extend a result, make existing research reproducible), then spend real time writing it up and shipping it publicly.
- Talent Needs of Technical AI Safety Teams — yams, Carson Jones, deus_ex_maki & Ryan Kidd, MATS (2024) · 15 min · Thirty-one interviews with hiring leaders, distilled into the Iterator / Connector / Amplifier archetypes and their very different ramp times. The course asks you to read to the end of "So How Do You Make an AI Safety Professional?" — that section is where the ramp-time argument lands.
- Technical AI Safety RFP: Research Areas — Coefficient Giving, formerly Open Philanthropy (2025) · optional · ~21 funded research areas with eligibility notes, from interpretability and scalable oversight to control, evals, and robustness. The best single map of what is currently fundable; see the RFP overview for grant sizes and process.
- UK AISI's Alignment Team: Research Agenda — Benjamin Hilton, Jacob Pfau, Marie Davidsen Buhl & Geoffrey Irving (2025) · optional · Safety-case decomposition, honesty as the first necessary condition, asymptotic guarantees, and automating alignment research. The most theory-forward of the four agendas and the clearest about what would count as sufficient evidence.
- Recommendations for Technical AI Safety Research Directions — Anthropic's Alignment Science team (2025) · optional · Seventeen concrete directions across evaluation, model cognition, monitoring, oversight, and robustness. Note that inclusion signals interest, not authorship — several directions originate elsewhere. Mirrored on the Alignment Forum.
- Recent Redwood Research project proposals — Ryan Greenblatt, Buck Shlegeris, Julian Stastny, Josh Clymer et al. (2025) · optional · Deliberately unpolished project ideas across control, training-time alignment and reward hacking, interpretability, and alignment-faking extensions. The shortest distance from reading an agenda to starting a project.
- I'm an experienced software engineer. How can I contribute to AI safety? — Li-Lian Ang (2025) · 5 min, optional · Three routes that do not require ML research credentials: scale existing safety experiments, build researcher tooling and infrastructure, or replicate and extend published work. The Amplifier archetype, made concrete.
Exercises
- Prioritise a single intervention — From everything the course has covered, commit to one technical AI safety technique: evals, interpretability, control protocols, scalable oversight, unlearning, adversarial robustness, monitoring, or another you can name precisely. The constraint that makes this exercise work is the link back to unit 5: choose the technique you believe is most effective against the specific threat model you developed there, not the one that is most interesting in the abstract. Use the four optional agendas as your candidate pool. What a good answer has: the technique named at a level of precision you could search for (not "interpretability" but "probing for deceptive-reasoning features in open-weight chat models"); an explicit statement of the unit-5 threat it addresses; one sentence on the causal path from the technique working to that threat being reduced; and at least one named alternative you rejected, with the reason. If your justification would survive having the technique swapped for a different one, it isn't a justification yet.
- Do your own research — Spend roughly an hour investigating the technique you chose and answer three questions. What does success look like? — describe the concrete state of the world in which this technique has worked, in terms someone could check. What is the current status? — are governments, frontier labs, or independent orgs already doing this, at what scale, and if the answer is "barely anyone," work out why: access constraints, missing prerequisites, or genuine neglect. Who could you join or contribute to? — name specific organisations, teams, or open repositories. What a good answer has: at least three cited primary sources published in the last 12–18 months (agenda documents, papers, org pages) rather than a summary of the course; a dated snapshot of who is working on it; an explicit "why not already solved" claim you would defend; and a named gap where your particular background gives you an edge. Write it in a form you can paste straight into chapter 3's one-pager.
- Agenda-overlap diff code — Field map extra, not part of the course. Turn the four agendas into a comparison table instead of reading them linearly, so that convergence and disagreement become visible rather than remembered. Build a small script that extracts each agenda's list of research areas, normalises the names to a shared vocabulary, and emits a matrix of area × organisation plus a ranked list of areas by how many organisations name them. What a good answer has: a table where at least one area is claimed by all four sources, at least one is claimed by exactly one, and a short written note on what each pattern implies for a newcomer's marginal value. Start here: (1) save the four agenda pages locally with
curl— note that the Coefficient Giving page blocks automated fetches, so copy that one out of a browser by hand; (2) strip to text withtrafilaturaorreadability-lxml; (3) pull candidate headings with a regex over<h2>/<h3>tags, or ask a small local model such as Qwen2.5-7B-Instruct viaollamato emit one JSON array of area names per document; (4) normalise synonyms by hand into a mapping dict — "mech interp" / "understanding model cognition" / "interpretability" are one row; (5) build the matrix withpandasand print it; (6) write three sentences on the areas with a count of 4 and the areas with a count of 1. Runs on a laptop in under two hours and produces an artifact you can put in the one-pager.
Go deeper
- Weak-to-Strong Generalization — Burns et al., OpenAI (2023). The primary source for a direction that appears on other labs' agenda lists; a useful worked example of checking authorship before you attribute a technique.
- MATS Program — the organisation behind the talent-needs report, and one of the main structured on-ramps for people at exactly this stage. Their mentor list doubles as a live directory of who is actively taking on research directions.
- Anthropic Alignment Science blog — the recommended-directions post ages; the blog is where the updates and follow-on results land, so check it before citing the 2025 list as current.
- BlueDot Technical AI Safety Project — the structured follow-on if your chosen focus turns into a portfolio project, and AGI Strategy if the exercise revealed that your uncertainty is about theory of change rather than technique.
- 80,000 Hours career advising — free one-to-one calls; the highest-leverage use is testing your chosen focus against someone who tracks the hiring market, not asking them to choose for you.