Your next steps
Units 1 through 5 built a map of the problem: what goes wrong, which techniques exist, where they fail. Unit 6 is the turn from understanding to action, and chapter 1 is its framing — why acting now beats waiting until you feel qualified, and a pointer at everything the remaining seven chapters detail.
The number that does the arguing
Most career advice opens by telling you the field is competitive. This one opens by telling you it is tiny: fewer than 2,000 people worldwide full-time, so for almost any specific safety question the headcount is between zero and a dozen. Three things follow, and they generate the rest of the chapter. Marginal people matter more than marginal credentials — in a field of 200,000 the tenth-best candidate is a rounding error; in a field of 2,000 they are a meaningful share of the effort. Nobody has the map — no accreditation body, no settled curriculum, no consensus on which agenda is right, which is uncomfortable but also means you need no permission to hold an opinion. And neglectedness is legible — you can read the major labs' and funders' open-problem lists in an afternoon and see what nobody has claimed. Try that in machine learning generally.
Five heuristics, and where each one breaks
The chapter distils several well-known essays into five rules. Each is useful, and each has a failure mode worth naming, because "good advice applied without judgment" is how people waste a year.
1. When in doubt, apply. An application costs hours, a rejection costs nothing durable, and committees predict fit worse than they think. Where it breaks: the rule is scoped to reversible actions — a replication is cheap to be wrong about, noisy claims in a policy fight are not, because reputational damage in a small field compounds.
2. Work in public. A write-up gets feedback, substitutes for institutional signal, and is how collaborators find you. Where it breaks: volume without a quality bar builds an anti-portfolio. One careful replication with honest negative results beats twelve hot takes.
3. Don't wait for permission. The instinct that someone smarter is already on your concern is usually wrong, and even when right, teams are short-handed. Where it breaks: "nobody is working on this" sometimes means "this was tried and it doesn't work" — ten minutes of literature search is the cost of finding out.
4. Weigh neglectedness against career capital. Visible paths carry the best career capital precisely because they are visible, which is also why they are crowded. The counterweight is the Pareto-frontier argument: you may not be the best interpretability researcher in the world, but you might be best at the intersection of your unusual background and safety — and intersections are where uncontested problems live.
"If your arguments can justify anything, then your arguments imply nothing."— Galaxy brain resistance, Vitalik Buterin (2025)
5. Prefer galaxy-brain-resistant plans. The sharpest of the five. Long inferential chains concluding "and therefore the highest-impact thing happens to be the thing I already wanted to do" are evidence about the reasoner, not the world; the chapter's canonical example is joining a capabilities team to influence it from inside, an argument constructible for almost any position. Ask whether you could derive the opposite conclusion from equally plausible steps. If yes, you have a rationalisation, not a strategy. Where it breaks: pushed too far it argues against all non-obvious action — the test is whether the reasoning is load-bearing and unfalsifiable, not whether it is unusual.
Three doors, plus the one the chapter adds at the end
Research produces new knowledge about model behaviour: training models, running experiments, or importing structure from other fields — model organisms is a biology idea, red-teaming a security one. Deep ML fluency helps enormously; a non-ML background is not disqualifying if it gives you a niche.
Engineering builds the infrastructure that lets researchers run experiments at all: eval harnesses, auditing tools, sandboxes, data pipelines. It is undersold relative to its leverage — one good framework multiplies every researcher who adopts it, and "evals research" in practice is mostly engineering and operations. If you already ship software, this is the shortest path from your current skills to real contribution. Founding starts the thing when no existing org will house it. Highest variance, and the door where agency substitutes most directly for credentials.
The boundary the chapter draws
The final note is honest about a limit of the whole course: technical safety work is necessary but not sufficient. A perfect alignment technique only helps if developers are made to use it (governance), if the weights are not stolen (AI security), and if leaked harms hit a society with defences (biosecurity, cybersecurity). Read it as a warning against treating your chosen technique as the whole solution — not as a reason to abandon technical work.
Readings, linked
This chapter assigns no formal reading block; these are the essays it cites inline as the source of its advice. If you read one, read the 80,000 Hours career review — it is the widest survey. If you are choosing between "go independent" and "get mentorship", read Hobbhahn.
- AI safety technical research — 80,000 Hours career review · the broadest treatment of paths, fit tests and failure modes; the chapter's five heuristics are a compression of advice like this. Note: the course links
/career-reviews/ai-safety-researcher/, which now redirects here. - Some advice on independent research — Marius Hobbhahn (2022) · when independent research is the right call and what it demands: early feedback, deliberate collaboration, and self-imposed accountability through planned outputs. Written before Hobbhahn founded Apollo Research.
- How To Become A Mechanistic Interpretability Researcher — Neel Nanda (2025) · the most concrete "here is the sequence" essay on this list: learn minimal basics by coding, then a month of learning the ropes, then 2–4 week mini-projects, then real projects. Interp-specific but the shape generalises.
- How I Formed My Own Views About AI Safety — Neel Nanda (2022) · the antidote to "I'll start once I understand the field properly." Inside views are a spectrum, they develop through doing, and you can act without one.
- How to be more agentic — Cate Hall (2024) · the chapter's source for "high agency". Practical rather than inspirational: the specific habits that separate people who make things happen from people who wait.
- Galaxy brain resistance — Vitalik Buterin (2025) · the source of heuristic 5, and the best single tool on this page for auditing your own career reasoning.
- Being the (Pareto) Best in the World — johnswentworth (2019) · why "best in the world at the intersection of three things" is achievable when "best in the world at one thing" is not. Underwrites the chapter's advice to lean on an unusual background.
"…forming an inside view by going out in the world and doing things - not just by hiding away and thinking really hard."— How I Formed My Own Views About AI Safety, Neel Nanda (2022)
The doors, annotated
Every organisation, list, tool and program this chapter names, with a verified link and what it actually is. Dates and deadlines are as the linked pages state them at the time of writing (September 2026) — always re-check on the page itself.
Open problems you can start on today
- TAIS RFP: Research Areas — Coefficient Giving · a major funder's request-for-proposals, split into the technical areas it wants to fund. Who it's for: anyone deciding what to work on — funder RFPs tell you what is both important and unclaimed. The course labels this link "Open Philanthropy"; Open Philanthropy renamed itself Coefficient Giving in November 2025, which is why the URL looks unfamiliar.
- Recent Redwood Research project proposals — Ryan Greenblatt, Buck Shlegeris, Julian Stastny, Josh Clymer and colleagues (July 2025) · roughly forty concrete proposals across AI control, training-time alignment and interpretability, most with a linked doc spelling out the experiment. Who it's for: the single best source on this page if you want a project you could start this weekend.
- AISI Research Agenda — UK AI Security Institute · the priorities of the largest government AI safety team, with the risk models behind them. Who it's for: people who want to know what a state actor thinks the pressing problems are. Note: the institute renamed from AI Safety Institute to AI Security Institute; the course still writes "UK AISI", which is the same body.
- UK AISI's Alignment Team: Research Agenda — Benjamin Hilton, Jacob Pfau, Marie_DB, Geoffrey Irving (May 2025) · the narrower, more technical companion: safety-case methodology, honesty with asymptotic guarantees, debate and scalable oversight, and automated alignment research, each with open problems attached. Who it's for: theory-inclined readers who want problems with crisp statements.
- Recommendations for Technical AI Safety Research Directions — Anthropic Alignment Science (2025) · open problems collected from the team, grouped into categories. The post describes itself as a tasting menu rather than a roadmap, and explicitly is not a list of directions Anthropic is staffing — which makes it a good source of unclaimed work. Who it's for: anyone who wants lab-flavoured problems without joining a lab.
Mentored research programs
- MATS — an independent research fellowship pairing scholars with mentors in alignment, transparency and security; 12 weeks in Berkeley and London with the option to apply for a funded 6–12 month extension. Who it's for: the default answer for "how do I get supervised research experience"; broad enough to fit most backgrounds. Winter 2027 applications listed as open.
- Pivotal Research Fellowship — 15 weeks in London at the London Initiative for Safe AI, with weekly 1:1s, dedicated research managers, and a £6,000–£8,000 stipend plus travel, housing, meals and compute. Who it's for: people who want heavy research-management support rather than sink-or-swim autonomy. Listed dates: 18 January – 30 April 2027; the site takes expressions of interest between rounds.
- LASR Labs — a 13-week London programme where teams of three or four, supervised by an experienced researcher, take one project from proposal to publication; £15,000 stipend, food, office and travel covered. Week 0 is spent evaluating candidate projects, which is unusually good training in research taste. Who it's for: people who want a paper out the other end — LASR work has landed at NeurIPS, ICLR and ICML. Winter 2027 cohort runs 11 January – 9 April; the site lists a deadline of 20 September, 23:59 GMT.
- PIBBSS Fellowship — Principles of Intelligence · a three-month interdisciplinary program pairing roughly twenty fellows from the complex-systems sciences with AI safety mentors, on projects at the intersection of the two. Who it's for: physicists, biologists, economists and others whose field has structure worth importing. The course writes "PIBBS"; the fellowship is PIBBSS, and the parent org now trades as Principles of Intelligence — hence the
princint.aidomain. - Anthropic Fellows Program (launch post, 2024) — the announcement the course links: funding plus Anthropic mentorship for a small cohort working full-time on safety research. Who it's for: strong engineers or researchers wanting a direct line into a frontier lab's safety team. This post is now superseded — it carries a July 2026 update pointing at the current cohort announcement, which reports that over 80% of the first cohort produced papers and over 40% joined Anthropic full-time. Check the newer post for live dates; applications for the November 2026 cohort closed on 26 July.
Self-serve upskilling and short sprints
- ARENA — the Alignment Research Engineer Accelerator: a four-to-five week in-person bootcamp at LISA in London, run two or three times a year, covering the engineering skills safety research actually uses. Who it's for: people who can code but have not built transformers, probes or eval harnesses from scratch. ARENA 9.0 runs 5 October – 6 November 2026 with applications closed; the curriculum is free online, so "work through ARENA independently" is a real option, not a consolation prize.
- Technical AI Safety Project Sprint — BlueDot Impact · a 30-hour project-based course: pick a paper to extend or research code to improve, get weekly check-ins with an expert and a peer group, publish a write-up at the end. Past participants have reproduced METR findings, fixed TransformerLens issues and replicated evals in Inspect. Who it's for: graduates of this course who want one public artefact rather than another curriculum. Listed deadline: apply by 27 September.
- Apart Sprints — Apart Research · monthly weekend hackathons, online and in person, with code starters, live mentorship on Discord, and a path into the Apart Lab Fellowship for teams whose work is worth continuing. Who it's for: the lowest-commitment way to produce something real — a weekend, not a quarter. Upcoming as listed: AI Incident Response Sprint, 11–13 September 2026; AI Collusion Research Sprint, 23–25 October 2026.
Engineering: the tools to contribute to
- Inspect — UK AISI and Meridian Labs · an open-source Python framework for LLM evaluations, with composable datasets, solvers, tools and scorers, a CLI, a log viewer and a VS Code extension, plus more than 200 pre-built benchmark implementations. Who it's for: the reference example of engineering leverage in this field, and a realistic first open-source contribution target.
- Petri — Anthropic, released 6 October 2025 · the Parallel Exploration Tool for Risky Interactions. You describe in natural language what behaviour you want to probe; Petri spins up auditor agents that run multi-turn conversations against the target with simulated users and tools, then uses judge models to surface the concerning transcripts. Who it's for: anyone who wants to run behavioural audits without hand-writing every scenario.
Founding
- Incubator Week — BlueDot Impact · five days in San Francisco, expenses paid, ending in a Friday pitch; up to $100k in grants if they back you. The week runs threat models → build and test interventions → pitch prep → pitch. Past cohorts total 38 participants and $2.5M+ raised externally. Who it's for: people with a specific thing they want to exist and no org to build it in. Cohort v5 is listed as running 24–28 August. Apply through the linked application form — it is a JavaScript-rendered form that shows an empty page to crawlers, so open it in a real browser.
- Seldon Lab — a program and investor building a portfolio focused on AI security, organised around four theses: demonstrations, detection and response, resilience, and hardware. Batch #2 was announced in December 2025 and the site publishes a request for startups. Who it's for: founders whose angle is security — model theft, export controls, hardware-level governance — rather than alignment research.
- 5050 AI RFS — Fifty Years · a program that turns scientists, researchers and engineers into founders (78 companies to date), with an AI track and five written-up startup ideas it would fund now, starting with scalable oversight infrastructure for multi-agent systems. Who it's for: technical people who want a concrete brief rather than a blank page. The page states applications for the current cohort are closed with the next cohort starting in March 2026 — a date already past as of September 2026, so treat the interest form and the listed contact address as the live route.
Community, and the courses next door
- BlueDot Impact events calendar — the events the chapter means when it says "go to AI safety events and talk to people". Cheapest possible action for heuristic 2. Who it's for: everyone; especially people who want collaborators rather than a job.
- Biosecurity course — BlueDot Impact · named in the chapter's closing note as an example of the "societal defences" layer that technical safety work does not cover. Who it's for: readers who accept the chapter's argument that alignment alone is not sufficient.
Exercises
This chapter ships no exercises — it is the unit's overview, and the graded work lives in 6.2 (choose your focus) and 6.3 (the 1-pager). The three below are field map extras, built from the chapter's own "this week you could" prompts so that the ideas turn into artefacts.
- The neglectedness audit (field map extra) — Open three of the open-problem lists above (Redwood, Anthropic's recommended directions, the UK AISI alignment agenda). Pull out every problem that you could state clearly to a peer, and for each, spend five minutes searching for existing work on it. Sort the result into three buckets: crowded, tried-and-stalled, and apparently untouched. What a good answer has: at least fifteen problems triaged, with a URL as evidence for each "crowded" verdict rather than an impression; a written note on why the untouched ones might be untouched (genuinely neglected, or quietly known to be dead ends); and one problem you would actually pick, with the reason stated in one sentence.
- Galaxy-brain audit of your own plan (field map extra) — Write down the career step you are currently most drawn to, then write the argument for it in explicit numbered steps. Now attempt to construct an argument of the same shape, using steps of equal plausibility, that concludes you should do the opposite. What a good answer has: both chains written out, an honest verdict on whether the second one was easy to build, and — if it was — a replacement plan whose case survives being stated in two sentences with no clever steps in it. Buterin's test is the standard: an argument that can justify anything implies nothing.
- Ship one artefact this week code (field map extra) — Convert the reading into one public, runnable thing: a small replication or a small eval. Pick a result from any unit of this course that a laptop can touch, reproduce it, and publish the notebook plus a short honest write-up including what failed. What a good answer has: a repo or gist anyone can run, a stated hypothesis fixed before you ran it, at least one negative or surprising result reported rather than buried, and a paragraph on what you would do with ten times the compute. Start here: (1) pick the target — a behavioural claim from unit 3 or an interpretability claim from unit 4 — and write the hypothesis down first; (2) choose a model that fits free Colab or a laptop:
Qwen2.5-0.5B/1.5B-Instruct,Llama-3.2-1B-Instruct, orgpt2-smallif you need a model that interpretability tooling knows well; (3) install the stack —transformersandtorchfor behaviour,transformer_lensornnsightfor internals,inspect_aiif you are writing an eval rather than a probe; (4) build the smallest dataset that could show the effect, 50–200 prompts, and hold out a control set that should not show it — this is the step that separates a result from an artefact of your prompt wording; (5) run it, plot it, and check whether the control behaves as predicted; (6) publish the repo and post the write-up somewhere public, then send it to one person who will tell you if it is wrong.
Go deeper
- inspect_ai on GitHub — the actual codebase behind Inspect. Reading its solver and scorer abstractions is the fastest way to understand how modern evals are structured, and its issue tracker is a genuine on-ramp for the "pick an issue on an open-source AI safety tool" suggestion.
- safety-research/petri on GitHub — Petri's source. Worth reading even if you never run it: the auditor/target/judge decomposition is a reusable pattern for automated behavioural testing.
- ARENA_3.0 curriculum on GitHub — the full ARENA material, free and self-servable. This is what makes "work through ARENA independently" a real plan rather than a fallback for people who did not get in.
- Anthropic Fellows Program: current cohort announcement — the live version of the program the course links via its 2024 launch post, with the current research areas (scalable oversight, adversarial robustness and AI control, model organisms, mechanistic interpretability, AI security, model welfare) and outcome statistics.