TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 6 · START CONTRIBUTINGchapter 1 · reading

Your next steps

BlueDot Impact · Technical AI Safety · unit 6, chapter 1
TL;DR — The course estimates fewer than 2,000 people work full-time on making AI systems safer, and that number is the chapter's whole argument: in a field that small the binding constraint is not credentials but people willing to pick something and start. It offers five heuristics for choosing and three broad doors — research, engineering, founding — then names specific places to knock. The one thing to remember: the doors are the content here, so this page's job is to make every one of them a verified, annotated link.

Units 1 through 5 built a map of the problem: what goes wrong, which techniques exist, where they fail. Unit 6 is the turn from understanding to action, and chapter 1 is its framing — why acting now beats waiting until you feel qualified, and a pointer at everything the remaining seven chapters detail.

The number that does the arguing

Most career advice opens by telling you the field is competitive. This one opens by telling you it is tiny: fewer than 2,000 people worldwide full-time, so for almost any specific safety question the headcount is between zero and a dozen. Three things follow, and they generate the rest of the chapter. Marginal people matter more than marginal credentials — in a field of 200,000 the tenth-best candidate is a rounding error; in a field of 2,000 they are a meaningful share of the effort. Nobody has the map — no accreditation body, no settled curriculum, no consensus on which agenda is right, which is uncomfortable but also means you need no permission to hold an opinion. And neglectedness is legible — you can read the major labs' and funders' open-problem lists in an afternoon and see what nobody has claimed. Try that in machine learning generally.

Five heuristics, and where each one breaks

The chapter distils several well-known essays into five rules. Each is useful, and each has a failure mode worth naming, because "good advice applied without judgment" is how people waste a year.

1. When in doubt, apply. An application costs hours, a rejection costs nothing durable, and committees predict fit worse than they think. Where it breaks: the rule is scoped to reversible actions — a replication is cheap to be wrong about, noisy claims in a policy fight are not, because reputational damage in a small field compounds.

2. Work in public. A write-up gets feedback, substitutes for institutional signal, and is how collaborators find you. Where it breaks: volume without a quality bar builds an anti-portfolio. One careful replication with honest negative results beats twelve hot takes.

3. Don't wait for permission. The instinct that someone smarter is already on your concern is usually wrong, and even when right, teams are short-handed. Where it breaks: "nobody is working on this" sometimes means "this was tried and it doesn't work" — ten minutes of literature search is the cost of finding out.

4. Weigh neglectedness against career capital. Visible paths carry the best career capital precisely because they are visible, which is also why they are crowded. The counterweight is the Pareto-frontier argument: you may not be the best interpretability researcher in the world, but you might be best at the intersection of your unusual background and safety — and intersections are where uncontested problems live.

"If your arguments can justify anything, then your arguments imply nothing."— Galaxy brain resistance, Vitalik Buterin (2025)

5. Prefer galaxy-brain-resistant plans. The sharpest of the five. Long inferential chains concluding "and therefore the highest-impact thing happens to be the thing I already wanted to do" are evidence about the reasoner, not the world; the chapter's canonical example is joining a capabilities team to influence it from inside, an argument constructible for almost any position. Ask whether you could derive the opposite conclusion from equally plausible steps. If yes, you have a rationalisation, not a strategy. Where it breaks: pushed too far it argues against all non-obvious action — the test is whether the reasoning is load-bearing and unfalsifiable, not whether it is unusual.

Three doors, plus the one the chapter adds at the end

Research produces new knowledge about model behaviour: training models, running experiments, or importing structure from other fields — model organisms is a biology idea, red-teaming a security one. Deep ML fluency helps enormously; a non-ML background is not disqualifying if it gives you a niche.

Engineering builds the infrastructure that lets researchers run experiments at all: eval harnesses, auditing tools, sandboxes, data pipelines. It is undersold relative to its leverage — one good framework multiplies every researcher who adopts it, and "evals research" in practice is mostly engineering and operations. If you already ship software, this is the shortest path from your current skills to real contribution. Founding starts the thing when no existing org will house it. Highest variance, and the door where agency substitutes most directly for credentials.

The chapter's closing move is the one most readers skip and shouldn't: research organisations are not made of researchers. They run on operations staff, research managers, recruiters, designers, product people and domain experts, and those roles are as understaffed as the technical ones. "I am not going to out-compete a PhD at interpretability" is not a reason to leave the field — it is a reason to read the org chart.

The boundary the chapter draws

The final note is honest about a limit of the whole course: technical safety work is necessary but not sufficient. A perfect alignment technique only helps if developers are made to use it (governance), if the weights are not stolen (AI security), and if leaked harms hit a society with defences (biosecurity, cybersecurity). Read it as a warning against treating your chosen technique as the whole solution — not as a reason to abandon technical work.

Readings, linked

This chapter assigns no formal reading block; these are the essays it cites inline as the source of its advice. If you read one, read the 80,000 Hours career review — it is the widest survey. If you are choosing between "go independent" and "get mentorship", read Hobbhahn.

"…forming an inside view by going out in the world and doing things - not just by hiding away and thinking really hard."— How I Formed My Own Views About AI Safety, Neel Nanda (2022)

The doors, annotated

Every organisation, list, tool and program this chapter names, with a verified link and what it actually is. Dates and deadlines are as the linked pages state them at the time of writing (September 2026) — always re-check on the page itself.

Open problems you can start on today

Mentored research programs

Self-serve upskilling and short sprints

Engineering: the tools to contribute to

Founding

Community, and the courses next door

Exercises

This chapter ships no exercises — it is the unit's overview, and the graded work lives in 6.2 (choose your focus) and 6.3 (the 1-pager). The three below are field map extras, built from the chapter's own "this week you could" prompts so that the ideas turn into artefacts.

  1. The neglectedness audit (field map extra) — Open three of the open-problem lists above (Redwood, Anthropic's recommended directions, the UK AISI alignment agenda). Pull out every problem that you could state clearly to a peer, and for each, spend five minutes searching for existing work on it. Sort the result into three buckets: crowded, tried-and-stalled, and apparently untouched. What a good answer has: at least fifteen problems triaged, with a URL as evidence for each "crowded" verdict rather than an impression; a written note on why the untouched ones might be untouched (genuinely neglected, or quietly known to be dead ends); and one problem you would actually pick, with the reason stated in one sentence.
  2. Galaxy-brain audit of your own plan (field map extra) — Write down the career step you are currently most drawn to, then write the argument for it in explicit numbered steps. Now attempt to construct an argument of the same shape, using steps of equal plausibility, that concludes you should do the opposite. What a good answer has: both chains written out, an honest verdict on whether the second one was easy to build, and — if it was — a replacement plan whose case survives being stated in two sentences with no clever steps in it. Buterin's test is the standard: an argument that can justify anything implies nothing.
  3. Ship one artefact this week code (field map extra) — Convert the reading into one public, runnable thing: a small replication or a small eval. Pick a result from any unit of this course that a laptop can touch, reproduce it, and publish the notebook plus a short honest write-up including what failed. What a good answer has: a repo or gist anyone can run, a stated hypothesis fixed before you ran it, at least one negative or surprising result reported rather than buried, and a paragraph on what you would do with ten times the compute. Start here: (1) pick the target — a behavioural claim from unit 3 or an interpretability claim from unit 4 — and write the hypothesis down first; (2) choose a model that fits free Colab or a laptop: Qwen2.5-0.5B/1.5B-Instruct, Llama-3.2-1B-Instruct, or gpt2-small if you need a model that interpretability tooling knows well; (3) install the stack — transformers and torch for behaviour, transformer_lens or nnsight for internals, inspect_ai if you are writing an eval rather than a probe; (4) build the smallest dataset that could show the effect, 50–200 prompts, and hold out a control set that should not show it — this is the step that separates a result from an artefact of your prompt wording; (5) run it, plot it, and check whether the control behaves as predicted; (6) publish the repo and post the write-up somewhere public, then send it to one person who will tell you if it is wrong.

Go deeper

Next: Choose your focus · Back to the map.