TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 6 · START CONTRIBUTINGchapter 4 · reading

Next steps: Apply to roles

BlueDot Impact · Technical AI Safety · unit 6, chapter 4
TL;DR — The question this chapter answers is "where do I actually send an application?" The answer is a list of twelve organisations that hire technical people today, spanning nonprofits, VC-backed startups, an insurance company and a government institute. The thing to remember is the chapter's one argument: these orgs are not screening for years of AI safety experience — almost nobody has that — they are screening for someone who can own a workstream from day one and who obviously cares about the risk. If you have shipped hard things somewhere else, the missing piece is context, and context is cheaper to acquire than a career.

Unit 6 has walked you from "I finished a course" to "I have a focus and a one-pager." This chapter is where the abstraction ends and URLs begin. It exists because the most common failure mode at this exact point is not rejection — it is never applying, on the theory that six more months of reading will make you legible. The chapter's implicit claim is that the field is small enough and young enough that legibility comes from doing, and the fastest way to do is to get hired by one of the roughly two dozen organisations that already have the funding and the problems.

Why "not ready" is usually the wrong diagnosis

The single link in the chapter's body text is a survey BlueDot ran across more than ten safety organisations, asking what they were actually short of. The answer was not "people with alignment PhDs." It was people who can take a poorly-specified problem, own it end to end, and not need a manager to convert it into tickets — plus genuine, non-performative concern about the risk. That combination is rare in any field. It is not, however, a thing you acquire by reading more papers.

"Given that relatively few people can claim 5+ years of direct AI safety experience, orgs are looking for the next best thing"— We asked 10+ AI safety orgs about their hiring needs, Li-Lian Ang (2026)

The corollary is the useful part. If the scarce thing is ownership rather than domain tenure, then your five years of distributed-systems work, or your track record of shipping production ML, or your experience running a government programme, is not a handicap you have to apologise for in a cover letter. It is the qualification. What you are missing is the field's vocabulary and its open questions — the survey names context as the real barrier:

"The biggest barrier for capable talent entering the field is context"— We asked 10+ AI safety orgs about their hiring needs, Li-Lian Ang (2026)

Context is what a course, a weekend replication, and a careful reading of one org's last three papers buys you. It is measured in weeks. Deciding to wait for it costs you a hiring cycle.

The field is not one kind of employer

Reading the twelve organisations below as a flat list hides the most decision-relevant thing about them: they are funded differently, and funding shapes what the job feels like day to day. Five rough shapes, and the list maps onto them cleanly.

Philanthropically funded nonprofits — METR, Redwood, Epoch, FAR.AI, CivAI, MATS, AI Digest. Freedom to work on things with no buyer; the flip side is that the work is only as durable as the next grant cycle, and headcount moves in steps rather than smoothly. Compensation at the larger ones is nonetheless competitive with industry: METR publishes technical-staff bands starting in the low-to-mid six figures.

Venture-backed companies whose product is safety research — Goodfire, and to a different degree Apollo, which sells a security product alongside its scheming research. The bet here is commercial-incentive alignment: if customers pay for interpretability, the company gets to do more interpretability. Worth stress-testing in interviews — ask what happens to the research agenda when a large customer wants something adjacent to it.

Liability and assurance plays — AIUC. This is the least AI-safety-shaped org on the list and arguably the most interesting structural bet: if agents cannot be deployed at scale until someone underwrites the downside, then the underwriter's audit standard becomes a de facto safety regulation with a market behind it.

Government — the UK AI Security Institute. Different currency entirely: statutory access to frontier models pre-deployment, and a research output (Inspect) that half the field now runs its evals on. You trade equity upside and speed for reach.

Profit-subsidised research — AE Studio, which funds an in-house alignment team from consulting revenue. The unusual property is that the research team is not fundraising, so it can chase directions nobody else will fund.

What a strong application looks like from the other side of the table

Most of these teams are between 3 and 50 people. There is no requisition machinery, no keyword screen, and frequently the person reading your application is the person you would report to. Three practical consequences.

First, specificity beats enthusiasm. "I care about AI risk" is the baseline, not a differentiator — every applicant writes it. "Your time-horizons methodology assumes X, and I think the assumption breaks for agentic coding tasks because Y" is a differentiator, and it takes an afternoon.

Second, artefacts beat claims. A small public replication, a plot, a short writeup, a merged PR against an open eval harness — anything that lets a reader verify your reasoning without trusting you. Unit 6's one-pager exercise is the minimum version of this; a repo is the better one.

Third, apply broadly and in parallel. The variance between orgs on what they select for is enormous — Redwood is hiring for one role at extreme seniority, Apollo has fourteen open across governance, security engineering and control research, FAR.AI is hiring red-teamers remotely and researchers on-site. Self-rejecting from the whole list because you do not fit the first entry is a category error.

Carry this away: the field's hiring bottleneck is ownership plus context, in that order. You cannot fake ownership and you do not need to fake context — you need one concrete artefact showing you have engaged with a specific org's actual open problem. Build that artefact first, then send it to five places in the same week.

Readings, linked

The course budgets no fixed reading time here — it is a directory, and the intended use is to skim all twelve, pick three, and apply. Read the hiring-needs survey first; it calibrates everything below. Org descriptions are the course's; role counts, locations and salary bands were checked against each employer's live job board in September 2026 and are marked as such, because these change weekly.

Exercises

The course sets no exercises on this page — it is a resource directory, and the intended action is simply to apply. The three below are field map extras, built so that finishing them leaves you with something you can attach to an application.

  1. The five-column shortlist (field map extra) — Take all twelve organisations above and put them in a table with five columns: what they'd hire me to do, funding model (grant / VC / revenue / government), location constraint, what I'd have to learn in month one, and one specific thing I disagree with or don't understand about their work. Fill the last column from their own writing, not from summaries. Then cut to three and apply this week. What a good answer has: a last column that is specific enough to be wrong — "I don't see how METR's time-horizon metric handles tasks where the bottleneck is tool latency rather than reasoning" beats "I'd like to understand their methodology better." That column is your cover letter's opening paragraph, and it is the part a hiring manager cannot skim past.
  2. Ship one verifiable artefact code (field map extra) — Pick one org from your shortlist and produce a small public artefact aimed at its actual problem: a mini-eval, a probe, a replicated plot, a red-team transcript. Target a weekend, not a quarter. What a good answer has: a public repo with a README stating the question, the method, the result, and — critically — what the result does not show; a plot or table someone can read in ten seconds; and honest limitations. Reviewers trust a small clean negative result far more than a sprawling positive one. Start here: (1) install Inspect (pip install inspect-ai) — it is the framework AISI built and much of the field uses, so the code itself is a signal; (2) pick a narrow capability question you can score automatically, e.g. "does the model follow an injected instruction embedded in a tool result?"; (3) write 30–50 task samples by hand or generate and then hand-check them; (4) run against a small open model you can host on a laptop or free Colab — Qwen3-4B or Llama-3.2-3B via transformers or ollama — plus one frontier API model if you have credits, so you have a contrast; (5) plot the two, write 300 words on why they differ; (6) publish and link it in the application. If interpretability is your target instead, swap Inspect for TransformerLens and replicate a single figure from a Goodfire or Anthropic paper.
  3. The rejection budget (field map extra) — Decide, in writing and before you send anything, how many rejections you will collect before you change strategy. Then send that many applications. What a good answer has: a number greater than one and a written trigger — "after eight rejections with no first-round interview, the problem is my artefact, not my targeting; after eight first-rounds with no offer, it is the interview." This exercise exists because the failure this chapter is written against is not being rejected; it is applying once, treating a single "no" as evidence about the field, and quietly stopping.

Go deeper

Next: Next steps: Technical fellowships · Back to the map.