TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 3 · DETECTING DANGERchapter 2 · overview

How do AI companies test for safety?

BlueDot Impact · Technical AI Safety · unit 3, chapter 2
TL;DR — Every frontier lab publishes a document saying, in effect, "if our model can do X, we will do Y before shipping it" — Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, DeepMind's Frontier Safety Framework, Meta's Advanced AI Scaling Framework. They agree on the shape (name a capability, set a threshold, run evals, attach safeguards) and differ on everything that gives it teeth: which risks count, how hard the evals try, who checks, and what happens when the answer is bad. Read one like a spec you have to implement, hunting for triggers and escape hatches.

Chapter 1 asked what dangerous capabilities look like in the abstract; this chapter grounds that in the artifacts the industry actually produces. If you want to work on evals, this is the demand side of the market — the frameworks are where a lab writes down which measurements it has promised to care about, and the system cards are where it reports what those measurements returned. Knowing that landscape is what turns "I built an eval" into "I built the eval that closes a stated gap in someone's threshold."

The if-then shape, and why it's the shape

You cannot certify a large model as safe: no proof, no exhaustive test set, and the capability surface keeps moving. So the industry converged on a weaker but tractable commitment — a conditional policy. Instead of "this model is safe," a lab asserts "we measured capability C, it is below threshold T, and here is the safeguard package we would attach if it crossed." METR's 2023 write-up is the cleanest statement of that logic, and the reading to start with: it gives you five slots — limits, protections, evaluation, response, accountability — to check any real framework against.

"An RSP aims to keep protective measures ahead of whatever dangerous capabilities AI systems have."— Responsible Scaling Policies, METR (2023)

The engineering intuition is a tripwire, not a seal of approval. Its value lives entirely in whether the wire sits low enough to fire before the harm, is sensitive enough to fire at all, and is connected to something that actually stops the train.

Four labs, four dialects

Anthropic — Responsible Scaling Policy. Oldest of the four (v1.0, September 2023) and most revised: the changelog shows a full rewrite at v3.0 in February 2026 and four revisions since, with v3.4 effective 8 July 2026. Discrete AI Safety Levels (ASL-2, ASL-3) bundle deployment and security safeguards together; Capability Thresholds in CBRN uplift and AI R&D acceleration decide which level applies. Watch the security half — ASL-3 is as much about who can steal the weights as about what a user can extract from the chat box.

OpenAI — Preparedness Framework. Version 2 collapses the old four-colour scale to two thresholds, High and Critical, over three Tracked Categories: bio/chem, cybersecurity, AI self-improvement. High demands sufficient safeguards before deployment; Critical demands them during development too. Its best invention is the Research Categories tier — an explicit holding pen for risks judged real but not yet measurable.

Google DeepMind — Frontier Safety Framework. Now at v3.1 (17 April 2026), organised around Critical Capability Levels across misuse, ML R&D and misalignment — and the only one of the four with a CCL for harmful manipulation. The v3 line also added Tracked Capability Levels, a lower rung meant to catch a trend before it reaches the critical one: the closest thing in the set to a leading indicator.

Meta — Advanced AI Scaling Framework. Note the rename: what the course calls Meta's "Frontier AI Framework" was superseded on 8 April 2026 by the Advanced AI Scaling Framework v2, which adds Loss of Control alongside cyber and chem-bio. Its structure is the odd one out, and worth the contrast: outcomes-led, decomposing catastrophic outcomes into threat scenarios, then enabling capabilities, then evaluations — the inverse of starting from a benchmark and asking what it implies.

Where the claims get tested — and where they don't

A framework is a promise; a system card is the report against it — the eval suites actually run, the elicitation effort applied (weak elicitation makes any threshold look comfortably far away), and the red-team findings. Three questions separate a serious card from a press release. Was the model tested with scaffolding, tools and fine-tuning, or bare? Did anyone outside the company get pre-deployment access and time to be adversarial? Is there a number you could recompute, or only an adjective?

The weak points are structural. Every threshold is self-set and self-assessed, "sufficient safeguards" is graded by the party that wants to ship, and the same party revises the document — which is why a version history showing both tightenings and loosenings tells you more than the current text. And the whole apparatus targets catastrophic misuse: ordinary harms, and most misalignment, are out of scope by construction.

Carry this away: read a framework by grepping its verbs, not its risk taxonomy. "We will pause development" is a commitment; "we will consider", "as appropriate", "where feasible" are the seams. The gap between a lab's threshold list and its escape-hatch language is the best single predictor of whether an eval you build will change a shipping decision.

A note on the course's first reading

AI Lab Watch scores seven companies on seven axes — risk assessment, scheming-risk prevention, safety research, misuse prevention, security, info sharing, planning. Still the fastest side-by-side view of the axes, but every score is a 2025 snapshot; use a current tracker for present-day claims.

"I stopped maintaining this website in September 2025."— About, AI Lab Watch, Zach Stein-Perlman (2025)

Readings, linked

The course lists two resources and sets no time budget for either. Read METR first — it gives you the checklist; the dashboard is more useful once you know what the columns mean.

Exercises

  1. Audit one company against one capability — Take the dangerous capability you picked in the previous chapter (bio/chem uplift, offensive cyber, autonomous replication, AI R&D acceleration, manipulation) and pick exactly one of the four labs. Then work through that lab's three artifacts in order: the framework, to find the threshold that governs your capability and the safeguards promised above it; the most recent system card for a frontier model, to find what was actually measured against that threshold and by whom; and any red-team or external-evaluation reporting, to see whether an adversary with access agreed. Write it up as a short memo. What a good answer has: the exact threshold language quoted, with a version number and date; the specific eval or benchmark used and whether elicitation included tools/scaffolding/fine-tuning; a named external evaluator or an explicit note that testing was internal-only; and — the part most people skip — one concrete gap stated as a testable claim, e.g. "the threshold covers uplift to a novice but no eval measures uplift to a trained expert." Recommended pairing: one lab from {Anthropic, OpenAI, DeepMind} against Meta, because Meta's outcomes-led structure makes the others' capability-led assumptions visible.
  2. Field map extra — the verb diff — Pull two consecutive versions of the same framework (Anthropic's changelog makes this easiest: v3.0 vs v3.4) and diff the commitment language, not the taxonomy. Classify each change as a tightening, a loosening, or a clarification. What a good answer has: a table of the changes with your classification and a one-line justification each, plus a verdict on whether revisions have net-tightened or net-loosened, and whether any revision immediately preceded a launch.
  3. Field map extra — build a threshold probe code — Pick one threshold you found in exercise 1 and build a small eval that measures it, then report where your own measurement is weakest. What a good answer has: 15–30 tasks with a deterministic grader, a documented refusal-vs-inability distinction, at least two elicitation conditions (bare prompt vs. tool-using agent scaffold) so you can show the elicitation gap, and an honest section on what your eval would miss. Start here: (1) use Inspect from the UK AI Security Institute, or plain pytest plus a provider SDK if you want no framework; (2) choose a capability whose ground truth is checkable — CTF-style cyber tasks or a code-exploit reproduction beat bio, which you should not build tasks for; (3) write tasks at three difficulty tiers so the score is a curve, not a pass/fail; (4) run condition A, single prompt, and condition B, a ReAct-style loop with a shell or Python tool, on a model you can hit cheaply — a small open model on a free Colab GPU, or a hosted mini/flash tier; (5) plot both conditions and note the gap; (6) write the paragraph a system card would need to include your result, and the paragraph a skeptic would write to dismiss it.

Go deeper

Next: Option 1: Anthropic · Back to the map.