TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 3 · DETECTING DANGERchapter 3 · 1h

Option 1: Anthropic

BlueDot Impact · Technical AI Safety · unit 3, chapter 3
TL;DR — This is the first of four "pick a lab and audit its framework" chapters. Anthropic's version is the Responsible Scaling Policy (RSP): name the capability levels that would make continued scaling dangerous, commit in advance to the mitigations each demands, and commit to slowing down if they aren't ready. Version 3, February 2026, tore out the thing everyone remembers the RSP for — the ASL ladder of pre-specified controls — and replaced it with four prose capability or usage thresholds, a public Frontier Safety Roadmap, and periodic Risk Reports reviewed by independent outsiders. Carry this: an if-then framework is only as strong as the evaluation deciding whether the "if" has fired, and Anthropic's own documents now say those calls are getting more subjective, not less.

Unit 3's earlier chapters covered the technical half of detecting danger — evals, red-teaming, what a dangerous-capability test measures. These four chapters cover the institutional half, because an evaluation is only useful if someone committed in advance to what its result would change. Every frontier lab has published a document attempting that: Anthropic's RSP, OpenAI's Preparedness Framework, Google DeepMind's Frontier Safety Framework, Meta's Frontier AI Framework. Pick one, read it as an auditor rather than a fan, answer the same five questions. This is the Anthropic option.

What a frontier safety framework is trying to buy

Nobody can say today what safeguards a 2028 model will need, so every framework uses the same move: conditional commitment. Not "we will do X," but "if our models reach capability level C, mitigations M will be in place before we train or deploy further." That settles the argument early, when nobody has a shipping deadline riding on the answer. It buys a forcing function (safety work with no deadline gets deprioritised forever; work that gates a launch gets staffed), an auditable artifact (checking whether a lab did what it said is far easier than checking whether a model is safe), and a template regulators can adopt. The cost: it is voluntary and self-scored. The lab writes the thresholds, runs the evaluations deciding whether they've been crossed, and grades its own compliance. Everything interesting about auditing an RSP lives in how hard the document makes it to quietly say "not yet."

What version 3 changed, and why it matters more than it sounds

If you learned about the RSP before 2026, you learned about AI Safety Levels: ASL-2 for current models, ASL-3 for models that meaningfully uplift a non-expert building chemical or biological weapons, ASL-4 above that — each with a list of required deployment and security controls. Anthropic activated ASL-3 in May 2025 and has shipped frontier models under it since.

Version 3.0, effective 24 February 2026, retired that ladder as a forward-looking device. ASLs survive only as shorthand for mitigations currently in force; they no longer define what future models require. In their place sit four capability or usage thresholds stated in prose, each specifying what kind of argument a developer should be able to make rather than which controls to install.

"…when defining the risk mitigations needed for future levels of AI capability, we have found that providing a specific list of controls is overly rigid, and we instead prefer to focus on what sort of argument an AI developer should make (and what sorts of actors it should address) regarding the risk level from its systems."— Responsible Scaling Policy v3.4, Anthropic (2026), Appendix B

Two more changes came with v3. Each threshold splits into what Anthropic plans as a company and what it recommends the industry do — a concession to a collective-action problem, on the reasoning that one lab pausing while others don't makes the world less safe, not more. And it added two recurring artifacts: a Frontier Safety Roadmap of public, gradeable goals, and Risk Reports every three to six months stating an overall risk judgement across the deployed fleet, shared unredacted with at least 200 employees and sent to independent external reviewers asked to publish commentary within 30 days.

The policy has moved fast since — v3.1 clarified the AI R&D threshold, v3.2 let the Long-Term Benefit Trust demand external review and approve who conducts it, v3.3 and v3.4 revised the bioweapons and automated-R&D thresholds. As of September 2026 the live document is version 3.4, effective 8 July 2026. Audit the current PDF and its redline, not a summary — including this one.

The four thresholds, in plain terms

Note what is missing: there is no formal cyber threshold, in any version. That becomes a live problem two paragraphs from now.

The policy meeting a real model

Section 8 of the Claude Opus 4.6 system card (February 2026) is the honest part of the exercise: the RSP applied under time pressure, still in the older ASL vocabulary because it shipped days before v3.0 took effect. The determination is that Opus 4.6 crosses neither the AI R&D-4 nor the CBRN-4 threshold, and it deploys under ASL-3. Read how that rule-out is reached. On autonomy, the model had "roughly reached the pre-defined thresholds we set for straightforward ASL-4 rule-out based on benchmark tasks" — so the rule-out rests instead on qualitative impressions plus a survey of 16 internal staff, none of whom believed it could fully automate an entry-level remote researcher. On cyber, Opus 4.6 saturated every evaluation Anthropic had — roughly 100% on Cybench (pass@30), 66% on CyberGym (pass@1) — in the one domain where the RSP defines no threshold at all. The framework ran out of instrument before it ran out of capability. Section 1.2.4.4 adds the second-order version: the team used Opus 4.6 via Claude Code to debug the infrastructure evaluating Opus 4.6.

Anthropic's response is the pattern to look for anywhere: rather than keep leaning on a subjective rule-out, they committed at the Opus 4.5 launch to write a sabotage risk report meeting the higher standard for every subsequent frontier model, threshold crossed or not. When a test stops discriminating, the fix is to owe the work unconditionally.

Mythos Preview: the framework under load

The third reading, the Claude Mythos Preview system card (7 April 2026), matters for one reason: it documents a model Anthropic did not ship. Mythos Preview could autonomously find and exploit zero-days in major operating systems and browsers, so instead of a release it went to a restricted set of security partners under Project Glasswing. Whatever else you conclude, that is a framework producing a decision that costs money. It is also the first system card under RSP v3, and narrates the shift: "Release decision process" becomes "RSP risk assessment process," rule-in/rule-out language is de-emphasised, and the model is judged against the four prose thresholds with an overall risk assessment instead of an ASL label. The conclusion: catastrophic risks remain low — held, on automated R&D, "with less confidence than for any prior model." Then it says the quiet part:

"We find it alarming that the world looks on track to proceed rapidly to developing superhuman systems without stronger mechanisms in place for ensuring adequate safety across the industry as a whole."— System Card: Claude Mythos Preview, Anthropic (2026), §1.2.2

Where it gives

Voluntary and self-graded. The RSP is not a regulatory filing — Anthropic keeps a separate Frontier Compliance Framework for statutory obligations like California SB 53. Nothing here is enforceable from outside. Accountability is internal-plus-adjacent: the Responsible Scaling Officer, the Board, the Long-Term Benefit Trust, an anonymous noncompliance channel with anti-retaliation protection, no non-disparagement clauses that muzzle safety concerns, and an annual third-party review of procedural compliance — explicitly not of outcomes.

Prose thresholds are harder to falsify than checklists. Trading a control list for "make a strong argument" buys adaptability and pays in legibility. Under the ASL ladder you asked whether a control existed; under v3 you must judge whether an argument is good — exactly the judgement Anthropic says is getting more subjective.

The pause is conditional on competitors. Appendix A makes the delay commitment contingent: Anthropic holds back when it has a clear lead, or when competitors demonstrably hold a strong safety posture it must match. Defensible, given a real collective-action problem — and also the escape hatch. Ask what evidence triggers it, and who judges that evidence.

Frequent amendment cuts both ways. Five versions in five months is a living document absorbing operational learning — and five chances to move a threshold about to bind. The mitigation is procedural: changes need Board approval in consultation with the LTBT, and ship with a redline and changelog. Read the redlines.

Carry this: auditing any lab's framework, don't start with the threshold prose — start with the evaluation that decides whether a threshold has fired, and ask what happens when it saturates. Opus 4.6 saturated the cyber suite in the one domain with no threshold, and hit the ceiling of the autonomy benchmarks, falling back on a 16-person survey. Threshold language is the part labs write carefully; measurement is the part that quietly stops working.

Readings, linked

The course budgets one hour and recommends 45 minutes reading plus 15 minutes writing. Start with the RSP itself — read Section 1 (the threshold table) and Appendix B, which is where the v3 rewrite is explained — then jump straight to Section 8 of the Opus 4.6 card to see the same framework applied to a shipping model.

Note: the course's resource block lists five items (three core, two optional) though the page data carries six records; if a sixth resource appears for you, it is not rendered in the public view.

Exercises

  1. Safety testing — The course's exercise for all four lab options is the same five-question audit, applied to your chosen company and one dangerous capability. Pick a capability first (bioweapons uplift, autonomous cyber operations, or automated AI R&D are the tractable ones for Anthropic) and then answer, with citations to the specific section you're relying on:

    1. Limits. Which concrete observations about this capability would indicate that it is, or strongly might be, unsafe to keep scaling? Quote the threshold text and then say what a passing measurement would actually look like.
    2. Protections. Which parts of the current protective measures are load-bearing for containing catastrophic risk from this capability — not the full list, the ones whose removal would break containment?
    3. Evaluation. What are the procedures for catching early warning signs before the limit is crossed? Cadence, who runs them, what happens when they saturate.
    4. Response. If the capability goes past the limit and protections can't be improved quickly, is the developer actually prepared to pause capability work and treat existing dangerous models with caution? Who has the authority to order that, and what is it conditional on?
    5. Accountability. How does the developer ensure the commitments are executed as intended; how can outsiders verify it or notice failure; where are the openings for third-party critique; and what stops the framework itself from being changed in a rushed or opaque way?

    What a good answer has: a specific section citation for every claim, and at least one place where you say "the document does not answer this." For Anthropic the sharp findings are available: on (1) the automated-R&D threshold has a numeric operationalisation (double the historical rate of progress, ~81× effective scaleup on the paper's own worked example) while cyber has no threshold at all; on (3) Opus 4.6 saturated the entire cyber evaluation suite and hit the ceiling of the autonomy rule-out benchmarks, forcing a fallback to a 16-person internal survey; on (4) the pause commitment in Appendix A is conditional on competitor behaviour rather than unconditional; on (5) the stack is RSO → Board → Long-Term Benefit Trust, plus published Risk Reports with independent external reviewers on a 30-day public-comment clock, an anonymous noncompliance channel, and an annual third-party review that checks procedure and explicitly not outcomes. A weak answer paraphrases the marketing summary; a strong one names what would have to be true for the framework to fail silently.
  2. Threshold-crossing tabletop code — field map extra. The course exercise is written; this one makes the "when does the evaluation stop discriminating" problem concrete. Build a tiny saturation tracker: take a public capability benchmark with several model generations of scores, fit the trend, and compute the date at which the benchmark ceiling is reached — the date after which that benchmark can no longer tell you anything about a new model.

    What a good answer has: a chart with the ceiling drawn as a horizontal line, an explicit saturation date with error bars, and one paragraph on what an RSP-style framework should be required to do at that date. Bonus: repeat for two benchmarks with different ceilings and show that the framework's coverage has holes at different times in different domains.

    Start here: (1) Pull scores from Epoch AI's benchmarking dashboard or straight out of the Opus 4.6 and Mythos Preview system cards — Cybench and CyberGym are the obvious pair, since the Opus 4.6 card already reports ~100% and 66%. (2) Put them in a pandas DataFrame keyed by release date. (3) Fit a logistic curve with scipy.optimize.curve_fit, not a line — capability curves against a capped metric are sigmoid, and a linear fit will lie to you about the ceiling. (4) Solve for the date at which the fit reaches 95% of max and bootstrap a confidence interval by resampling the points. (5) Plot with matplotlib; annotate each model. (6) Write the paragraph: on your saturation date, what should the RSP have committed to — a harder benchmark, a mandatory uplift trial, an unconditional risk report, a pause? No GPU needed; this runs on a laptop in a free Colab.

Go deeper

Next: Option 2: OpenAI · Back to the map.