Option 1: Anthropic
Unit 3's earlier chapters covered the technical half of detecting danger — evals, red-teaming, what a dangerous-capability test measures. These four chapters cover the institutional half, because an evaluation is only useful if someone committed in advance to what its result would change. Every frontier lab has published a document attempting that: Anthropic's RSP, OpenAI's Preparedness Framework, Google DeepMind's Frontier Safety Framework, Meta's Frontier AI Framework. Pick one, read it as an auditor rather than a fan, answer the same five questions. This is the Anthropic option.
What a frontier safety framework is trying to buy
Nobody can say today what safeguards a 2028 model will need, so every framework uses the same move: conditional commitment. Not "we will do X," but "if our models reach capability level C, mitigations M will be in place before we train or deploy further." That settles the argument early, when nobody has a shipping deadline riding on the answer. It buys a forcing function (safety work with no deadline gets deprioritised forever; work that gates a launch gets staffed), an auditable artifact (checking whether a lab did what it said is far easier than checking whether a model is safe), and a template regulators can adopt. The cost: it is voluntary and self-scored. The lab writes the thresholds, runs the evaluations deciding whether they've been crossed, and grades its own compliance. Everything interesting about auditing an RSP lives in how hard the document makes it to quietly say "not yet."
What version 3 changed, and why it matters more than it sounds
If you learned about the RSP before 2026, you learned about AI Safety Levels: ASL-2 for current models, ASL-3 for models that meaningfully uplift a non-expert building chemical or biological weapons, ASL-4 above that — each with a list of required deployment and security controls. Anthropic activated ASL-3 in May 2025 and has shipped frontier models under it since.
Version 3.0, effective 24 February 2026, retired that ladder as a forward-looking device. ASLs survive only as shorthand for mitigations currently in force; they no longer define what future models require. In their place sit four capability or usage thresholds stated in prose, each specifying what kind of argument a developer should be able to make rather than which controls to install.
"…when defining the risk mitigations needed for future levels of AI capability, we have found that providing a specific list of controls is overly rigid, and we instead prefer to focus on what sort of argument an AI developer should make (and what sorts of actors it should address) regarding the risk level from its systems."— Responsible Scaling Policy v3.4, Anthropic (2026), Appendix B
Two more changes came with v3. Each threshold splits into what Anthropic plans as a company and what it recommends the industry do — a concession to a collective-action problem, on the reasoning that one lab pausing while others don't makes the world less safe, not more. And it added two recurring artifacts: a Frontier Safety Roadmap of public, gradeable goals, and Risk Reports every three to six months stating an overall risk judgement across the deployed fleet, shared unredacted with at least 200 employees and sent to independent external reviewers asked to publish commentary within 30 days.
The policy has moved fast since — v3.1 clarified the AI R&D threshold, v3.2 let the Long-Term Benefit Trust demand external review and approve who conducts it, v3.3 and v3.4 revised the bioweapons and automated-R&D thresholds. As of September 2026 the live document is version 3.4, effective 8 July 2026. Audit the current PDF and its redline, not a summary — including this one.
The four thresholds, in plain terms
- Non-novel chemical/biological weapons production. The model meaningfully helps someone with an undergraduate STEM background build and deploy a known weapon. This is what ASL-3 protections were built for: classifier guards, vetted-user access controls, jailbreak bug bounties, threat intelligence.
- Novel chemical/biological weapons production. The model substitutes for the world-leading specialists who are currently the bottleneck on new agents, end to end. The recommended bar is security "roughly in line with RAND SL4" — resistant to well-resourced state actors, not just opportunistic theft.
- Misaligned AI systems in high-stakes settings. Systems with sensitive access and enough autonomy that sabotage becomes plausible. The policy is candid that this bites first inside AI companies, where models help train their successors and help produce the evidence used to judge them.
- Automated R&D in key domains. Systems that fully automate or dramatically accelerate top-tier research teams. v3.4 operationalises it as either (1) substituting for Anthropic's entire population of Research Scientists and Engineers at within 5× the cost, or (2) "dramatic acceleration" — double the historical rate of progress in aggregate AI capabilities, plausibly attributable to automated R&D. The footnote does the work: doubling the rate means as much progress in one year as two years at baseline, so a 9× baseline effective scaleup becomes roughly 81× — explicitly not "researchers are twice as productive."
Note what is missing: there is no formal cyber threshold, in any version. That becomes a live problem two paragraphs from now.
The policy meeting a real model
Section 8 of the Claude Opus 4.6 system card (February 2026) is the honest part of the exercise: the RSP applied under time pressure, still in the older ASL vocabulary because it shipped days before v3.0 took effect. The determination is that Opus 4.6 crosses neither the AI R&D-4 nor the CBRN-4 threshold, and it deploys under ASL-3. Read how that rule-out is reached. On autonomy, the model had "roughly reached the pre-defined thresholds we set for straightforward ASL-4 rule-out based on benchmark tasks" — so the rule-out rests instead on qualitative impressions plus a survey of 16 internal staff, none of whom believed it could fully automate an entry-level remote researcher. On cyber, Opus 4.6 saturated every evaluation Anthropic had — roughly 100% on Cybench (pass@30), 66% on CyberGym (pass@1) — in the one domain where the RSP defines no threshold at all. The framework ran out of instrument before it ran out of capability. Section 1.2.4.4 adds the second-order version: the team used Opus 4.6 via Claude Code to debug the infrastructure evaluating Opus 4.6.
Anthropic's response is the pattern to look for anywhere: rather than keep leaning on a subjective rule-out, they committed at the Opus 4.5 launch to write a sabotage risk report meeting the higher standard for every subsequent frontier model, threshold crossed or not. When a test stops discriminating, the fix is to owe the work unconditionally.
Mythos Preview: the framework under load
The third reading, the Claude Mythos Preview system card (7 April 2026), matters for one reason: it documents a model Anthropic did not ship. Mythos Preview could autonomously find and exploit zero-days in major operating systems and browsers, so instead of a release it went to a restricted set of security partners under Project Glasswing. Whatever else you conclude, that is a framework producing a decision that costs money. It is also the first system card under RSP v3, and narrates the shift: "Release decision process" becomes "RSP risk assessment process," rule-in/rule-out language is de-emphasised, and the model is judged against the four prose thresholds with an overall risk assessment instead of an ASL label. The conclusion: catastrophic risks remain low — held, on automated R&D, "with less confidence than for any prior model." Then it says the quiet part:
"We find it alarming that the world looks on track to proceed rapidly to developing superhuman systems without stronger mechanisms in place for ensuring adequate safety across the industry as a whole."— System Card: Claude Mythos Preview, Anthropic (2026), §1.2.2
Where it gives
Voluntary and self-graded. The RSP is not a regulatory filing — Anthropic keeps a separate Frontier Compliance Framework for statutory obligations like California SB 53. Nothing here is enforceable from outside. Accountability is internal-plus-adjacent: the Responsible Scaling Officer, the Board, the Long-Term Benefit Trust, an anonymous noncompliance channel with anti-retaliation protection, no non-disparagement clauses that muzzle safety concerns, and an annual third-party review of procedural compliance — explicitly not of outcomes.
Prose thresholds are harder to falsify than checklists. Trading a control list for "make a strong argument" buys adaptability and pays in legibility. Under the ASL ladder you asked whether a control existed; under v3 you must judge whether an argument is good — exactly the judgement Anthropic says is getting more subjective.
The pause is conditional on competitors. Appendix A makes the delay commitment contingent: Anthropic holds back when it has a clear lead, or when competitors demonstrably hold a strong safety posture it must match. Defensible, given a real collective-action problem — and also the escape hatch. Ask what evidence triggers it, and who judges that evidence.
Frequent amendment cuts both ways. Five versions in five months is a living document absorbing operational learning — and five chances to move a threshold about to bind. The mitigation is procedural: changes need Board approval in consultation with the LTBT, and ship with a redline and changelog. Read the redlines.
Readings, linked
The course budgets one hour and recommends 45 minutes reading plus 15 minutes writing. Start with the RSP itself — read Section 1 (the threshold table) and Appendix B, which is where the v3 rewrite is explained — then jump straight to Section 8 of the Opus 4.6 card to see the same framework applied to a shipping model.
- Anthropic's Responsible Scaling Policy — Anthropic (2026) · ~30 min · The framework itself. The course links the v3.0 announcement post (24 Feb 2026), which explains the rewrite's reasoning; the live policy is now v3.4, effective 8 July 2026, and the hub page carries every prior version plus redlines. Read the current PDF, not just the announcement.
- System Card: Claude Opus 4.6 — Anthropic (February 2026) · ~20 min for §8 · The course says skim Section 8, which is "RSP evaluations": process, CBRN, autonomy, cyber, third-party assessments. Also read §1.2.4 ("Conclusions"), four pages that state the rule-outs and admit how thin they are getting. The course's link is the direct CDN PDF (~14 MB); the anthropic.com URL above is the stable redirect.
- System Card: Claude Mythos Preview — Anthropic (7 April 2026) · ~25 min for §1–2 · The first system card published under RSP v3, and the first published for a model deliberately not made generally available. §2.1.1 narrates the v2→v3 vocabulary shift; §1.2.2 gives the per-threshold determinations. The course links a Google Drive copy; the anthropic.com link above is the canonical one.
- Optional — Threat Intelligence Report: August 2025 — Anthropic (2025) · ~15 min · Real misuse Anthropic caught and disrupted: a large-scale extortion operation run through Claude Code, a North Korean fraudulent-employment scheme, AI-generated ransomware sold by someone with only basic coding skills. Useful ballast — it shows what the detection layer catches below the catastrophic thresholds. Full PDF.
- Optional — The Long-Term Benefit Trust — Anthropic (2023, with later trustee updates) · ~8 min · The governance backstop the RSP leans on. Five financially disinterested trustees hold a special class of stock that lets them elect and remove a growing share of the Board. RSP v3.2 gave the Trust the power to demand external review of a Risk Report and to approve who conducts it — that is the concrete link between this page and the accountability question in the exercise.
Note: the course's resource block lists five items (three core, two optional) though the page data carries six records; if a sixth resource appears for you, it is not rendered in the public view.
Exercises
- Safety testing — The course's exercise for all four lab options is the same five-question audit, applied to your chosen company and one dangerous capability. Pick a capability first (bioweapons uplift, autonomous cyber operations, or automated AI R&D are the tractable ones for Anthropic) and then answer, with citations to the specific section you're relying on:
1. Limits. Which concrete observations about this capability would indicate that it is, or strongly might be, unsafe to keep scaling? Quote the threshold text and then say what a passing measurement would actually look like.
2. Protections. Which parts of the current protective measures are load-bearing for containing catastrophic risk from this capability — not the full list, the ones whose removal would break containment?
3. Evaluation. What are the procedures for catching early warning signs before the limit is crossed? Cadence, who runs them, what happens when they saturate.
4. Response. If the capability goes past the limit and protections can't be improved quickly, is the developer actually prepared to pause capability work and treat existing dangerous models with caution? Who has the authority to order that, and what is it conditional on?
5. Accountability. How does the developer ensure the commitments are executed as intended; how can outsiders verify it or notice failure; where are the openings for third-party critique; and what stops the framework itself from being changed in a rushed or opaque way?
What a good answer has: a specific section citation for every claim, and at least one place where you say "the document does not answer this." For Anthropic the sharp findings are available: on (1) the automated-R&D threshold has a numeric operationalisation (double the historical rate of progress, ~81× effective scaleup on the paper's own worked example) while cyber has no threshold at all; on (3) Opus 4.6 saturated the entire cyber evaluation suite and hit the ceiling of the autonomy rule-out benchmarks, forcing a fallback to a 16-person internal survey; on (4) the pause commitment in Appendix A is conditional on competitor behaviour rather than unconditional; on (5) the stack is RSO → Board → Long-Term Benefit Trust, plus published Risk Reports with independent external reviewers on a 30-day public-comment clock, an anonymous noncompliance channel, and an annual third-party review that checks procedure and explicitly not outcomes. A weak answer paraphrases the marketing summary; a strong one names what would have to be true for the framework to fail silently. - Threshold-crossing tabletop code — field map extra. The course exercise is written; this one makes the "when does the evaluation stop discriminating" problem concrete. Build a tiny saturation tracker: take a public capability benchmark with several model generations of scores, fit the trend, and compute the date at which the benchmark ceiling is reached — the date after which that benchmark can no longer tell you anything about a new model.
What a good answer has: a chart with the ceiling drawn as a horizontal line, an explicit saturation date with error bars, and one paragraph on what an RSP-style framework should be required to do at that date. Bonus: repeat for two benchmarks with different ceilings and show that the framework's coverage has holes at different times in different domains.
Start here: (1) Pull scores from Epoch AI's benchmarking dashboard or straight out of the Opus 4.6 and Mythos Preview system cards — Cybench and CyberGym are the obvious pair, since the Opus 4.6 card already reports ~100% and 66%. (2) Put them in a pandas DataFrame keyed by release date. (3) Fit a logistic curve withscipy.optimize.curve_fit, not a line — capability curves against a capped metric are sigmoid, and a linear fit will lie to you about the ceiling. (4) Solve for the date at which the fit reaches 95% of max and bootstrap a confidence interval by resampling the points. (5) Plot with matplotlib; annotate each model. (6) Write the paragraph: on your saturation date, what should the RSP have committed to — a harder benchmark, a mandatory uplift trial, an unconditional risk report, a pause? No GPU needed; this runs on a laptop in a free Colab.
Go deeper
- Redacted Risk Report, August 2026 — the v3 artifact in the flesh, covering risks through a 15 July coverage date. Read it next to the February 2026 report to see what changed in six months. This is the document the external reviewers actually review.
- Frontier Safety Roadmap — the public, gradeable goals across security, alignment, safeguards and policy. Anthropic commits to not quietly weakening goals it can't hit, which makes this the easiest place to check whether the forcing function is working.
- Assessing Claude Mythos Preview's cybersecurity capabilities (Frontier Red Team) — the technical write-up behind the decision not to ship. Concrete detail on autonomous zero-day discovery, which is the best available answer to "what does an undefined threshold being crossed look like?"
- Announcing Anthropic's Responsible Scaling Policy (September 2023) — v1.0's original framing. Worth 10 minutes purely as a diff against v3: what the authors thought they could specify in advance, and what three years of contact with frontier models made them give up on.
- METR — Frontier AI Safety Policies — a neutral index of every published framework (Anthropic, OpenAI, Google DeepMind, Meta, Microsoft, Amazon, xAI, Nvidia and more), plus METR's "Common Elements of Frontier AI Safety Policies" and the Frontier Model Forum's components write-up. Exactly the muscle the next three chapters build: read one framework closely, then look at the shape of the whole field.