Option 2: OpenAI
Chapter 3.2 argued that frontier safety frameworks are a genre with a common grammar: capability thresholds, mitigations tied to each threshold, and a commitment to stop if the mitigations aren't ready. Chapters 3.3–3.6 make you read one instance of that genre properly. This is the OpenAI instance. Read it the way you would read an API contract someone is about to depend on in production: what is promised, what is merely described, who can change it, and what happens when the promise and the shipping schedule disagree.
Deciding where to look before deciding what to do
Most of a safety framework's real content is in scoping, because a framework that tracked every possible harm would be unfalsifiable. OpenAI narrows aggressively and says so out loud: the document targets "severe harm," defined in a footnote as the death or grave injury of thousands of people, or hundreds of billions of dollars of economic damage. Everything below that bar — bias, dependency, ordinary abuse, most of a trust-and-safety team's day job — is explicitly a different part of the safety stack.
Inside that bar, a capability becomes a Tracked Category only if it passes five tests: the harm pathway is plausible, the capability is measurable by evaluations that track the harm, the harm is severe by the definition above, the capability is net new relative to what a threat actor could already do with 2021-era tools, and the outcome is instantaneous or irremediable. Three areas currently pass: biological and chemical uplift, cybersecurity, and AI self-improvement.
Two criteria do most of the work. "Net new" keeps ordinary misuse out of scope — if a determined actor could already do it with a search engine and a library card, an AI that speeds it up doesn't count. "Measurable" is self-referential: a risk earns close tracking only if OpenAI can already build an eval for it, so the hardest-to-measure risks are structurally the least likely to be tracked. Version 2 also dropped Persuasion, tracked in the December 2023 beta, as systemic rather than the kind of discrete severe event this framework is built around.
Research Categories: the honest holding pen
The framework's most interesting structural move is admitting a second tier. Five capabilities sit in Research Categories — Long-range Autonomy, Sandbagging, Autonomous Replication and Adaptation, Undermining Safeguards, and Nuclear and Radiological — where OpenAI says the threat models or the measurement science aren't mature enough for a commitment, so it commits instead to threat-modelling, eval development, and publishing what it finds.
Two of these are not risk categories at all; they are attacks on the framework itself. Sandbagging is a model that performs differently when it thinks it is being evaluated, which invalidates every number the Capabilities Report rests on. Undermining Safeguards is a model working against the mitigations chosen from those numbers. A framework whose validity depends on capabilities it has declined to track carries a large load-bearing assumption, and the stated response is proportionate but weak: beat sandbagging with better elicitation, or use a conservative upper bound. Note too that Model Autonomy, one tracked category in the beta, was split here — self-improvement promoted to Tracked, long-range autonomy and autonomous replication demoted to Research.
Two thresholds, two different obligations
Each Tracked Category carries thresholds rather than a score. High means the capability significantly amplifies an existing risk vector; a system that crosses it must have safeguards that sufficiently minimise the risk before external deployment, plus security controls during development. Critical is stronger in both senses:
"Critical capability thresholds mean capabilities that present a meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent. Critical capabilities require safeguards even during the development of the covered system, irrespective of deployment plans."— Preparedness Framework, Version 2, OpenAI (2025)
The thresholds are written concretely, which is what makes the document readable as a spec. High bio/chem is meaningful counterfactual uplift to a novice — someone with a basic technical background — for creating known threats. High cyber is removing the bottlenecks that limit scaled operations: end-to-end automation against reasonably hardened targets, or automated discovery and exploitation of operationally relevant vulnerabilities. High self-improvement is defined by analogy rather than benchmark — impact equal to giving every OpenAI researcher a high-performing mid-career research engineer, against a 2024 baseline. Critical self-improvement is the recursive one: a superhuman research-scientist agent, or generational improvements at five times the 2024 wall-clock rate, sustained for months.
Read as an engineer, these are thresholds on capability, not risk, and the response is never "don't build it." High means deploy-with-safeguards. Critical means halt further development until a Critical-standard safeguard and security spec exists — a pause defined by the readiness of mitigations, not the danger of the capability.
Measurement, and why the number is a floor
Every covered deployment runs Scalable Evaluations — automated proxies with pre-agreed indicative thresholds — optionally backed by Deep Dives: expert red-teaming, third-party evaluators, wet-lab studies. Evals run against a maximally elicited model: highest capability tier, best available scaffolds, and a variant with safety refusals stripped out, because the question is what an adversary could extract, not what a polite user gets. Even so, the framework concedes that any one-time elicitation is a lower bound, since scaffolding improves after release.
Results go into a Capabilities Report for the Safety Advisory Group (SAG), an internal cross-functional body that decides whether a threshold was crossed, orders a deep dive, or declines. If it was crossed, a Safeguards Report follows: how harm could be realised, the safeguard mapped to each pathway, efficacy evidence, residual risk, stated limitations. That is a safety case in the aviation sense — an argument with evidence that a specific claim holds — and the framework enumerates the claim families you may argue: against malicious users, robustness / usage monitoring / trust-based access; against a misaligned model, lack of autonomous capability, value alignment, instruction alignment, reliable oversight, or an architecture that denies the model the channels it would need. SAG recommends; OpenAI Leadership decides; the Board's Safety and Security Committee oversees.
Where it gives, and where the seams are
Three seams. First, the marginal risk clause in §4.3: if a competitor ships a High- or Critical-capable system without comparable safeguards, OpenAI may lower its own requirements — bounded by three conditions (no meaningful increase in overall risk, public acknowledgement, and staying more protective than the competitor, with information shared to prove it). It is a candid description of a race dynamic and a documented mechanism for ratcheting down. Second, every load-bearing decision is internal: SAG composition, threshold determination, safeguard sufficiency, and the text itself are OpenAI's to change, with no external veto or audit requirement — the point pressed by an affordance analysis from Coggins et al. (2025), which argues the framework guarantees no specific mitigation practice at all. Third, the severe-harm bar is set so high that a great deal of real harm sits outside the document by construction.
What the framework looks like when it fires
The reason the course pairs the framework with Section 5 of the GPT-5 system card is that the abstraction only becomes legible when you watch it run. For GPT-5, OpenAI treated the launch as High in Biological and Chemical — and was explicit that this was precautionary rather than evidenced:
"We do not have definitive evidence that this model could meaningfully help a novice to create severe biological harm, our defined threshold for High capability, and the model remains on the cusp of being able to reach this capability. We are treating the model as such primarily to ensure organizational readiness for future updates."— GPT-5 System Card, OpenAI (2025)
That triggered the safeguard stack: a biothreat taxonomy, safety training against it, always-on classifiers covering 100% of production traffic on the affected models, and account-level enforcement. Cyber came in below High (comparable to o3 on capture-the-flag and cyber-range work) and self-improvement showed modest sub-threshold gains, with third parties reporting alongside — METR measured a ~2h17m 50% time horizon on agentic software tasks against a 40-hour concern threshold, and Apollo Research probed Sandbagging directly, finding covert-action rates near 4% versus o3's 8% and a model that sometimes notices it is being evaluated.
It has kept firing since. The GPT-5.6 assessment rates all three family members High in Biological and Chemical and High in Cybersecurity — the first cyber High — with per-model tailored safeguards and below-High self-improvement. In September 2026 OpenAI reported Astra crossing the Critical cyber threshold, the first Critical determination in company history, with frontier training runs paused, infrastructure hardened, and access restricted. Note the tension that leaves in your reading: §4.4 says OpenAI expects to update the framework before any model reaches Critical, yet the assigned document is still Version 2 of April 2025.
Readings, linked
The course budgets about an hour: read the Preparedness Framework first (it is the spec), then skim Section 5 of the GPT-5 system card (it is the spec executed once). Both course links point at OpenAI's own CDN and are public — no login needed.
- OpenAI's Preparedness Framework (Version 2) — OpenAI (15 April 2025) · ~40 min · The primary document: Tracked and Research Categories with the five inclusion criteria, the High/Critical threshold table, the Capabilities Report → SAG → Safeguards Report pipeline, the marginal-risk clause, and appendices of illustrative safeguards and security controls. Read §2 and §4 closely; the appendices are skimmable.
- GPT-5 System Card — OpenAI (2025) · ~20 min, skim Section 5 · The framework applied to a real launch: the precautionary High call in Biological and Chemical, the cyber and self-improvement evaluations that came in below threshold, the biothreat taxonomy and always-on classifier stack, and the external evaluations from METR and Apollo Research.
Exercises
- Safety testing — Pick one dangerous capability from your chosen company's framework (here: Biological and Chemical, Cybersecurity, or AI Self-improvement) and answer five questions about it in writing. Limits: which specific, observable results would tell you it is unsafe to keep scaling? Protections: which parts of the current protective measures are actually necessary to contain catastrophic risk from this capability — and which are decoration? Evaluation: what procedures catch early warning signs promptly, before the limit is crossed rather than after? Response: if the capability passes the limit and protections cannot be improved quickly, is the developer genuinely prepared to pause capability work and handle the dangerous model with caution? Accountability: how does anyone — employees, government, the public — verify the commitments are being executed, critique them from outside, and notice if the framework itself is changed in a rushed or opaque way? The course suggests 45 minutes reading and 15 minutes writing. What a good answer has: section-and-page citations rather than impressions; a distinction drawn everywhere between what the document commits to and what it merely describes; at least one concrete failure scenario the framework would not catch (sandbagged evals, a capability that only appears with better post-release scaffolding, a threat below the severe-harm bar); an explicit reading of the marginal-risk clause and of who can amend the document; and a comparison point from a sibling chapter — Anthropic's RSP ties obligations to ASL levels where OpenAI ties them to per-category thresholds, which changes what "pause" means in practice.
- Threshold-to-eval traceability code — Field map extra. The framework's promises route through evaluations, so trace one threshold end to end and see what the evidence actually supports. Take High Cybersecurity ("removes existing bottlenecks to scaling cyber operations… OR automating the discovery and exploitation of operationally relevant vulnerabilities") and map it to the concrete evaluations reported in the GPT-5 system card — capture-the-flag at high-school/collegiate/professional tiers with pass@12 over 16 rollouts, and the five cyber-range scenarios — then argue in writing whether a model saturating those evals would in fact satisfy the written threshold, and whether a model that fails them could still satisfy it. What a good answer has: a table mapping threshold clause → eval → metric → reported result → gap; at least one named construct-validity failure (scaffold sensitivity, contamination, wall-clock or context limits, pass@k inflating a single lucky rollout); and a proposal for the eval you would add. Start here: (1)
pdftotextthe two PDFs and grep for the threshold text and the eval names, so you are quoting rather than remembering; (2) pick a small open-weight model you can run locally withollamaor on a free Colab; (3) build a five-task mini-benchmark from public beginner CTF challenges using UK AISI's Inspect or a 100-line harness of your own; (4) score pass@1 and pass@8 and watch the gap — that gap is the elicitation uncertainty the framework calls a lower bound; (5) re-run one task with a better prompt or a tool-use scaffold and record how much capability the scaffold alone unlocked; (6) write the one paragraph you would add to a Capabilities Report about what your numbers do not establish.
Go deeper
- GPT-5.6 Preparedness assessment (OpenAI Deployment Safety Hub) — the framework a year on: High in Biological and Chemical and High in Cybersecurity across Sol, Terra and Luna, with safeguards tailored per model rather than per rating. Useful for seeing how a threshold call scales to a model family.
- Path to Astra: critical capabilities and frontier safeguards (OpenAI, Sept 2026) — the first Critical determination, in cybersecurity. Note: openai.com blocks automated fetches, so this link is unverified by fetch; the reporting is corroborated by CNBC's coverage, which does resolve.
- The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices — Coggins, Saeri, Daniell, Ruster, Liu & Davis (2025). An affordance analysis of the exact document you just read; the strongest published version of the "procedure, not promise" critique, and a template you can reuse on any lab's framework.
- METR's evaluation of GPT-5 — what a third-party pre-deployment assessment looks like, including the time-horizon methodology and an explicit statement of what the access METR was granted did and did not let it conclude.
- Anthropic's Responsible Scaling Policy — the chapter 3.3 counterpart. Reading the two side by side is the fastest way to see which parts of a frontier safety framework are genre convention and which are a real design choice.