Option 3: Google DeepMind
Chapter 3.2 argued that frontier safety frameworks are a genre: every major lab publishes one promising to measure dangerous capability, tie mitigations to thresholds, and stop if the mitigations are not ready. Chapters 3.3–3.6 are where the genre meets specific text. Reading one closely beats skimming four: the differences live in whether a threshold is defined by an observable, whether a mitigation is a commitment or a recommendation, and whether anyone outside the company can check.
Thresholds, not products
Most safety documents are organised around products. The FSF is organised around thresholds defined without reference to any model — the inversion everything else follows from.
"CCLs are determined by identifying and analyzing the main foreseeable paths through which a model could result in severe harm: we then define the CCLs as the minimal set of capabilities a model must possess to do so."— Frontier Safety Framework 3.1, Google DeepMind (2026)
Read backwards, that is a threat-modelling procedure: enumerate the routes to catastrophe, then ask the minimum capability that puts a model on one. The threshold is the floor of the dangerous region, not a score — which is why FSF thresholds read as sentences, not numbers, and why no determination is ever mechanical. DeepMind says as much, calling its assessments subjective and expert-led.
Version 3.1 stacks three tiers, and confusing them is the commonest way to misread it. CCLs mark heightened risk of severe harm; reaching one triggers the heavy machinery — safety case, residual risk assessment, named security level. Tracked Capability Levels (TCLs), new in 3.1, are the same shape at a lower bar (significant but not severe; so far CBRN and ML R&D/misalignment), and exist because a framework that only fires at "catastrophe" gives you no practice runs and no early data. Alert thresholds sit marginally below a CCL and are purely temporal: capability moves faster than the assessment cycle, so the alert fires while there is time to react. Crossing one means the safety margin is consumed and someone must look again — not that a CCL was reached.
The alert threshold is the piece worth stealing — page on the leading indicator, not the outage. It also creates the framework's most honest failure mode: when an alert fires and the analysis finds risk still acceptable, a permitted response is to move the alert threshold.
Four domains, and what each one recommends
- CBRN uplift level 1 — uplift to low-to-medium-resourced actors in reference scenarios, at severe scale. Recommended Security Level 2+; a CBRN TCL sits below it.
- Cyber uplift level 1 — sufficient uplift on high-impact attacks for severe-scale harm. SL2+, reasoning that automated cyber-defence and social adaptation blunt the exfiltration case.
- Harmful manipulation level 1 (exploratory) — enough to systematically and substantially shift beliefs and behaviour in identified high-stakes contexts over the course of interaction. SL2+; DeepMind labels the domain exploratory and its research base nascent.
- ML R&D acceleration level 1 (AI progress substantially above historical rates) — SL3; ML R&D automation level 1 (fully automates any Google team working on AI capabilities, at comparable all-in cost) — SL4, and only if the whole field adopts it.
Those levels are aligned to RAND's Securing AI Model Weights, which grades weight-theft resistance from amateurs to top-tier state operations. DeepMind's addition is SL2+: RAND SL2 plus insider-threat controls — insider risk teams, background checks, training-data tamper review, sandboxed untrusted input, APT-simulating red teams, 24/7 threat hunting. Notice the implication: every misuse CCL recommends a level explicitly scoped below resisting a well-resourced state actor. Argued for in the same table as the threshold — but still a trade your audit should surface.
What a crossed threshold actually triggers
Two mitigation families on different logic. Security mitigations protect the weights, because exfiltrated weights can have their safeguards trained back out; these are the RAND-aligned levels. Deployment mitigations — safety post-training, input/output and chain-of-thought monitoring, jailbreak detection, account moderation, user verification, bug bounties — counter misuse or misaligned expression in a live deployment. That side is deliberately not a fixed list, since you cannot pre-specify defences for capabilities that do not exist yet; 3.1 gives a process: build safeguards, red-team their robustness, produce a residual risk assessment, and — if a CCL rather than merely a TCL was reached — add a safety case, an assessable argument that severe risk is now acceptable. Deployment waits on a governance function accepting that risk. One consideration a safety case may lean on is "what capabilities and mitigations are available on other publicly available models" — real counterfactual reasoning, and structurally a race-to-the-bottom clause.
Misalignment: the part still under construction
Version 3.1 sets a Stealth and Situational Awareness TCL: instrumental reasoning good enough — situational awareness (using details of its own deployment) plus stealth (circumventing basic oversight) — that, absent extra mitigations, undermining human control cannot be ruled out. Above it, DeepMind commits to periodic misalignment risk assessments covering internal high-risk deployments, naming chain-of-thought monitoring as an example safeguard. Two flags: the claim is inability-shaped — "cannot scheme effectively", not "will not" — weaker than a propensity argument and expiring the moment the capability arrives; and the domain has no paired misalignment CCL, so it is handled by assessment rather than a named severe-harm threshold. The eval suite behind the TCL is public — Evaluating Frontier Models for Stealth and Situational Awareness (Phuong et al., 2025) — more than most frameworks offer.
The framework in contact with a real model
"We ran our full suite of early warning evaluations on Gemini 3 Pro. We found that Gemini 3 Pro did not reach any of our FSF CCLs."— Gemini 3 Pro Frontier Safety Framework Report, Google DeepMind (2025)
Underneath the headline: on cyber the alert threshold was met — 11 of 12 hard v1 challenges solved, 0 of 13 on the harder v2 set — and the response was to keep deploying while strengthening cyber safeguards. On CBRN, an alert threshold had earlier been met by Gemini 2.5 Deep Think and the CCL could not immediately be ruled out; later analysis cleared it, and under the recalibrated framework that alert threshold is no longer met. On misalignment, the agent solved 3/11 situational-awareness and 1/4 stealth challenges. The report also uses a "cannot rule out being at the CCL" designation when evidence is insufficient to clear a model — a default-to-caution rule worth copying.
Sit with that recalibration: a threshold moved and a domain left the alert region. The report is transparent for saying so, but no audit can verify the reasoning, because the supporting evidence is not public. That is the structural fact about all four frameworks in this unit: the lab writes the exam, sits it, grades it.
Where it breaks
- No pause commitment. Search 3.1 for "pause" and you get nothing; the exercise's fourth question has no direct answer in the text, and the closest thing gates release, not training.
- Security levels are recommendations, and conditional ones — framed as the minimum the field should apply, with a reserved right to apply less if the capability is not meaningfully ahead of public models. And alert thresholds are adjustable by the party they constrain: a permitted response plan is to move one. Reasonable in principle, unfalsifiable from outside.
- Accountability is the thinnest section. Section 4 runs a few sentences: no named committee, no board veto, no scheduled public reporting beyond launch documentation, disclosure to government framed as an aim conditioned on unmitigated material risk. External review happens — the 2026 review of DeepMind's scheming-inability safety case — but voluntarily, not as a standing requirement.
- Harmful manipulation is self-labelled exploratory. A threshold whose owner says the underlying science is nascent is not a tested tripwire.
None of this makes the FSF a bad document — it is one of the more precise entries in the genre. The exercise is to separate its commitments from its intentions.
Readings, linked
The course budgets 1h for this chapter: 45 minutes reading, 15 minutes writing. Start with the framework itself — but read the current 3.1 rather than the 2.0 PDF the course links — then use the Gemini 3 Pro card to see the framework applied.
- Google DeepMind's Frontier Safety Framework (v2.0, as linked by the course) — Google DeepMind (2025) · core · The version BlueDot assigns. Still live, but two revisions behind: read Frontier Safety Framework 3.1 (April 2026) instead, announced in Strengthening our Frontier Safety Framework. Everything the exercise asks about — thresholds, mitigations, response, governance — is in Sections 1–5 and the glossary.
- Gemini 2.5 Pro Model Card — Google DeepMind (2025) · optional ("maybe") · Course URL moved. The link served by the course (
modelcards.withgoogle.com/assets/documents/gemini-2.5-pro.pdf) now redirects to DeepMind's model-cards index; the live PDF is the one linked here, found via deepmind.google/models/model-cards. Useful only as a before-and-after against the 3 Pro card — 2.5 Pro is where the cyber and CBRN alert thresholds were first met. - Gemini 3 Pro Model Card — Google DeepMind (2025) · core · The framework applied. Its Frontier Safety section carries the per-domain CCL table and states the model was assessed against the September 2025 framework; the long form is the separate Gemini 3 Pro Frontier Safety Framework Report, which is where the evaluation detail, the alert-threshold results and the external-testing notes actually live. Read the report, not just the card.
Exercises
- Safety testing — Pick one dangerous capability from the lab you chose in 3.2 (for DeepMind: CBRN, cyber, harmful manipulation, or ML R&D/misalignment) and answer five questions about it from the framework's own text. Limits: which specific, observable findings about that capability would indicate it is — or plausibly might be — unsafe to keep scaling? Protections: which parts of the current protective measures are actually necessary to contain catastrophic risk from that capability, as opposed to nice to have? Evaluation: what are the procedures for catching early warning signs promptly, before the limit is crossed? Response: if the capability goes past the limit and protections cannot be improved fast enough, is the developer prepared to halt further capability gains until they can, and to handle the dangerous model with corresponding caution? Accountability: how does the developer ensure its commitments are executed as written; how can outsiders verify that, or notice when it fails; where is third-party critique invited; and what stops the framework itself from being rewritten quickly and quietly? The course budgets 45 minutes reading and 15 minutes writing. What a good answer has: a quoted threshold with a page or section reference for each claim, not a summary from memory; a clear separation of what the document commits to from what it recommends or aims to do; at least one place where you could not answer the question from the text, named as a gap rather than filled with charity; and, for DeepMind specifically, an explicit note that you audited version 3.1 rather than the 2.0 the course links, plus what changed between them. The strongest answers close with the single change to the framework that would most improve it — and say what it would cost.
- Field map extra: diff the framework against itself code — Frameworks change, and the changes are where the real commitments show. Build a small diff harness over the FSF version history (1.0 May 2024, 2.0 Feb 2025, 3.0 Sep 2025, 3.1 Apr 2026) and characterise the drift: which thresholds were added, which were reworded, and — the interesting one — which obligations got weaker. What a good answer has: a per-version table of risk domains and threshold names; a list of modal-verb changes (will → may, commit → recommend) with surrounding sentences; and one paragraph on whether the trend is toward more precision or more discretion. Start here: (1) fetch the version PDFs from deepmind.google/frontier-safety and the FSF blog posts; (2) extract text with
pdftotext -layoutorpypdf; (3) segment by heading and align sections across versions by title similarity (rapidfuzz); (4) rundifflib.unified_diffper aligned section and render side by side; (5) grep the modal verbs —will|shall|must|commit|may|aim|recommend|intend— and count them per version, normalised by word count; (6) sanity-check by hand that your aligner did not silently drop a renamed section, which is exactly where a weakened obligation would hide. Runs on a laptop in minutes, no model required. Extend it across labs — Anthropic's RSP and OpenAI's Preparedness Framework are versioned too — and you have a comparison the published literature mostly lacks.
Go deeper
- Frontier safety at Google DeepMind — the hub page, and the canonical place to check which FSF version is current before you quote one.
- An Approach to Technical AGI Safety and Security — Shah et al., Google DeepMind (2025). The research agenda the FSF sits on top of: misuse, misalignment, mistakes and structural risk, and the technical defences DeepMind expects to build against each. Read it if you want the "why these four domains".
- Evaluating Frontier Models for Stealth and Situational Awareness — Phuong et al. (2025). The concrete eval suite behind the Stealth and Situational Awareness TCL: 5 stealth and 11 situational-awareness evaluations. The best available answer to "what does a dangerous-capability eval for scheming actually look like".
- Lessons from External Review of DeepMind's Scheming Inability Safety Case — Barrett et al. (2026). Independent reviewers picking apart one of DeepMind's safety cases, including where the argument's scope does not support the decision it is used for. Exactly the third-party critique the accountability question asks about — and evidence of how rare and ad hoc it still is.
- Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models — Nevo et al., RAND (2024). The source of the SL1–SL5 ladder the FSF's security recommendations are pinned to; 38 attack vectors graded by attacker class. Needed to know what "Security Level 2+" is actually promising and, more importantly, what it is not.