TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 3 · DETECTING DANGERchapter 5 · 1h

Option 3: Google DeepMind

BlueDot Impact · Technical AI Safety · unit 3, chapter 5
TL;DR — One of four lab-specific deep dives: pick a lab, audit its safety framework against five questions. DeepMind's answer is the Frontier Safety Framework, built on Critical Capability Levels — named thresholds in CBRN, cyber, harmful manipulation, and ML R&D/misalignment, each with a recommended security level and a mitigation process. The machinery that does the work sits a layer below: alert thresholds that fire early, early warning evaluations that measure proximity, a safety case owed before a model at a CCL ships. It is specific about detection and vague about consequence: "pause" does not appear in it, and the body that signs off is named only as "the appropriate governance function."

Chapter 3.2 argued that frontier safety frameworks are a genre: every major lab publishes one promising to measure dangerous capability, tie mitigations to thresholds, and stop if the mitigations are not ready. Chapters 3.3–3.6 are where the genre meets specific text. Reading one closely beats skimming four: the differences live in whether a threshold is defined by an observable, whether a mitigation is a commitment or a recommendation, and whether anyone outside the company can check.

Version check first. The course links FSF 2.0 (Feb 2025); DeepMind has shipped 3.0 (22 Sep 2025) and 3.1 (17 Apr 2026) since — current text at deepmind.google/frontier-safety. Not cosmetic: 3.0 added a harmful-manipulation CCL; 3.1 added Tracked Capability Levels and folded the standalone misalignment domain into a combined ML R&D and misalignment domain. Audit the current version and note the drift.

Thresholds, not products

Most safety documents are organised around products. The FSF is organised around thresholds defined without reference to any model — the inversion everything else follows from.

"CCLs are determined by identifying and analyzing the main foreseeable paths through which a model could result in severe harm: we then define the CCLs as the minimal set of capabilities a model must possess to do so."— Frontier Safety Framework 3.1, Google DeepMind (2026)

Read backwards, that is a threat-modelling procedure: enumerate the routes to catastrophe, then ask the minimum capability that puts a model on one. The threshold is the floor of the dangerous region, not a score — which is why FSF thresholds read as sentences, not numbers, and why no determination is ever mechanical. DeepMind says as much, calling its assessments subjective and expert-led.

Version 3.1 stacks three tiers, and confusing them is the commonest way to misread it. CCLs mark heightened risk of severe harm; reaching one triggers the heavy machinery — safety case, residual risk assessment, named security level. Tracked Capability Levels (TCLs), new in 3.1, are the same shape at a lower bar (significant but not severe; so far CBRN and ML R&D/misalignment), and exist because a framework that only fires at "catastrophe" gives you no practice runs and no early data. Alert thresholds sit marginally below a CCL and are purely temporal: capability moves faster than the assessment cycle, so the alert fires while there is time to react. Crossing one means the safety margin is consumed and someone must look again — not that a CCL was reached.

The alert threshold is the piece worth stealing — page on the leading indicator, not the outage. It also creates the framework's most honest failure mode: when an alert fires and the analysis finds risk still acceptable, a permitted response is to move the alert threshold.

Four domains, and what each one recommends

Those levels are aligned to RAND's Securing AI Model Weights, which grades weight-theft resistance from amateurs to top-tier state operations. DeepMind's addition is SL2+: RAND SL2 plus insider-threat controls — insider risk teams, background checks, training-data tamper review, sandboxed untrusted input, APT-simulating red teams, 24/7 threat hunting. Notice the implication: every misuse CCL recommends a level explicitly scoped below resisting a well-resourced state actor. Argued for in the same table as the threshold — but still a trade your audit should surface.

What a crossed threshold actually triggers

Two mitigation families on different logic. Security mitigations protect the weights, because exfiltrated weights can have their safeguards trained back out; these are the RAND-aligned levels. Deployment mitigations — safety post-training, input/output and chain-of-thought monitoring, jailbreak detection, account moderation, user verification, bug bounties — counter misuse or misaligned expression in a live deployment. That side is deliberately not a fixed list, since you cannot pre-specify defences for capabilities that do not exist yet; 3.1 gives a process: build safeguards, red-team their robustness, produce a residual risk assessment, and — if a CCL rather than merely a TCL was reached — add a safety case, an assessable argument that severe risk is now acceptable. Deployment waits on a governance function accepting that risk. One consideration a safety case may lean on is "what capabilities and mitigations are available on other publicly available models" — real counterfactual reasoning, and structurally a race-to-the-bottom clause.

Misalignment: the part still under construction

Version 3.1 sets a Stealth and Situational Awareness TCL: instrumental reasoning good enough — situational awareness (using details of its own deployment) plus stealth (circumventing basic oversight) — that, absent extra mitigations, undermining human control cannot be ruled out. Above it, DeepMind commits to periodic misalignment risk assessments covering internal high-risk deployments, naming chain-of-thought monitoring as an example safeguard. Two flags: the claim is inability-shaped — "cannot scheme effectively", not "will not" — weaker than a propensity argument and expiring the moment the capability arrives; and the domain has no paired misalignment CCL, so it is handled by assessment rather than a named severe-harm threshold. The eval suite behind the TCL is public — Evaluating Frontier Models for Stealth and Situational Awareness (Phuong et al., 2025) — more than most frameworks offer.

The framework in contact with a real model

"We ran our full suite of early warning evaluations on Gemini 3 Pro. We found that Gemini 3 Pro did not reach any of our FSF CCLs."— Gemini 3 Pro Frontier Safety Framework Report, Google DeepMind (2025)

Underneath the headline: on cyber the alert threshold was met — 11 of 12 hard v1 challenges solved, 0 of 13 on the harder v2 set — and the response was to keep deploying while strengthening cyber safeguards. On CBRN, an alert threshold had earlier been met by Gemini 2.5 Deep Think and the CCL could not immediately be ruled out; later analysis cleared it, and under the recalibrated framework that alert threshold is no longer met. On misalignment, the agent solved 3/11 situational-awareness and 1/4 stealth challenges. The report also uses a "cannot rule out being at the CCL" designation when evidence is insufficient to clear a model — a default-to-caution rule worth copying.

Sit with that recalibration: a threshold moved and a domain left the alert region. The report is transparent for saying so, but no audit can verify the reasoning, because the supporting evidence is not public. That is the structural fact about all four frameworks in this unit: the lab writes the exam, sits it, grades it.

Carry this away: in the FSF, detection is specified and consequence is discretionary. Thresholds, early warning evaluations, safety cases and residual risk assessments are defined in operational detail; what must happen once a threshold is crossed routes through "the appropriate governance function determines the residual risk to be acceptable" — a body the document never names. When auditing any framework, read the verbs: will versus may, commit versus recommend, determine versus disclose.

Where it breaks

None of this makes the FSF a bad document — it is one of the more precise entries in the genre. The exercise is to separate its commitments from its intentions.

Readings, linked

The course budgets 1h for this chapter: 45 minutes reading, 15 minutes writing. Start with the framework itself — but read the current 3.1 rather than the 2.0 PDF the course links — then use the Gemini 3 Pro card to see the framework applied.

Exercises

  1. Safety testing — Pick one dangerous capability from the lab you chose in 3.2 (for DeepMind: CBRN, cyber, harmful manipulation, or ML R&D/misalignment) and answer five questions about it from the framework's own text. Limits: which specific, observable findings about that capability would indicate it is — or plausibly might be — unsafe to keep scaling? Protections: which parts of the current protective measures are actually necessary to contain catastrophic risk from that capability, as opposed to nice to have? Evaluation: what are the procedures for catching early warning signs promptly, before the limit is crossed? Response: if the capability goes past the limit and protections cannot be improved fast enough, is the developer prepared to halt further capability gains until they can, and to handle the dangerous model with corresponding caution? Accountability: how does the developer ensure its commitments are executed as written; how can outsiders verify that, or notice when it fails; where is third-party critique invited; and what stops the framework itself from being rewritten quickly and quietly? The course budgets 45 minutes reading and 15 minutes writing. What a good answer has: a quoted threshold with a page or section reference for each claim, not a summary from memory; a clear separation of what the document commits to from what it recommends or aims to do; at least one place where you could not answer the question from the text, named as a gap rather than filled with charity; and, for DeepMind specifically, an explicit note that you audited version 3.1 rather than the 2.0 the course links, plus what changed between them. The strongest answers close with the single change to the framework that would most improve it — and say what it would cost.
  2. Field map extra: diff the framework against itself code — Frameworks change, and the changes are where the real commitments show. Build a small diff harness over the FSF version history (1.0 May 2024, 2.0 Feb 2025, 3.0 Sep 2025, 3.1 Apr 2026) and characterise the drift: which thresholds were added, which were reworded, and — the interesting one — which obligations got weaker. What a good answer has: a per-version table of risk domains and threshold names; a list of modal-verb changes (will → may, commit → recommend) with surrounding sentences; and one paragraph on whether the trend is toward more precision or more discretion. Start here: (1) fetch the version PDFs from deepmind.google/frontier-safety and the FSF blog posts; (2) extract text with pdftotext -layout or pypdf; (3) segment by heading and align sections across versions by title similarity (rapidfuzz); (4) run difflib.unified_diff per aligned section and render side by side; (5) grep the modal verbs — will|shall|must|commit|may|aim|recommend|intend — and count them per version, normalised by word count; (6) sanity-check by hand that your aligner did not silently drop a renamed section, which is exactly where a weakened obligation would hide. Runs on a laptop in minutes, no model required. Extend it across labs — Anthropic's RSP and OpenAI's Preparedness Framework are versioned too — and you have a comparison the published literature mostly lacks.

Go deeper

Next: Option 4: Meta · Back to the map.