TECHNICAL AI SAFETY // FIELD MAP
← field map
UNIT 3 · DETECTING DANGERchapter 6 · 1h

Option 4: Meta

BlueDot Impact · Technical AI Safety · unit 3, chapter 6
TL;DR — The fourth of four lab-specific reads in unit 3, and the one that puts the most pressure on the frontier-safety-framework genre, because Meta ships model weights. Its framework — renamed the Advanced AI Scaling Framework in version 2 — starts from catastrophic outcomes rather than capability levels, works backwards through threat scenarios to the enabling capabilities, and sorts a model into moderate / high / critical risk before mitigations. The interesting move: the adversary model is a function of release type, so for an open-weight release you must assume an adversary who can fine-tune your safety training away — every mitigation living in the weights scores roughly zero. A framework's teeth are in its irreversibility, not its prose: once weights are public, "deploy with mitigations" has no undo.

Chapters 3.3 to 3.6 are the same assignment run four times: read one lab's frontier safety policy closely enough to grade it. Meta is the option worth taking if you want to find where the genre's assumptions are load-bearing, because Anthropic, OpenAI and Google DeepMind all reason about models they can claw back, and Meta mostly does not.

Outcomes first, capabilities second

Most frontier safety policies are organised around capability tiers: a model reaches a level, the level maps to required mitigations. Meta inverts that. It names a small set of catastrophic outcomes it must prevent, runs threat-modelling workshops to enumerate threat scenarios — causal stories about how such an outcome could actually happen — and only then asks which capabilities would let a system materially contribute to one.

The stated reason is durability: capabilities churn every six months, "a cyberattack causing large-scale casualties" does not. The practical consequence is that the evaluations are designed as rule-outs — cheap, sensitive tests whose job is to establish the absence of risk, escalating to uplift studies and expert red-teaming only when a cheap test trips. Score below a stated bar on simple capture-the-flag challenges and the framework concludes you do not exceed moderate cyber risk, and stops. Defensible engineering — the same logic as a fast path in front of an expensive validator — and also where the framework is most falsifiable, because a rule-out is only as good as the correlation between the cheap test and the real scenario.

Four filters, and what falls through them

An outcome only enters the framework if it clears four criteria: plausible (you can draw a causal pathway and build a simulatable scenario on it), catastrophic (large-scale, potentially irreversible), net new (not already achievable at that scale, by that actor, for that cost, without general-purpose AI), and instantaneous or irremediable (once it happens you cannot undo it).

Read those filters as a scope statement. "Net new" is the marginal-risk argument that has run through the open-weights debate since Kapoor et al. — the question is never "could this help a bad actor" but "how much more than a search engine, a textbook, and last year's open model." That is the right question, and one the framework answers with its own evaluations; the honest position in the literature is that marginal risk for open models is under-measured rather than measured and found small. "Instantaneous or irremediable" does quieter work, excluding slow cumulative harms — information-ecosystem erosion, dependence, labour displacement — by routing them to other internal processes. Whether that is sensible scoping or the quiet removal of the likeliest risks is one of the real judgement calls here.

Three domains, and the odd one out

The framework covers cybersecurity, chemical and biological risk, and — new in version 2 — loss of control. The first two are conventional, with outcomes graded by actor skill and enabling capabilities spelled out as things like PhD-level knowledge in a threat-relevant sub-specialty or the ability to drive biodesign tools.

Loss of control is handled differently, and that is the most interesting technical content in the reading. Rather than enumerate the harms an uncontrolled system might cause — a list nobody can write credibly — Meta defines the outcomes as failures of control mechanisms: losing the ability to evaluate a system before deployment, and losing the ability to monitor it during operation. Both are measurable today, which is the point.

The capability checkpoints that operationalise this are the closest thing in any lab's framework to a concrete tripwire for scheming. For the first outcome: can the model autonomously and consistently finish software tasks that would take a competent human two or more hours, and can it tell it is being evaluated and change its behaviour accordingly? For the second: can it run a hidden side task alongside the user's task, with competing objectives, without the side task degrading the main one — the setup where a monitor sees nothing wrong. Clear a checkpoint and you graduate to enhanced evaluations in realistic conditions plus propensity measures, with illustrative acceptance thresholds named against the MASK honesty benchmark and an agent-misalignment eval. That is unit 3's capability-versus-propensity distinction as a gating rule rather than a slogan, and it takes eval-gaming seriously: dual-use and propensity evals are held out from training teams, and evaluation-awareness counts as a risk factor rather than a nuisance.

Note where each domain's risk lives. Misuse risk — cyber, chem-bio — is a property of who can reach the model. Loss-of-control risk is not:
"Loss of Control risks may occur with similar probability with any type of deployment, including internal deployment."— Advanced AI Scaling Framework v2, Meta (2026)
Access controls are the wrong tool for that domain, which is why unit 5's control agenda exists.

The thresholds, in plain English

Three levels, assigned before mitigations and aggregated across all threat scenarios. Moderate or lower: relevant capabilities present but no substantial contribution to any threat scenario across plausible deployment paths — ship. High: deployment could substantially contribute — do not ship until mitigations are defined, implemented and validated back down to moderate. Critical: continued development could substantially contribute, or deployment would uniquely enable a scenario whose risk cannot be mitigated in the proposed context — lock down weight access under the Chief AI Officer and stop developing until risk returns to moderate.

Two details separate this from marketing copy. The critical tier is the only one that bites on training rather than release, for a real reason: the access restrictions needed to safely hold a uniquely-enabling model are incompatible with a team continuing to work on it. And thresholds are assigned on pre-mitigation capability under maximum elicitation — task-optimised scaffolds, generous token budgets, helpful-only fine-tunes, and, where weights or a fine-tuning API are on the table, deliberate domain-specific capability training to find the ceiling. That is the right call, and also the step most likely to be quietly softened under schedule pressure — which is why the exercise's accountability question matters more than it looks.

What the open-weight case actually changes

Here is what makes this the right chapter to study: the adversary model is parameterised by release type.

"For open-weight releases or deployments with fine-tuning APIs, we would additionally consider adversaries capable of modifying model behavior through continued training."— Advanced AI Scaling Framework v2, Meta (2026)

Follow that through and most of the mitigation toolkit evaporates. Refusal training, safety post-training, RLHF-installed reluctance all live in the weights, and cheap fine-tuning strips them. If the adversary can retrain, the only mitigations that survive live outside the weights — the deployment-side classifier stack a first-party product enforces and a downstream fine-tuner cannot. Which is why the second reading is a model card for a filter. Llama Guard 4 is a 12B dense classifier pruned from Llama 4 Scout, natively multimodal, scoring prompts and responses against a fourteen-category hazard taxonomy — a separate, swappable guard a good-faith deployer runs in front of and behind the generator.

The honest reading: a real mitigation with a stated scope. It reduces harm from ordinary product misuse at scale and gives the open ecosystem a defensible default. It does nothing against the adversary the framework just told you to assume, because that adversary does not run the guard. Which is the structural point — for open weights the release decision is the mitigation decision. No staged rollout, no revocation, no rate limit, no patch. Every other lab in this unit can degrade gracefully; Meta gets one irreversible commit, which makes the pre-mitigation threshold and the willingness to not ship the whole of the safety case.

What version 2 changed, and where it is still soft

Version 1 landed in February 2025 as the Frontier AI Framework, with two domains and thin governance. Version 2 (April 2026) renames it, adds loss of control, names the officers who assign thresholds and approve deployment, commits to a published preparedness report per frontier release, promises a model spec covering intended propensities including acquiescence to shutdown, adds whistleblower protections, and flags radiological/nuclear and physical autonomy as not-yet-measurable. Most of that runs in the direction external critics asked for: named owners, published artefacts, and an explicit "we cannot measure this yet" list.

Where it stays soft, and where your exercise answer should press: the thresholds are qualitative — "substantially contribute" carries the whole load, defined only as "a material factor" — so the same evidence supports several defensible gradings, and the grader reports up the same chain as the shipper. Preparedness reports are published after the decision and are redactable; there is no third-party audit or pre-release outside review. And the pause commitment is narrower than it reads: critical risk stops development of that model, not scaling in general, and the officers who assign the threshold also decide what "sufficient mitigation" means. Compare against the other three labs via METR's index — the differences between them are smaller than the gap between all of them and a policy with binding external review.

Readings, linked

The course budgets 1 hour: roughly 45 minutes reading and 15 minutes writing. Read the framework first and in full — the model card only makes sense once you understand what it is a mitigation for. Skim the framework's Section 3 (Outcomes & Thresholds) and Section 4.2 (Evaluation and mitigation) most carefully; those two carry the argument.

Exercises

  1. Safety testing — The course's single exercise, carried over from chapter 3.2 where you picked one lab and one dangerous capability to follow. Holding that pair fixed, answer five questions about your chosen developer's framework, and answer them with citations to the document rather than impressions of it. Limits: which specific, observable findings about your dangerous capability would mean it is — or plausibly might be — unsafe to keep scaling? Protections: which parts of the current protective measures are actually load-bearing for containing a catastrophe from that capability, as opposed to decorative? Evaluation: what procedure is supposed to catch early warning signs before the limit is crossed, and how often does it run? Response: if the capability blows past the limit and protections cannot be improved quickly, is the developer prepared to pause further capability improvements and treat the existing dangerous model with corresponding caution? Accountability: how does the developer ensure the commitments are executed as written, that outsiders can verify execution or notice its absence, that there is a route for third-party critique, and that the framework itself cannot be rewritten in a rush or in the dark? What a good answer has: a named threshold and the exact wording that defines it, quoted; a distinction between what the document commits to and what it merely describes as current practice; at least one place where you can construct two defensible gradings from the same evidence; a specific test for the accountability question — name who decides, who they report to, what gets published, and whether anyone outside the company sees anything before the release rather than after. For Meta specifically, the strongest answers notice that the response commitment is scoped to the individual model rather than to scaling in general, and that for open-weight releases the response option set collapses to "release or don't" because there is no rollback. Cross-check your answer against a second lab's framework via METR's index — the contrast is what makes the gaps visible.
  2. Strip the safety training code (field map extra) — The framework asserts that for open-weight releases the adversary can modify behaviour through continued training. Test the claim's cheap end yourself, on a small model, and measure how much refusal behaviour survives a trivial fine-tune. Do this on benign-but-refused content — the point is to measure the durability of the refusal mechanism, not to produce harmful outputs, and you should pick a category like "explain how a common household chemical reaction works" that safety training over-refuses rather than anything genuinely dangerous. What a good answer has: a refusal rate before and after, on a held-out prompt set you wrote yourself; the number of training examples and GPU-minutes it took to move the number; and an honest paragraph on what this does and does not show about the framework's adversary model. Start here: (1) pick a small open-weight instruct model that fits a free Colab T4 or a laptop — Llama 3.2 1B/3B Instruct or Qwen2.5 1.5B Instruct; (2) write 40–60 prompts in your chosen benign-but-over-refused category and score baseline refusal rate with a simple string/LLM judge; (3) build a 100–200 example instruction dataset in the same category with compliant answers, generated by hand or by a model that does not refuse them; (4) LoRA fine-tune with peft + trl's SFTTrainer (or unsloth if you want it to fit in less memory), one to two epochs, rank 16; (5) re-score the same held-out prompts and plot refusal rate against training-example count by re-running step 4 at 25 / 50 / 100 / 200 examples; (6) then run Llama Guard (or the smaller Llama Guard 3 1B if 12B will not fit) in front of the fine-tuned model and observe that the external filter's behaviour is completely unchanged by anything you did to the weights — which is exactly the asymmetry the chapter is about.

Go deeper

Next: How does AI think? · Back to the map.