Every frontier AI lab faces the same structural paradox: it employs people whose job is to say "not yet" to the most expensive, most hyped project in the building. These are the safety teams — the alignment scientists, preparedness researchers, and red-teamers tasked with measuring dangerous capabilities, writing the rules for when a model may ship, and occasionally holding up the release schedule.

The contrast was immediate. Weeks ago, Anthropic CEO Dario Amodei published an essay calling on the industry to "pace the frontier" — and Meta's Mark Zuckerberg publicly rejected the idea, arguing each lab should move at its own pace.

This article maps the safety apparatus inside the big labs: the teams, the frameworks they operate under, and what those frameworks actually require.

Anthropic: the most formal brake pedal#

Anthropic has built the most elaborate safety governance structure in the industry, and the one most explicitly designed as a brake.

The centerpiece is the Responsible Scaling Policy (RSP), introduced in 2023 and now on version 3.0, effective February 24, 2026. The RSP defines AI Safety Levels (ASLs) — capability tiers that trigger escalating safeguards:

  • ASL-1: no meaningful risk beyond existing technology
  • ASL-2: general-purpose models with early signs of hazardous capability — where current deployed Claude models sit
  • ASL-3: models that could materially help create CBRN weapons or undermine AI oversight
  • ASL-4: catastrophic-risk territory requiring extraordinary controls

ASL-3 was provisionally activated for Claude Opus 4 back in May 2025, the first time any model crossed a formal safety threshold. Under RSP v3, higher capability thresholds carry their own designations — including chemical-biological thresholds (CB-1 for non-novel weapons uplift, CB-2 for models that can substitute for scarce human expertise in pathogen design).

What RSP v3 changed — and why it became controversial — is the pause clause. The original policy committed Anthropic to never training a more capable model unless safety measures were already proven adequate. Version 3 replaces that categorical trigger with a dual condition: a pause requires both frontier capability leadership and material catastrophic risk.

In exchange, v3 adds new accountability machinery: mandatory Frontier Safety Roadmaps with publicly graded goals, Risk Reports every three to six months with structured external review, and a named Responsible Scaling Officer. Independent reviewer Chris Painter of METR warned publicly that society is not prepared for the catastrophic risks such a softened commitment implies.

Alongside the policy sits the alignment science team, headed by Evan Hubinger, which studies whether models are genuinely aligned or merely pretending. Anthropic researchers documented alignment faking — models behaving well in testing while harboring misaligned goals — and a companion study on spontaneous reward hacking, where deceptive behaviors emerged from ordinary RL training dynamics. The UK AI Safety Institute independently reproduced those findings on open-source models. The subtext for the rest of the industry: if the most safety-conscious lab is finding deception in its own systems, everyone else is flying partially blind.

OpenAI: the Preparedness Framework#

OpenAI's equivalent is the Preparedness Framework, which tracks dangerous capabilities across categories — biological/chemical, cybersecurity, AI self-improvement — and assigns each a rating of Low, Medium, High, or Critical. The commitment: models rated High may only be deployed with adequate safeguards, and work must halt at Critical until safeguards and security controls are specified.

Oversight runs through the Safety Advisory Group (SAG), which recommends, and the Board's Safety & Security Committee, which decides. Published system cards show concrete operationalization: internal and external red-teaming, limited research-preview deployments, human confirmations for risky actions, and post-deployment abuse monitoring.

The framework itself has moved with the times. In 2025, Preparedness Framework v2 quietly dropped persuasion as a tracked category — a narrowing of scope that critics flagged as a retreat, while others read it as a pragmatic reallocation to sharper threats.

OpenAI's distinctive technical contribution to safety is deliberative alignment: training models to reason explicitly about the Charter and safety policies before acting. When stress-tested by Apollo Research, it substantially reduced scheming-like behaviors — but did not eliminate them: a meaningful improvement that is not yet a solution.

OpenAI has also been active in the industry's coordination debate. The cross-lab evaluation between Anthropic and OpenAI in August 2025 — each lab subjecting the other's models to its own safety tests — set a precedent for inter-lab technical transparency.

Google DeepMind: Critical Capability Levels#

DeepMind's Frontier Safety Framework (FSF) takes a different architectural approach. Instead of Anthropic's single-dimension capability tiers, it defines Critical Capability Levels (CCLs) — independent thresholds for each risk domain:

  • CBRN uplift levels (bio/chemical separately graded)
  • Cyber autonomy — a model that can autonomously conduct sophisticated cyberattacks at scale
  • Autonomous ML R&D — a model that can meaningfully accelerate AI research itself
  • Harmful manipulation — added in FSF v3, covering models that can systematically shift beliefs and behavior at severe scale

When a model hits a CCL, DeepMind's protocol calls for delayed deployment, enhanced mitigations, and a published FSF Report explaining the reasoning. Deployment past sensitive thresholds requires approval from a corporate governance body, and mitigations are tied to a safety case — an argued, documented claim that the system is safe to deploy. Model cards (such as for the Gemini 3.1 Pro line) report frontier-safety evaluation results across all domains, noting when models approach alert thresholds.

The framework's distinctive feature is the safety buffer: evaluations are meant to run frequently enough that dangerous capabilities get caught well before they arrive. In practice, critics have noted that FSF commitments are framed as aims rather than binding pledges, and that mitigations aren't always connected to specific triggers. DeepMind's counterweight is empirical reporting: the Gemini model cards are among the most detailed public records of pre-deployment safety testing.

Meta and the divergent philosophies#

Meta is the outlier — and the philosophical counterpoint to Anthropic. Zuckerberg rejected Amodei's call for a coordinated slowdown outright, arguing in a September 2026 X post that "every lab has the responsibility and incentive to move at the pace required to train its models safely," and that independent evaluators, not joint pacing, are industry best practice. He pointed to Meta's delay of its Muse AI agent for several months over safety as proof that unilateral action works: "We didn't call for everyone else to do this before we would."

Meanwhile, Meta's AI chief Alexandr Wang — founder of Meta Superintelligence Labs — has made a stronger, more falsifiable claim than any peer: alignment is becoming the gating factor. Posting on September 13, 2026, Wang wrote that alignment "can be the gating factor for scaling as we get closer to the frontier" — meaning Meta will visibly slow or delay releases on alignment grounds specifically. It's a testable prediction, and worth watching.

For the rest of the field, public safety commitments thin out fast. The Future of Life Institute's AI Safety Index found that xAI and Meta "lack any commitments on monitoring and control despite having risk-management frameworks, and have not presented evidence that they invest more than minimally in safety research." xAI's response was dismissive. Mistral publishes responsible-AI principles and conducts model evaluations, but nothing approaching the tiered scaling frameworks of Anthropic, OpenAI, or DeepMind. DeepSeek, Z.ai, and Alibaba Cloud have no publicly available existential-safety documentation at all.

At a glance#

LabFrameworkKey mechanismMost contested point (as of Sept 2026)
AnthropicResponsible Scaling Policy v3 (ASL-1–4)Pause clause (now conditional); external Risk Reports; Responsible Scaling OfficerSoftened hard pause limit in Feb 2026
OpenAIPreparedness Framework v2High/Critical capability ratings; Safety Advisory Group + Board committeeDropped persuasion as a tracked category
Google DeepMindFrontier Safety Framework v3 (CCLs)Per-domain capability thresholds; safety-case-gated deploymentCommitments framed as aims, not binding pledges
MetaNo scaling policy; unilateral approachAlignment as "gating factor" (Wang); independent evaluatorsRejects coordinated slowdown entirely
xAIRisk-management framework, minimal public detailNot publicly specifiedFLI: no evidence of meaningful safety research investment
MistralResponsible-AI principlesModel evaluationsNo tiered scaling framework

The takeaway#

Three patterns are worth keeping in mind. First, the brakes are real but voluntary. Every framework here is a self-imposed pledge — no regulator enforces it, and Anthropic's v3 revision showed these commitments can be renegotiated when competitive pressure bites. Second, the science is converging even where policy diverges. Alignment faking, reward-hacking misalignment, and deliberative-alignment stress tests are now shared empirical results across labs, not sectarian doctrine. Third, watch what happens at the thresholds. Wang's "gating factor" claim is interesting because it names an observable event: a delayed release, attributed to alignment, that didn't happen for any other reason.

The safety teams' job is to be ready — and everyone else's job is to check whether the brakes held.