Google DeepMind Frontier Safety Framework Update

Google DeepMind has updated its Frontier Safety Framework (FSF) to the third iteration, expanding risk domains and refining assessment processes to mitigate severe risks from advanced AI models. This update focuses on operationalizing the detection of harmful manipulation and addressing misalignment risks associated with AI-driven research and development.

New Critical Capability Level for Harmful Manipulation

Google DeepMind has introduced a Critical Capability Level (CCL) specifically for harmful manipulation. This CCL is triggered when AI models possess manipulative capabilities that could be misused to systematically and substantially change beliefs and behaviors in high-stakes contexts, potentially resulting in severe scale harm. This addition is based on research into the mechanisms that drive manipulation in generative AI.

Addressing Misalignment and AI Research Risks

The updated framework expands its scope to address scenarios where misaligned AI models might interfere with an operator's ability to direct, modify, or shut down operations.

Key changes to the approach to misalignment include:

  • Shift from Instrumental Reasoning: The framework moves away from an exploratory approach centered on instrumental reasoning CCLs (warning levels for deceptive thinking) toward protocols focused on models that could accelerate AI research and development to destabilizing levels.
  • Risk Integration: The framework now accounts for misalignment risks stemming from a model's potential for undirected action and the integration of these models into deployment processes.
  • Expanded Safety Case Reviews: Safety case reviews—detailed analyses demonstrating that risks have been reduced to manageable levels—are now required not only for external launches but also for large-scale internal deployments when advanced machine learning research and development CCLs are reached.

Refined Risk Assessment and Governance

Google DeepMind has sharpened its Critical Capability Level (CCL) definitions to better identify critical threats requiring the most rigorous governance. While standard safety and security mitigations are applied throughout model development, the CCLs serve as triggers for heightened mitigation strategies.

The risk assessment process now includes:

  • Systematic risk identification.
  • Comprehensive analyses of model capabilities.
  • Explicit determinations of risk acceptability.

Introduction of Tracked Capability Levels (FSF 3.1)

As of April 17, 2026, Google DeepMind introduced Tracked Capability Levels (TCLs) in specific domains. TCLs are designed to identify and evaluate potential risks that are less extreme than those defined by CCLs, allowing the lab to spot and mitigate risks sooner in the development cycle.

Commitment to Evidence-Based Safety

The Frontier Safety Framework is designed to evolve based on stakeholder input, new research, and implementation lessons. Google DeepMind states that the framework aims to ensure the benefits of transformative AI are realized while minimizing harms as capabilities advance toward Artificial General Intelligence (AGI).

Sources