Anthropic Responsible Scaling Policy Update

Anthropic has updated its Responsible Scaling Policy (RSP), a risk governance framework designed to mitigate catastrophic risks from frontier AI systems. The update introduces a more flexible approach to assessing and managing risks, ensuring that safety and security safeguards scale proportionally with the capabilities of the models being developed.

Proportional Safeguards and AI Safety Levels (ASL)

Anthropic employs a system of AI Safety Level (ASL) Standards, which are graduated sets of security and safety measures. These standards increase in stringency as model capabilities grow, similar to Biosafety Levels used in biological research.

Currently, all Anthropic models operate under ASL-2 Standards, which the company states reflect current industry best practices. The updated policy defines two primary Capability Thresholds that would trigger a transition to higher ASL standards:

  • Autonomous AI Research and Development: If a model can independently conduct complex AI research tasks typically requiring human expertise, Anthropic requires elevated security standards (potentially ASL-4 or higher) and additional safety assurances.
  • Chemical, Biological, Radiological, and Nuclear (CBRN) weapons: If a model can meaningfully assist an individual with a basic technical background in creating or deploying CBRN weapons, the company requires ASL-3 standards, which include enhanced security and deployment controls.

ASL-3 Safeguards and Deployment Controls

When a model reaches the ASL-3 threshold, Anthropic implements a multi-layered approach to prevent misuse. Security measures include more robust protection of model weights and stricter internal access controls. Deployment safeguards include:

  • Real-time and asynchronous monitoring.
  • Rapid response protocols.
  • Rapid pre-deployment red teaming.

Implementation, Oversight, and Governance

To ensure the RSP is executed effectively, Anthropic has established several operational processes:

  • Capability Assessments: Routine evaluations to determine if current safeguards remain appropriate based on the defined Capability Thresholds.
  • Safeguard Assessments: Regular evaluations to verify that security and deployment safety measures meet the required ASL bar.
  • Documentation: The use of safety case methodologies, common in high-reliability industries, to document assessments and decision-making.
  • Internal and External Review: The framework is supported by internal stress-testing, an internal reporting process for safety issues, and the solicitation of external expert feedback.

Lessons Learned from Initial Implementation

Following a year of implementing the original RSP, Anthropic identified several procedural shortcomings, including completing evaluations three days later than scheduled and a lack of clarity regarding placeholder evaluations. The company concluded that these instances posed minimal risk but highlighted the need for greater flexibility in policy and improved compliance tracking.

Organizational Changes and Leadership

Anthropic has updated its leadership roles to manage the scaling of the RSP:

  • Responsible Scaling Officer: Co-Founder and Chief Science Officer Jared Kaplan succeeds Sam McCandlish in this role.
  • Head of Responsible Scaling: Anthropic is opening a new position for a Head of Responsible Scaling to coordinate the teams responsible for RSP compliance and iteration.

Several internal teams, including the Frontier Red Team, Trust & Safety, Security and Compliance, Alignment Science, and a dedicated RSP Team, contribute to the ongoing risk management process.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch