Anthropic Responsible Scaling Policy
Anthropic has published its Responsible Scaling Policy (RSP), a framework of technical and organizational protocols designed to manage the catastrophic risks associated with the development of increasingly capable AI systems. The policy establishes a structured approach to scaling AI capabilities only when corresponding safety and security measures are proven effective.
The AI Safety Level (ASL) Framework
Anthropic's RSP centers on the AI Safety Level (ASL) system, which is modeled after the US government's biosafety level (BSL) standards. This framework requires that safety, security, and operational standards scale in proportion to a model's potential for catastrophic risk.
ASL Definitions and Criteria
- ASL-1: Systems that pose no meaningful catastrophic risk, such as a 2018 LLM or an AI system dedicated to playing chess.
- ASL-2: Systems showing early signs of dangerous capabilities (e.g., providing instructions on building bioweapons) but where the information is not yet useful due to insufficient reliability or because the information is already available via search engines. Current LLMs, including Claude, are categorized as ASL-2.
- ASL-3: Systems that substantially increase the risk of catastrophic misuse compared to non-AI baselines (such as textbooks or search engines) or systems that demonstrate low-level autonomous capabilities.
- ASL-4 and ASL-5+: These levels are not yet defined due to their distance from current system capabilities, but they are expected to involve qualitative escalations in autonomy and catastrophic misuse potential.
Safety Measures and Implementation
Safety requirements increase in rigor as a model moves up the ASL scale. ASL-2 measures align with Anthropic's previous White House commitments. ASL-3 requires stricter standards, including enhanced security requirements and a commitment not to deploy models if they exhibit meaningful catastrophic misuse risk during adversarial testing by world-class red-teamers.
For ASL-4, Anthropic commits to defining the measures before reaching ASL-3. These measures may involve solving currently unsolved research problems, such as using interpretability methods to mechanistically demonstrate that a model is unlikely to engage in catastrophic behaviors.
Operational and Business Impact
The RSP is designed to balance risk mitigation with the pursuit of beneficial AI applications. It creates a mechanism where training for more powerful models may be temporarily paused if AI scaling outstrips the ability to comply with safety procedures. This structure is intended to incentivize the resolution of safety issues to unlock further scaling.
From a business perspective, the RSP does not alter the current use or availability of Claude. Anthropic compares the policy to pre-market testing in the aviation or automotive industries, where safety is rigorously demonstrated before a product is released to the market.
Governance and Iteration
Anthropic's RSP has been formally approved by its board, and any changes to the policy must be board-approved following consultations with the Long Term Benefit Trust. The policy is an early iteration and is expected to evolve through rapid course correction as the field of AI progresses.
Anthropic acknowledged the influence of ARC Evals, specifically regarding their expertise in AI risk assessment and the development of the broader ARC Responsible Scaling Policy framework.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch