Anthropic Responsible Scaling Policy Version 3.0
Anthropic has released Version 3.0 of its Responsible Scaling Policy (RSP), a voluntary framework designed to mitigate catastrophic risks as AI capabilities advance. The update shifts the policy from a strict set of "if-then" capability thresholds to a more pragmatic approach that distinguishes between what Anthropic can achieve unilaterally and what requires collective industry or government action.
Structural Shift: Unilateral vs. Multilateral Mitigations
Anthropic has restructured the RSP to separate its internal operational plans from its broader safety recommendations for the AI industry. This change addresses the reality that some high-level safeguards are impossible for a single company to implement alone.
- Unilateral Commitments: The RSP now outlines specific mitigations that Anthropic will pursue independently of other industry players.
- Industry Recommendations: The policy includes an ambitious capabilities-to-mitigations map intended as a blueprint for the entire AI industry to manage advanced risks collectively.
The Frontier Safety Roadmap
Version 3.0 introduces the Frontier Safety Roadmap, a public document detailing ambitious but achievable goals across four key domains: Security, Alignment, Safeguards, and Policy. While these goals are non-binding, they serve as an internal forcing function and a public benchmark for progress.
Key objectives within the roadmap include:
- Information Security: Launching "moonshot R&D" projects to achieve unprecedented levels of security.
- Automated Red-Teaming: Developing automated red-teaming methods that exceed the effectiveness of human bug bounty participants.
- Constitutional Adherence: Implementing systematic measures to ensure models behave according to Anthropic's constitution.
- Internal Oversight: Establishing centralized records of critical development activities to detect insider threats (both human and AI) and security vulnerabilities.
- Regulatory Guidance: Publishing a policy roadmap proposing a "regulatory ladder" where policies scale in proportion to increasing risk.
Risk Reports and External Review
To increase transparency, Anthropic is implementing a systematic practice of publishing Risk Reports every 3 to 6 months. These reports provide a comprehensive safety profile of current models, detailing the intersection of capabilities, threat models, and active mitigations.
External Validation
In specific circumstances, the RSP now requires external review of these Risk Reports. Anthropic will appoint third-party experts—free of conflicts of interest and familiar with AI safety research—to access unredacted reports and subject the company's reasoning and decision-making to public review. While current models do not yet trigger this requirement, Anthropic is currently piloting this process.
Retrospective: Successes and Failures of RSP v1 and v2
Anthropic's transition to Version 3.0 is informed by a two-and-a-half-year assessment of its previous policies.
What Worked
- Internal Incentives: The RSP successfully compelled the development of stronger safeguards, such as the input and output classifiers used to comply with ASL-3 (AI Safety Level 3) standards.
- Industry Influence: The framework encouraged other labs, including OpenAI and Google DeepMind, to adopt similar safety frameworks and informed early AI policy in jurisdictions like California (SB 53), New York (RAISE Act), and the EU (AI Act's Codes of Practice).
What Did Not Work
- Capability Thresholds: The "zone of ambiguity" made it difficult to determine exactly when a model had passed a specific capability threshold. For example, in biological risks, models may pass simple tests but fail to provide definitive evidence of high risk, making it difficult to build a multilateral case for action.
- Government Pace: Government action on AI safety has moved slower than anticipated, with political climates often prioritizing economic growth and competitiveness over safety regulation.
- Unilateral Limits: Some high-level security standards (such as the "SL5" standard mentioned in a RAND report) are currently deemed impossible to achieve without assistance from the national security community.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch