Anthropic: The Case for Targeted AI Regulation
Anthropic argues that governments must implement targeted AI policy within the next eighteen months to prevent catastrophic risks as AI capabilities advance. The company asserts that judicious, narrowly-focused regulation can mitigate these risks without hindering scientific and commercial progress, avoiding the danger of "knee-jerk" regulation that is both ineffective and burdensome.
Accelerating Risks in Cyber and CBRN Domains
AI capabilities in mathematics, reasoning, and coding are advancing rapidly, increasing the potential for both misuse and autonomous destructive behavior. Anthropic identifies two primary areas of urgent concern:
Cyber Capabilities
Model performance on software engineering tasks has scaled significantly. On the SWE-bench benchmark, capabilities have grown from 1.96% (Claude 2, October 2023) to 13.5% (Devin, March 2024) and to 49% (Claude 3.5 Sonnet, October 2024). Anthropic's Frontier Red Team reports that current models already assist in cyber offense tasks, with next-generation models expected to be more effective due to their ability to plan long, multi-step tasks.
CBRN Risks
AI systems are demonstrating expert-level knowledge in biology and chemistry. The UK AI Safety Institute found that some models provide science answers on par with PhD-level experts. This is reflected in GPQA benchmark scores on the hardest section, which rose from 38.8% (November 2023) to 59.4% (Claude 3.5 Sonnet, June 2024) and to 77.3% (OpenAI o1, September 2024), compared to human experts at 81.2%.
The Responsible Scaling Policy (RSP) Framework
Anthropic utilizes a Responsible Scaling Policy (RSP) as an adaptive framework to identify and mitigate catastrophic risks. The RSP is built on two core principles:
- Proportionality: Safety and security measures increase in strength as models meet defined "capability thresholds."
- Iteration: Capabilities are regularly measured, and safety approaches are updated based on new findings.
According to Anthropic, RSPs provide three primary organizational benefits: they force early investment in security and safety evaluations, require the development of specific threat models, and encourage transparency regarding misuse mitigation practices.
Proposed Principles for Effective AI Regulation
Anthropic suggests that RSPs should serve as prototypes for enforceable government regulation. They propose three key elements for any effective regulatory framework:
- Transparency: Governments should require companies to publish RSP-like policies and risk evaluations for each new generation of AI systems, with a mechanism to verify the accuracy of these claims.
- Incentivizing Safety: Regulation should encourage the development of effective RSPs. This can be achieved by regulators identifying threat models, specifying standards, or comparing RSPs to drive a "race to the top" in safety practices.
- Simplicity and Focus: Regulations must be "surgical," avoiding unnecessary burdens or illogical rules that could create confusion or hinder the cause of catastrophic risk prevention.
Implementation and Strategic Considerations
Anthropic addresses several strategic questions regarding the deployment of these regulations:
Jurisdiction and Scope
While federal legislation in the U.S. is the ideal vehicle for uniform regulation and international negotiation, Anthropic suggests state regulation may serve as a necessary backstop if the federal process is too slow. Internationally, they argue that common safety and security approaches could reduce the global cost of business through standardization and mutual recognition.
Model-Based vs. Use-Case Regulation
Anthropic argues against "regulation by use case" for general-purpose models. Because models like Claude.ai or ChatGPT are general products capable of a wide range of tasks, it is more practical to regulate the fundamental properties and safety measures of the underlying base model, which is also the easiest point of regulation due to the resource-intensive nature of its training.
Impact on Innovation and Open Source
Anthropic contends that proportionate, flexible regulation will not significantly hinder innovation. They suggest that safety research often has spillover benefits for general AI science and that robust security prevents IP exfiltration. Regarding open-weights models, they argue that regulation should be based on empirically measured risks rather than the weight-sharing status of the model.
Sources
- OriginalThe case for targeted regulation
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch