Anthropic Third-Party Testing AI Policy Proposal
TL;DR
Anthropic advocates for the establishment of a third-party testing regime for frontier AI systems to validate safety claims and prevent societal harm. This approach aims to create a broadly trusted oversight infrastructure involving industry, government, and academia to manage risks that self-governance alone cannot address.
The Necessity of Third-Party Oversight
Frontier AI systems—large-scale generative models requiring substantial computational resources—function as "everything machines" that can be adapted to numerous downstream use cases. Because these systems inherit the capabilities and weaknesses of the base model, they do not fit neatly into existing sector-specific regulatory frameworks.
Anthropic argues that while self-governance frameworks, such as their own Responsible Scaling Policy (RSP), are essential, they are insufficient because they rely on decisions made by single private actors. A robust third-party regime is necessary to:
- Validate Safety: Provide independent verification of model behavior regarding election integrity, harmful discrimination, and national security misuse.
- Prevent Catastrophic Accidents: Identify emergent, autonomous behaviors, such as the insertion of vulnerabilities into code or actions that contradict human intentions.
- Avoid "Knee-Jerk" Regulation: Proactively design effective regulation to prevent major incidents that could lead to stifling and ineffective legislative reactions.
Framework for a Robust Testing Regime
An effective testing regime must be precisely scoped to apply only to the most computationally intensive, large-scale systems to avoid disadvantaging smaller companies. Anthropic proposes the following structural components:
Two-Stage Testing Process
- Automated Stage: A fast, broad testing phase biased toward avoiding false negatives to spot potential problems quickly.
- Expert Stage: A thorough secondary test utilizing expert human-led elicitation for flagged issues.
Diverse Testing Ecosystem
Testing should be conducted by a variety of actors to ensure legitimacy and competition:
- Private Companies: Subcontracted firms (e.g., Gryphon Scientific) or auditing firms similar to those used in financial accounting.
- Universities: Academic institutions administering testing initiatives, potentially supervised by government bodies.
- Governments: Direct testing for specific, legally mandated areas, particularly those involving classified national security risks.
Key Requirements for Implementation
- Shared Standards: A consensus across industry, government, and academia on what a safety testing framework should include.
- Government Funding: Increased resources for government agencies (e.g., NIST, US AI Safety Institute) to fund the technical work of building and analyzing tests.
- Public Infrastructure: The creation of "National Research Clouds" to allow academia and government to develop independent capacity to study and test frontier systems.
National Security and High-Risk Capabilities
Anthropic identifies national security as a primary area where third-party testing is critical. Specifically, systems should be tested for capabilities that could compromise national security, such as the ability to accelerate the creation of bioweapons or execute complex cyberattacks. If such capabilities are discovered, Anthropic suggests mitigations such as removing the capabilities from broadly deployed models or implementing "know your customer" (KYC) regimes.
Addressing Open-Source Models and Regulatory Capture
Openly Disseminated Models
Anthropic acknowledges the value of open-source research but suggests that as models become more capable, the culture of full open dissemination may conflict with societal safety. They propose that if a model is released openly, it must be resilient to fine-tuning that could enable misuse. Third-party testing is presented as the only legitimate way to define unacceptable misuses and apply those standards consistently across both proprietary and open-weight models.
Mitigating Regulatory Capture
To prevent well-resourced AI companies from using regulation to stifle competition, Anthropic advocates for testing capacity that exists independently of large corporations. By focusing on measurement and third-party capacity rather than high-cost compliance consortia, the policy aims to maintain a level playing field for smaller developers.
Anthropic's Commitment to Support
To help develop this regime, Anthropic commits to:
- Prototyping testing regimes via their Responsible Scaling Policy (RSP).
- Testing third-party assessments through contractors and government partners.
- Expanding frontier red teaming to clarify risks and mitigations.
- Advocating for government funding for agencies like NIST and the development of national research infrastructure.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch