Anthropic AI Accountability Recommendations for NTIA
Anthropic has submitted a formal response to the National Telecommunications and Information Administration’s (NTIA) Request for Comment on AI Accountability. The proposal outlines a structured approach to evaluating advanced artificial intelligence systems to ensure safety and accountability as model capabilities increase.
Standardized Evaluations and Funding
Anthropic recommends increasing government funding for AI model evaluation research to create rigorous, standardized benchmarks. Because developing these evaluations is resource-intensive, public funding is seen as a critical driver for progress.
Key proposals for evaluation include:
- Disclosure Requirements: AI companies should be mandated to disclose evaluation methods and results to regulators, though not necessarily to the public, to protect intellectual property (IP) and confidential information.
- Industry Standards: Government agencies, specifically the National Institute of Standards and Technology (NIST), should collaborate with industry to establish best practices and benchmarks for evaluating model capabilities, limitations, and risks.
Risk-Responsive Assessments and Thresholds
To ensure regulation is proportionate to the risk posed by a model, Anthropic proposes the development of standard capability evaluations targeting critical risks such as autonomy and deception.
This framework suggests a risk-based deployment pipeline:
- Risk Threshold Establishment: Research and funding should be used to define a specific risk threshold.
- Below Threshold: Models falling below this threshold can be deployed after verifying compliance with existing safety standards.
- Above Threshold: Models exceeding the risk threshold that lack sufficient safety mitigations must have their deployment halted, oversight strengthened, and regulators notified to determine necessary safeguards.
Pre-registration of Large Training Runs
Anthropic proposes a confidential registry for AI developers to pre-register details of large training runs with their national government before training begins. This registry would track:
- Model specifications and type
- Compute infrastructure
- Intended training completion date
- Safety plans
This process is intended to ensure regulators are aware of potential risks before a model is created, with aggregated data protected by high cybersecurity and privacy standards.
External Auditing and Red Teaming
To validate safety claims, Anthropic recommends mandating external red teaming and the empowerment of third-party auditors.
Third-Party Auditor Requirements
Auditors must possess three specific traits to be effective:
- Technical Literacy: Deep machine learning experience.
- Security Consciousness: The ability to protect sensitive IP that could pose national security threats.
- Flexibility: The capability to conduct lightweight assessments that identify threats without hindering US competitiveness.
External Red Teaming
External red teaming should be a precondition for releasing advanced AI systems. Anthropic suggests this be managed either through a centralized body like NIST or a decentralized approach via researcher API access. The company emphasizes that high-quality red teaming options must be established before this becomes a mandate, as the necessary talent currently resides primarily within private labs.
Interpretability and Industry Collaboration
Anthropic advocates for increased funding for interpretability research through government grants for universities, nonprofits, and companies. This allows safety work to be conducted on smaller models outside of frontier labs.
Additionally, the proposal addresses two systemic hurdles:
- Regulatory Feasibility: Anthropic notes that regulations demanding fully interpretable models are currently infeasible but may become possible as research advances.
- Antitrust Clarity: Regulators should provide clear guidance on how AI companies can coordinate on safety in the public interest without violating antitrust laws, reducing legal uncertainty for industry collaboration.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch