Anthropic Third-Party Model Evaluation Initiative

Anthropic has introduced a new initiative to fund the development of third-party evaluations designed to measure advanced capabilities and safety risks in AI models. This effort aims to address the shortage of high-quality, safety-relevant evaluations and provide the broader AI ecosystem with tools to assess model capabilities and risks more effectively.

Priority Focus Areas for Evaluation Development

Anthropic is prioritizing funding for evaluations in three primary categories: AI Safety Level (ASL) assessments, advanced capability and safety metrics, and the infrastructure required to build these evaluations.

AI Safety Level (ASL) Assessments

These evaluations are designed to measure the AI Safety Levels defined in Anthropic's Responsible Scaling Policy, which dictate the safety and security requirements for models based on their capabilities. Key focus areas include:

  • Cybersecurity: Assessments focusing on the "cyber kill chain," including vulnerability discovery, exploit development, and lateral movement. Anthropic is seeking evaluations that resemble novel Capture The Flag (CTF) challenges without public solutions to avoid measuring simple memorization.
  • CBRN Risks: Evaluations targeting Chemical, Biological, Radiological, and Nuclear risks, specifically the potential for models to assist non-experts or experts in creating threats or designing novel, more harmful CBRN threats.
  • Model Autonomy: Measuring proficiency in AI research and development (at junior to expert levels), advanced autonomous behaviors, and the ability to self-replicate or adapt by acquiring resources or exfiltrating weights.
  • National Security Risks: Development of early warning systems to identify complex emerging risks to defense and intelligence operations for both state and non-state actors.
  • Social Manipulation: Measuring the amplification of persuasion-related threats, such as disinformation and manipulation, and isolating the model's unique contribution to these risks.
  • Misalignment Risks: Monitoring for dangerous goals, motivations, and deceptive behaviors that could allow a model to bypass security or sabotage an organization.

Advanced Capability and Safety Metrics

This category focuses on broader metrics to understand model strengths and potential risks beyond the ASL framework:

  • Advanced Science: Funding for tens of thousands of new questions and tasks that challenge graduate-level knowledge, including knowledge synthesis, autonomous research execution, novel hypothesis generation, and tacit knowledge acquired through lab apprenticeship.
  • Harmfulness and Refusals: Improving classifiers to better distinguish between dual-use and non-dual-use information and identify harmful CBRN or cyber-related outputs.
  • Multilingual Evaluations: Expanding capability benchmarks to support a wider range of global languages.
  • Societal Impacts: Rigorous assessments of harmful biases, discrimination, over-reliance, psychological influence, and economic impacts.

Infrastructure, Tools, and Methods

Anthropic is seeking to improve the efficiency of evaluation development through:

  • No-code Platforms: Tools that allow subject-matter experts without coding skills to develop and export robust evaluations.
  • Model Grading: Developing diverse datasets (including questions, sample answers, ground truth scores, and rubrics) to improve the reliability of models acting as graders for other models.
  • Uplift Trials: Supporting the creation of networks for high-quality study populations and tooling to run large-scale controlled trials that compare task performance with and without model access.

Principles of High-Quality Evaluations

Anthropic identifies ten key characteristics of effective evaluations to avoid common pitfalls and ensure the results are indicative of real-world risk:

  1. Sufficiently Difficult: Must measure capabilities relevant to ASL-3 or ASL-4 or human-expert level behavior.
  2. Not in Training Data: Must capture generalization rather than memorization.
  3. Efficient and Scalable: Optimized for automation and easy deployment.
  4. High Volume: Preference for evaluations with 1,000 to 10,000 tasks, though high-quality low-volume tasks are also valuable.
  5. Domain Expertise: Developed or reviewed by subject-matter experts.
  6. Diverse Formats: Use of task-based evaluations, model-graded evaluations, or human trials rather than just multiple choice.
  7. Expert Baselines: Comparison of model performance against human experts.
  8. Documentation and Reproducibility: Use of standards like Inspect or METR.
  9. Iterative Development: Starting with small samples to refine the evaluation before scaling.
  10. Safety-Relevant Threat Modeling: High scores should correlate with a believable risk of a major incident.

Proposal Submission and Process

Interested parties can submit proposals via Anthropic's application form. Submissions are reviewed on a rolling basis, with various funding options available. Selected applicants will have the opportunity to collaborate with Anthropic's domain experts from the Frontier Red Team, Finetuning, and Trust & Safety teams to refine their evaluations.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch