Anthropic Frontier Threats Red Teaming for AI Safety

Overview

Anthropic has introduced a "frontier threats red teaming" approach to identify and mitigate AI capabilities that could pose risks to national security, such as biosecurity and cybersecurity. This adversarial testing framework is designed to uncover underlying model capabilities that could be exploited by bad actors, moving beyond general crowd-sourced red teaming to specialized, expert-led evaluations.

Methodology for Frontier Threats Red Teaming

Frontier threats red teaming requires intensive investments in subject matter expertise and time to uncover latent model capabilities. Anthropic's process involves the following key components:

  • Expert Collaboration: The process begins by partnering with domain experts with decades of experience to define threat models. These models identify what information is dangerous, how that information is combined to create harm, and the required accuracy and frequency for the information to be dangerous.
  • Intensive Probing: Subject matter and LLM experts spend substantial time (100+ hours) interacting with models to probe for capabilities. This includes learning how to "jailbreak" the models to bypass safety filters.
  • Automated Evaluations: A primary goal is to develop new, automated evaluations and tooling based on expert knowledge to make the testing process repeatable and scalable.
  • Secure Infrastructure: Because the information uncovered is highly sensitive, this work is conducted through partnerships with trusted third parties and strong information security protections.

Findings from Biosecurity Red Teaming

Anthropic conducted a test project focusing on biological risks, spending over 150 hours with biosecurity experts using a secure interface without public-facing trust and safety monitoring.

Key Risks Identified

  • Expert-Level Knowledge: Frontier models can occasionally produce sophisticated, accurate, and detailed knowledge at an expert level. This capability increases as models grow larger.
  • Acceleration of Harm: Unmitigated LLMs could accelerate a bad actor's ability to misuse biology compared to using a search engine alone, enabling tasks that would otherwise be impossible. While these effects are currently small, they are growing rapidly.
  • Near-Term Timeline: Anthropic estimates that if left unmitigated, these biological risks could be actualized within the next two to three years, rather than five or more.

Implemented Mitigations

Anthropic found that these risks can be substantially reduced through specific interventions:

  • Training Process Adjustments: Changes in the training process, such as those used in Constitutional AI, help models better distinguish between harmful and harmless biological uses.
  • Classifier-Based Filters: These filters make it harder for actors to obtain the chained, expert-level information necessary to cause harm. These mitigations are currently deployed in Anthropic's public-facing frontier models.

Future Research and Industry Implications

Anthropic argues that the current state of frontier models provides a "warning shot" for the industry. The company is focusing on the following future directions:

  • Comparative Speedup Analysis: Measuring the speedup LLMs provide toward producing harm compared to traditional search engines across next-generation, tool-using, and multimodal models.
  • Open-Weight Model Risks: Anthropic warns that bad actors could extract harmful biological capabilities from smaller, fine-tuned models adapted from the weights of openly available, sufficiently capable base models.
  • Collaborative Disclosure: Anthropic is establishing a disclosure process to report risks and mitigations to other labs and government agencies.
  • Third-Party Evaluation: The company advocates for the creation of impartial third-party organizations to conduct national security evaluations with appropriate safeguards for sensitive information.

Scaling the Research Agenda

Anthropic is expanding its frontier threats red teaming team to experiment with future capabilities and build scalable evaluations. This framework is intended to be applied to other long-term risks, such as model deception, by identifying future capabilities that should be avoided and building alignment techniques before those capabilities emerge.

Sources

Related