Anthropic Frontier Red Team Progress Report

Anthropic's Frontier Red Team has observed that AI models are displaying "early warning" signs of rapid progress in dual-use capabilities, specifically in cybersecurity and biology. While models are approaching or exceeding undergraduate-level skills in cyber and expert-level knowledge in some biological areas, they currently fall short of the thresholds required to generate substantially elevated risks to national security.

Rapid Advancement in Cybersecurity Capabilities

AI capabilities in the cyber domain have transitioned from basic to undergraduate-level proficiency within a single year. This progress is evidenced by performance on Capture The Flag (CTF) exercises and public benchmarks:

  • CTF Performance: In less than a year, Claude progressed from solving less than 25% of Intercode CTF challenges (designed for high schoolers) to solving nearly all of them.
  • Cybench Benchmark: Claude 3.7 Sonnet solves approximately one-third of Cybench challenges within five attempts, a significant increase from the 5% success rate of frontier models from the previous year.
  • Task Diversity: Improvements are noted across multiple categories, including discovering and exploiting vulnerabilities in remote servers ("pwn"), web applications ("web"), and cryptographic protocols ("crypto").

Despite these gains, Claude still lags behind expert humans in reverse engineering binary executables and performing reconnaissance and exploitation within network environments without assistance.

Network-Scale Cyber Operations

In collaboration with Carnegie Mellon University, Anthropic tested models on realistic cyber ranges consisting of approximately 50 hosts. While models cannot yet autonomously succeed in these environments, they can replicate large-scale thefts of personally identifiable information (similar to the Equifax breach) when equipped with specialized software tools like Incalmo. This infrastructure allows Anthropic to monitor for the moment autonomous capabilities improve.

Progress and Gaps in Biosecurity Knowledge

Claude has demonstrated swift progress in biological understanding, moving from underperforming world-class virology experts to comfortably exceeding that baseline on the VCT evaluation (troubleshooting virology tasks) within one year.

Biological Expertise vs. Practical Risk

Biological capabilities remain uneven across different tasks:

  • Strengths: Claude 3.7 Sonnet approaches human expert baselines in understanding biology protocols and manipulating DNA and protein sequences, and it exceeds human expert baselines in cloning workflows.
  • Weaknesses: Models remain inferior to human experts at interpreting scientific figures.

Regarding weaponization risk, small controlled studies indicate that while the most recent models provide some uplift to novices compared to those with only internet access, the resulting plans still contained critical mistakes that would lead to real-world failure. Expert red-teaming concluded that the number of critical failures in planning is currently too high to enable the end-to-end execution of an attack by a novice malicious actor.

Strategic Partnerships and Safety Frameworks

Anthropic utilizes a Responsible Scaling Policy (RSP) to commit to capability thresholds that trigger increased security levels. To validate these risks, the lab employs several external partnerships:

  • AI Safety/Security Institutes: Models undergo pre-deployment testing by the US and UK AI Safety Institutes (AISI), which informed the AI Safety Level (ASL) determination for Claude 3.7 Sonnet.
  • National Nuclear Security Administration (NNSA): Anthropic partnered with the US Department of Energy's NNSA to evaluate Claude in a classified environment for nuclear and radiological risk knowledge.

Future Outlook and Risk Mitigation

Anthropic is scaling up to more frequent, automated evaluations to detect risks earlier. The lab notes that as models improve in "extended thinking," the need for external toolkits like Incalmo may decrease as models become more capable "out of the box."

Based on current biological research, Anthropic believes its models are approaching the capabilities threshold that requires AI Safety Level 3 (ASL-3) safeguards, prompting increased investment in the necessary security measures.

Sources

Related