Anthropic Building AI for Cyber Defenders
Anthropic has developed Claude Sonnet 4.5 with a specific focus on enhancing its cybersecurity capabilities for defenders, enabling the model to match or eclipse the performance of the larger Claude Opus 4.1 in discovering code vulnerabilities. This shift reflects an inflection point where frontier AI is increasingly capable of executing advanced cyber tasks, necessitating an acceleration of AI-driven defensive tools to maintain parity with AI-augmented attackers.
Enhanced Cyber Capabilities in Claude Sonnet 4.5
Claude Sonnet 4.5 was developed by dedicating research specifically to key defensive skills, such as code vulnerability discovery and patching, rather than relying solely on general model scaling. While cybersecurity skills often emerge as byproducts of general training, Anthropic focused on these areas to ensure defenders are equipped with high-performance tools that are faster and less expensive than previous frontier models.
Performance on Cybench
In evaluations using Cybench, a benchmark based on Capture-the-Flag (CTF) challenges, Claude Sonnet 4.5 demonstrated significant improvements over previous iterations:
- Success Rates: When given 10 attempts per task, Sonnet 4.5 achieved a 76.5% success rate, more than doubling the 35.9% success rate of Sonnet 3.7 (released February 2025).
- Efficiency: Sonnet 4.5 can achieve a higher probability of success in a single attempt than Opus 4.1 can in ten attempts.
- Complex Workflows: The model successfully solved a complex task involving network traffic analysis, malware extraction, and decryption in 38 minutes—a task estimated to take a skilled human at least an hour.
Performance on CyberGym
Using the CyberGym benchmark, which tests the ability to find both known and new vulnerabilities in open-source software, Claude Sonnet 4.5 set new performance records:
- Known Vulnerabilities: With a $2 API query limit per vulnerability, Sonnet 4.5 achieved a state-of-the-art score of 28.9%. When constraints were removed and 30 trials were allowed, the success rate rose to 66.7%.
- New Vulnerability Discovery: Sonnet 4.5 discovered new vulnerabilities in 5% of cases in a single trial, increasing to over 33% of projects when given 30 trials. This is a significant increase over Sonnet 4, which discovered vulnerabilities in approximately 2% of targets.
Research into Automated Patching
Anthropic is conducting preliminary research into the model's ability to generate and review patches to fix vulnerabilities. Patching is identified as a more difficult task than discovery because it requires surgical changes that preserve original functionality without explicit specifications.
Initial experiments showed that 15% of Claude Sonnet 4.5's generated patches were judged to be semantically equivalent to human-authored reference patches. Manual analysis confirmed that some of the highest-scoring patches were functionally identical to those merged into open-source software, suggesting that patch generation is an emergent capability that can be further refined through focused research.
Real-World Application and Industry Feedback
Anthropic collaborated with security organizations to test Claude Sonnet 4.5 on real-world challenges. Key feedback includes:
- HackerOne: Reported that Claude Sonnet 4.5 reduced average vulnerability intake time for their Hai security agents by 44% while improving accuracy by 25%.
- CrowdStrike: Noted the model's promise for red teaming and generating creative attack scenarios to accelerate the study of attacker tradecraft across endpoints, identity, cloud, and AI workloads.
Strategic Implications for Cybersecurity
Anthropic identifies a critical need for organizations to adopt AI for defense to avoid ceding the advantage to malicious actors. This is evidenced by the disruption of threat actors using AI for large-scale data extortion and complex espionage operations targeting critical telecommunications infrastructure, some consistent with Chinese APT operations.
To combat this, Anthropic suggests integrating AI into CI/CD pipelines for automated security reviews and expanding its use in Security Operations Center (SOC) automation, Security Information and Event Management (SIEM) analysis, and active defense. The goal is to transition AI's role in cybersecurity from a future concern to a present-day imperative for securing digital infrastructure by design.
Sources
- OriginalBuilding AI for cyber defenders
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch