Anthropic Bug Bounty Program for ASL-3 Safety Defenses
Anthropic has launched a bug bounty program to stress-test safety classifiers designed to prevent the misuse of AI for chemical, biological, radiological, and nuclear (CBRN) weapons. This initiative is part of Anthropic's effort to meet the AI Safety Level-3 (ASL-3) Deployment Standard as defined in its Responsible Scaling Policy.
Stress-Testing Constitutional Classifiers
Anthropic is utilizing a bug bounty program, managed in partnership with HackerOne, to test an updated version of its Constitutional Classifiers system. Constitutional Classifiers are designed to guard against jailbreaks that could elicit information related to CBRN weapons by following a set of principles that define allowed and disallowed content based on specific harms.
For the initial phase of the program, participants were granted early access to test these classifiers on Claude 3.7 Sonnet. The program offered rewards of up to $25,000 for verified universal jailbreaks—vulnerabilities that consistently bypass safety measures across multiple topics—specifically those that enable misuse on CBRN-related topics.
Alignment with Responsible Scaling Policy (RSP)
The bug bounty program is a direct application of the Responsible Scaling Policy (RSP) framework, which governs how Anthropic develops and deploying increasingly capable models. The company states that as models become more capable, they may require the advanced security and safety protections outlined in the ASL-3 standard. This initiative serves as a method to iterate and stress-test those ASL-3 safeguards before they are deployed publicly.
Program Evolution and Expansion
Following the conclusion of the initial program on May 18, 2025, Anthropic announced an update on May 22, 2025, transitioning the initiative to a new phase. This new invite-only program focuses on:
- Claude Opus 4 Testing: Stress-testing the Constitutional Classifiers system on the new Claude Opus 4 model.
- General Safety Systems: Testing other safety systems currently under development.
- Public Report Collection: Anthropic is now accepting reports of universal jailbreaks for ASL-3 uses of concern—specifically those eliciting information related to biological threats—that are found on public platforms or forums, such as social media.
Participation Requirements
Participation in the bug bounty programs is invite-only to ensure timely feedback and response. Anthropic encourages experienced red teamers and researchers with demonstrated expertise in identifying language model jailbreaks to apply for an invitation.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch