Anthropic expands model safety bug bounty program
TL;DR
Anthropic announced an expanded, invite‑only bug bounty program that rewards up to $15,000 for discovering universal jailbreak attacks on its next‑generation AI safety mitigations, aiming to protect high‑risk domains such as CBRN and cybersecurity.
Program Overview
Anthropic is launching a new bug bounty initiative focused on universal jailbreak attacks—exploits that consistently bypass AI safety guardrails across many topics. The program targets vulnerabilities that could affect critical, high‑risk areas like chemical, biological, radiological, and nuclear (CBRN) threats and cybersecurity.
How the Initiative Works
- Early Access: Selected participants receive pre‑release access to Anthropic’s latest safety mitigation system, allowing them to test for weaknesses in a controlled environment.
- Scope and Rewards: Bounties of up to $15,000 are offered for novel universal jailbreaks that could compromise high‑risk domains. A jailbreak is considered universal if it enables the model to answer a defined set of harmful questions across topics.
- Evaluation Process: Detailed instructions and feedback will be provided to participants. Submissions are evaluated for novelty and impact, with timely, constructive feedback promised during the invite‑only phase.
Eligibility and Application
The initial phase is invite‑only and run in partnership with HackerOne. Researchers with proven experience in AI security or language‑model jailbreaks can apply via the application form by Friday, August 16. Selected applicants will be contacted in the fall for participation.
Ongoing Reporting Mechanism
Anthropic continues to accept safety issue reports for its currently deployed models. Researchers can email detailed findings to usersafety@anthropic.com. The company references its Responsible Disclosure Policy for guidance.
Alignment with Global AI Safety Commitments
The bounty program supports Anthropic’s commitments under the White House Voluntary AI Commitments and the G7 Code of Conduct for Organizations Developing Advanced AI Systems. By focusing on universal jailbreak mitigation, Anthropic aims to ensure that safety measures evolve in step with rapidly advancing AI capabilities.
Related Announcements
- Improving Fable 5's biology safeguards – see the linked announcement for details on domain‑specific safety enhancements.
- Mariano‑Florentino (Tino) Cuéllar joins Anthropic as Chief Global Affairs Officer – a leadership update relevant to the company’s global policy engagement.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch