OpenAI Safety Bug Bounty Program Launch
OpenAI has launched a public Safety Bug Bounty program to identify and mitigate AI abuse and safety risks across its products. This program is designed to complement the existing Security Bug Bounty by rewarding researchers for finding issues that pose meaningful safety risks, even if they do not qualify as traditional security vulnerabilities.
Agentic Risks and MCP
OpenAI is prioritizing the identification of risks associated with agentic AI products, including the Model Context Protocol (MCP). The program accepts reports on the following agentic vulnerabilities:
- Third-party prompt injection and data exfiltration: Reports must demonstrate that attacker text can reliably hijack a victim's agent (such as ChatGPT Agent or Browser) to perform harmful actions or leak sensitive user information. These behaviors must be reproducible at least 50% of the time.
- Unauthorized actions at scale: Cases where an agentic OpenAI product performs a disallowed action on OpenAI's own website at scale.
- General harmful actions: Any agentic behavior that results in plausible and material harm, even if not explicitly listed in the other categories.
Researchers testing for MCP risks must comply with the third-party terms of service.
OpenAI Proprietary Information
The program targets the exposure of OpenAI's internal data. In-scope issues include:
- Reasoning-related proprietary information: Model generations that return proprietary information regarding the model's reasoning processes.
- General proprietary leaks: Vulnerabilities that expose other OpenAI proprietary information.
Account and Platform Integrity
The program focuses on maintaining the integrity of the system's access controls and trust signals:
- Integrity signal bypasses: Vulnerabilities that allow users to bypass anti-automation controls, manipulate account trust signals, or evade account restrictions, suspensions, or bans.
- Permission-based issues: Issues allowing unauthorized access to features, data, or functionalities are redirected to the Security Bug Bounty program.
Scope Limitations and Exclusions
OpenAI specifies that general jailbreaks and content-policy bypasses are out of scope for this program. Specifically, "jailbreaks" that result in the model using rude language or returning information easily found via search engines are not eligible for rewards.
However, OpenAI may consider rewards on a case-by-case basis for flaws that facilitate direct paths to user harm with actionable remediation steps. Additionally, the lab periodically runs private campaigns for specific harm types, such as Biorisk content issues in ChatGPT Agent and GPT-5.
Participation Process
Researchers can apply to participate in the program through the Bugcrowd platform. Submissions are triaged by OpenAI's Safety and Security Bug Bounty teams, and reports may be rerouted between the two programs depending on the scope and ownership of the issue.