OpenAI Cyber Resilience Strategy

OpenAI has announced a comprehensive strategy to manage the dual-use risks of advancing AI cybersecurity capabilities. The company is deploying a layered safety stack and establishing new defensive tools and advisory bodies to ensure that as models become more capable, they provide a significant advantage to cybersecurity defenders over malicious actors.

Rapid Advancement of AI Cyber Capabilities

AI models are demonstrating a significant increase in their ability to perform cybersecurity tasks. OpenAI reports that capabilities measured through capture-the-flag (CTF) challenges improved from 27% on GPT-5 in August 2025 to 76% on GPT-5.1-Codex-Max in November 2025.

OpenAI is now preparing for models that reach "High" levels of cybersecurity capability according to its Preparedness Framework. This level is defined as models capable of developing working zero-day remote exploits against well-defended systems or providing meaningful assistance with complex, stealthy enterprise or industrial intrusion operations.

Layered Safety Stack for Mitigating Misuse

Because defensive and offensive cyber workflows often rely on the same underlying knowledge, OpenAI is employing a defense-in-depth approach to limit malicious uplift while maintaining utility for defenders.

Infrastructure and Monitoring

At the foundation, OpenAI utilizes a combination of:

  • Access controls
  • Infrastructure hardening
  • Egress controls
  • System-wide monitoring to detect potentially malicious activity

When unsafe activity is detected, the system may block output, route prompts to less capable models, or escalate the issue for human review based on severity and repeat behavior.

Model Training and Red Teaming

OpenAI is training frontier models to refuse requests that enable clear cyber abuse while remaining helpful for educational and defensive use cases. To validate these protections, the company works with expert red teaming organizations to conduct end-to-end evaluations, simulating determined and well-resourced adversaries to identify and close security gaps.

Ecosystem Initiatives for Defensive Resilience

OpenAI is launching several initiatives designed to accelerate responsible remediation and provide defenders with advanced tools.

Aardvark Agentic Security Researcher

Aardvark is an agentic security researcher currently in private beta. It is designed to help developers and security teams find and fix vulnerabilities at scale by scanning codebases and proposing patches. Aardvark has already identified novel CVEs in open-source software. OpenAI plans to provide free coverage to select non-commercial open-source repositories to secure the software supply chain.

Trusted Access and Advisory Councils

OpenAI is introducing a trusted access program to provide qualifying cyberdefense users with tiered access to enhanced capabilities in the latest models. Additionally, the company is establishing the Frontier Risk Council, an advisory group of experienced cyber defenders and security practitioners who will advise on the boundary between responsible capability and potential misuse.

Industry-Wide Collaboration

To address the risk of cyber misuse across all frontier models, OpenAI collaborates with other labs through the Frontier Model Forum. This nonprofit effort focuses on developing a shared understanding of threat models, identifying weaponization pathways, and identifying critical bottlenecks for threat actors. OpenAI is also engaging with external teams to develop independent cybersecurity evaluations to build a shared understanding of model capabilities across the industry.

Sources