OpenAI Astra: Addressing Critical Cyber Capabilities
OpenAI has announced that internal evaluations of its upcoming Astra model indicate significant advancements in agentic coding and cybersecurity. These results, lead OpenAI to conclude that it cannot rule out that Astra has reached the "Critical" cyber capabilities threshold as defined by the company's Preparedness Framework.
Critical Cybersecurity Threshold Definition
A model is classified as reaching the Critical cybersecurity threshold when it can perform the following without human intervention:
- Identifying and developing functional zero-day exploits of all severity levels in many hardened real-world critical systems.
- Devising and executing end-to-end novel strategies for cyberattacks against hardened targets based on only a high-level goal.
While evaluations are ongoing, preliminary results for Astra indicate performance strong enough that this Critical level cannot be ruled out. OpenAI explicitly stated that Astra was not involved in the exploiting of Hugging Face.
Security Controls and Mitigation Steps
To manage the risks associated with these capabilities, OpenAI is scaling up robustness testing of safeguards and security controls. The company is implementing several internal security measures for higher-capability models:
- Infrastructure Protections: Implementation of isolated testing environments, restricted network and tool access, enhanced encryption and protection of model weights, and sandboxed execution.
- Activity Pauses: Internal activities involving Astra that do not meet these strengthened security requirements are currently paused.
- Universal Monitoring: Implementation of monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. These monitors analyze the model's Chain of Thought to trigger security responses and interrupt high-risk activity.
- External Collaboration: OpenAI will work with government agencies and select AI safety organizations to test the model's capabilities.
- Third-Party Guidance: The company will provide recommended security controls to third-party testing partners to ensure high-risk evaluations and workloads are run safely.
Context and Preparedness Framework
OpenAI's Preparedness Framework, published in December 2023, serves as a guide for identifying progress in capabilities across domains such as biology, chemistry, cybersecurity, and AI self-improvement. Previous models, including GPT-5.6-Sol, were evaluated and assessed at the "High" (rather than "Critical") threshold for frontier cyber capabilities.
This response follows a similar protocol to a June 2025 announcement regarding models approaching the high capability threshold for biology, where OpenAI strengthened safeguards and expanded testing in order to deploy capabilities responsibly.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch