OpenAI Preparedness Framework Update

OpenAI has updated its Preparedness Framework, a structured process for tracking and preparing for advanced AI capabilities that could introduce risks of severe harm. This update focuses on sharpening risk prioritization, strengthening requirements for risk minimization, and providing clearer operational guidance for the evaluation and governance of safeguards.

Risk Prioritization and Criteria

OpenAI uses a structured risk assessment process to categorize capabilities based on whether they could lead to severe harm. A capability is prioritized for advance preparation if it meets five specific criteria: it must be plausible, measurable, severe, net new, and either instantaneous or irremediable.

Capability Categories

The framework divides AI capabilities into two primary categories to manage dual-use risks and emerging threats:

Tracked Categories

These are established areas with mature evaluations and ongoing safeguards. They include:

  • Biological and Chemical capabilities
  • Cybersecurity capabilities
  • AI Self-improvement capabilities

Research Categories

These areas are identified as potentially posing risks of severe harm but do not yet meet the criteria to be Tracked Categories. OpenAI is currently developing threat models and advanced evaluations for these areas, which include:

  • Long-range Autonomy
  • Sandbagging (intentionally underperforming)
  • Autonomous Replication and Adaptation
  • Undermining Safeguards
  • Nuclear and Radiological

Persuasion risks are managed outside of this framework via the Model Spec, restrictions on political campaigning and lobbying, and investigations into influence operations.

Capability Levels and Operational Commitments

OpenAI has streamlined its capability levels into two thresholds that trigger specific operational requirements:

  • High capability: Capabilities that could amplify existing pathways to severe harm. Systems reaching this level must have safeguards that "sufficiently minimize" the associated risk before they are deployed.
  • Critical capability: Capabilities that could introduce unprecedented new pathways to severe harm. These systems require safeguards that sufficiently minimize associated risks during both development and deployment.

The Safety Advisory Group (SAG), a cross-functional team of internal safety leaders, reviews these safeguards and provides recommendations to OpenAI Leadership for final deployment decisions. The SAG's guidance ranges from approving deployment to requesting further protections or evaluations.

Evaluation Scaling and Reporting

To keep pace with faster model improvement cycles—often driven by reasoning advances rather than major training runs—OpenAI has implemented a suite of automated evaluations to scale testing. These are supplemented by expert-led "deep dives" to ensure accuracy.

Reporting has also been refined into two distinct types of documentation:

  • Capabilities Reports: These assess whether a model has crossed a risk threshold.
  • Safeguards Reports: These provide detailed information on the design and verification of safeguards, following a "defense in depth" principle.

Response to the Frontier Landscape

OpenAI states that if another frontier AI developer releases a high-risk system without comparable safeguards, OpenAI may adjust its own requirements. Such adjustments would only occur after rigorous confirmation that the risk landscape has changed, public acknowledgment of the following adjustment, and an assessment that the change does not meaningfully increase the overall risk of severe harm while maintaining a protective level of safeguards.

Sources