OpenAI Frontier Risk and Preparedness Framework

OpenAI is implementing a structured approach to catastrophic risk preparedness by creating a specialized Preparedness team and a Risk-Informed Development Policy (RDP). This initiative aims to proactively identify and mitigate severe risks posed by frontier AI models—systems that exceed the capabilities of current advanced models—to ensure the safe development of artificial general intelligence (AGI).

The Preparedness Team and Risk Categories

The Preparedness team, led by Aleksander Madry, integrates capability assessment, evaluations, and internal red teaming to track and forecast catastrophic risks. The team focuses on four primary risk categories:

  • Chemical, Biological, Radiological, and Nuclear (CBRN) threats: Risks related to the creation or deployment of hazardous materials.
  • Cybersecurity: The potential for AI to enhance the scale or sophistication of cyberattacks.
  • Autonomous Replication and Adaptation (ARA): Risks associated with AI systems that can self-replicate or evolve independently.
  • Individualized Persuasion: The use of AI to manipulate or persuade individuals at scale.

Risk-Informed Development Policy (RDP)

To govern the development process, OpenAI is maintaining a Risk-Informed Development Policy (RDP). This policy establishes a formal framework for:

  1. Evaluation and Monitoring: Developing rigorous methods to evaluate the capabilities of frontier models.
  2. Protective Actions: Creating a spectrum of responses and safeguards based on the identified risk level.
  3. Governance: Establishing a structure for accountability and oversight throughout the development lifecycle.

The RDP is designed to complement existing risk mitigation and alignment work conducted both before and after model deployment.

AI Preparedness Challenge and "Unknown Unknowns"

As part of its "unknown unknowns" work stream, OpenAI conducted a Preparedness Challenge to surface non-obvious risk areas. The challenge awarded $25,000 in API credits to ten top submissions that identified plausible but unique risk scenarios.

Key Findings from the Challenge

Analysis of the submissions revealed that approximately 70% of entrants identified the enhancement of malicious persuasive capabilities—including online radicalization, polarization, and political influence—as a primary threat. Other identified risks included:

  • Financial Stability: Precipitating financial crises in strategically important countries.
  • Information Security: Increasing the likelihood of reverse-engineering classified information or identifying private information in public settings.
  • Public Safety: Disrupting flight paths via radio frequencies or interfering with medical dosages.
  • Cybercrime: Scaling cyberattacks that break computers for ransom.
  • Social Engineering: Identifying targets for blackmail and scams.

OpenAI noted that these submissions were evaluated based on technical rigor, uniqueness, scale of potential damage, and clarity, specifically highlighting cases where AI tools provided a distinct advantage over non-AI methods.

Sources