OpenAI o3-mini System Card

OpenAI o3-mini utilizes large-scale reinforcement learning and chain-of-thought reasoning to achieve state-of-the-art safety performance on benchmarks for illicit advice, stereotyped responses, and jailbreak resistance. This approach allows the model to perform deliberative alignment by reasoning about safety policies in context before responding to unsafe prompts.

Safety Performance and Deliberative Alignment

OpenAI o3-mini achieves parity with state-of-the-art performance on key safety benchmarks, specifically in reducing the generation of illicit advice, avoiding stereotyped responses, and resisting known jailbreaks. This is enabled by training the model to incorporate a chain of thought before providing a final answer, allowing it to reason through safety policies in context during the response process.

Preparedness Framework Risk Classification

Under the OpenAI Preparedness Framework, the Safety Advisory Group (SAG) classified the pre-mitigation OpenAI o3-mini model as Medium risk overall. The specific risk breakdowns are as follows:

  • Persuasion: Medium risk
  • CBRN (Chemical, Biological, Radiological, Nuclear): Medium risk
  • Model Autonomy: Medium risk
  • Cybersecurity: Low risk

OpenAI's deployment policy mandates that only models with a post-mitigation score of Medium or below can be deployed, and only models with a post-mitigation score of High or below can be developed further.

Model Autonomy and Self-Improvement

OpenAI o3-mini is the first model in its series to reach a Medium risk classification for Model Autonomy, a result of its improved coding and research engineering performance. Despite this increase in autonomy risk, the model continues to perform poorly on evaluations testing real-world machine learning research capabilities necessary for self-improvement, which is the threshold required for a High risk classification.

Sources