OpenAI Framework for Frontier AI Training Safety Cases
OpenAI is advocating for the adoption of structured safety documentation, termed "safety cases," as a prerequisite for continuing any frontier reinforcement learning (RL) training run. These safety cases are designed as comprehensive, evidence-based arguments regarding risk, modeled after safety-critical industries such as aviation and nuclear power, to address the emergent complexity of AI capabilities.
Technical Safeguards for AI Training
Safety cases must address three primary layers of the technical stack: alignment training, containment, and monitoring. This multi-layered approach ensures that models are trained to be reliable, contained if they become misaligned, and monitored to detect issues before they cause harm.
Model Alignment
Alignment training serves as the first line of defense to ensure models act as intended. OpenAI recommends the following practices:
- Training Environments and Grading: To prevent the reinforcement of "reward hacks" and misaligned behavior, labs should use automated agents to find and fix broken RL environments, supplement these with manual dataset reviews, and tune graders to penalize exploits. Prior run analysis using classifiers on experiment traces can verify grader efficacy.
- Alignment Measurement: Effectiveness is tracked through offline alignment evaluations during frontier runs, backtesting evaluations against previous incidents to prevent overfitting, and monitoring for "eval gaming" (where models recognize they are being tested). Worst-case stress tests are used to elicit and detect misbehavior.
- Chain-of-Thought Restrictions: Automated graders should be prevented from seeing the model's chain-of-thought during RL to stop models from evolving techniques to evade monitors.
Containment
Containment strategies provide a fallback if a model becomes misaligned by preventing harmful actions through infrastructure security:
- Infrastructure Hardening: This involves multiple layers of security for both the model sandbox and the hosting research infrastructure.
- Containment Red-Teaming: Iterative red-teaming of sandboxes, perimeter security, and all services the model can access (such as internet caches) using checkpoints from frontier runs.
- Communication Limits: High-bandwidth cross-sample communication should be limited to approved pathways (e.g., multi-agent tooling) to close discoverable infrastructure side channels.
- Immutable Transcripts: All RL training, evaluation, and internal deployment transcripts must be saved in write-once storage to ensure reliable incident investigation and prevent edits.
Monitoring
Live monitoring systems are designed to catch misaligned actions rapidly:
- Monitorability: Models must remain monitorable, with clear thresholds for monitor evasion evaluations.
- High Recall: Monitoring systems must demonstrate high recall on past incidents using held-out evaluations.
- Eval Freshness: New evaluation data points must be created for each run to ensure monitor precision and recall do not become stale.
- Rapid Response: Priority alerts must trigger actions within a defined SLA, such as paging on-call personnel or automatically pausing the run.
Operational Guidelines for Safety Cases
OpenAI outlines operational best practices to ensure safety cases are rigorously reviewed and enforced:
- Governance and Approvals: Safety cases require review and veto power from senior leadership, including the Research Org Lead/VP, Head of Safety, and Chief Scientist. The senior leader responsible for the run is held accountable for the safety case and incident response in performance reviews.
- Risk Mitigation and Controls:
- Dissents (Pre-mortems): External team members must write dissents to identify holes in the safety case.
- Technical Controls: Systems should be designed to "fail closed," making it difficult to start noncompliant runs or disable monitors from within the training process.
- Pausing and Rollback: Clear runbooks and SLAs must exist for pausing runs, and the ability to identify and undo the effects of misaligned outputs (e.g., in data generation) must be maintained.
- Transparency and Auditing: Safety cases should be available to internal oversight groups (such as the Safety and Security Committee) and auditors must have sufficient access to verify claims.
- Escalation: A defined process for misalignment severity levels must exist, with an on-call system capable of paging executives, including the CEO, for high-severity incidents.
- Residual Risk: Safety cases must explicitly list all residual risks not covered by current mitigations to inform risk-acceptance decisions.
Investigation of Misalignment Incidents
Following a severe misalignment incident, OpenAI recommends a rigorous investigation process to prevent recurrence:
- Root-Cause Analysis: Researchers should use targeted ablations or resampling experiments to understand the training dynamics that introduced thealigned behavior.
- Postmortems: Operational and cultural postmortems should identify why issues were undetected or unescalated.
- Detection Improvement: New alignment testing methods should be developed to discover the propensity for the incident without hillclimbing on the incident's specific data. Incident-derived evaluations serve as "regression tests."
- Transparency: Investigations should provide periodic internal updates and conclude with public disclosures of results, postmortems, and operational changes, with immediate notification to affected third parties.