OpenAI Safety Practices

OpenAI has outlined ten core safety practices to ensure the development and deployment of industry-leading AI models remain safe and reliable. This framework integrates safety measures into the development process from the outset, moving from aligning current models to preparing for far more capable future systems.

Model Evaluation and Risk Mitigation

OpenAI employs empirical model red-teaming and testing before any release. This process is governed by the Preparedness Framework, which establishes risk thresholds. A new model will not be released if it crosses a "Medium" risk threshold until safety interventions are implemented to bring the post-mitigation score back to "Medium."

For the release of GPT-4o, more than 70 external experts were engaged in red-teaming efforts to identify weaknesses in earlier checkpoints, which then informed the evaluations for later checkpoints.

Alignment and Safety Research

Safety improvements are driven by a combination of increased model intelligence—which reduces factual errors and harmful outputs—and targeted investments in practical alignment, safety systems, and post-training research. Key areas of focus include:

  • Fine-tuning data: Improving the quality of human-generated data.
  • Model Specifications: Refining the instructions models are trained to follow.
  • Robustness: Conducting and publishing research to improve system resilience against attacks such as jailbreaks.

Monitoring and Abuse Prevention

OpenAI utilizes a broad spectrum of tools to monitor for safety risks and abuse across its API and ChatGPT. This includes the use of dedicated moderation models and utilizing GPT-4 to automate content policy development and content moderation decisions, which reduces the amount of abusive material exposed to human moderators.

To increase industry-wide safety, OpenAI has shared findings on state actor abuse of its technology in a joint disclosure with Microsoft.

Lifecycle and Specialized Safety Focus

OpenAI implements safety measures across the entire model lifecycle, from pre-training to deployment. This includes investing in pre-training data safety, system-level behavior steering, and robust monitoring infrastructure.

Specific high-priority safety areas include:

Child Safety

OpenAI has integrated default guardrails into ChatGPT and DALL·E to mitigate harms to children. In 2023, the company partnered with Thorn’s Safer to detect and report Child Sexual Abuse Material (CSAM) to the National Center for Missing and Exploited Children when users attempt to upload such content to image tools.

Election Integrity

To prevent abuse and ensure transparency, OpenAI has introduced tools for identifying DALL·E 3 images and incorporated C2PA metadata to help users identify the source of media. Additionally, ChatGPT directs users to official voting information sources in the U.S. and Europe and the company supports the "Protect Elections from Deceptive AI Act" in the U.S. Senate.

Impact Assessment and Security

OpenAI conducts impact assessments to measure risks associated with AI, including chemical, biological, radiological, and nuclear (CBRN) risks. The company also researches the extent to which language models impact various occupations and industries, and analyzes the implications of language models for influence operations.

Security measures to protect intellectual property and data include:

  • API-based deployment: Controlling access via API to enable policy enforcement.
  • Infrastructure security: Restricting access to training environments and algorithmic secrets on a need-to-know basis.
  • Cybersecurity programs: Implementing internal and external penetration testing, a bug bounty program, and a the Cybersecurity Grant Program for third-party researchers.
  • Future controls: Exploring confidential computing for GPUs and AI-driven cyber defense.

Governance and Government Partnership

OpenAI partners with governments globally to inform AI safety policies, sharing learnings and piloting third-party assurance.

Internally, safety decision-making is managed through an operational structure: the cross-functional Safety Advisory Group reviews capability reports and makes recommendations, company leadership makes final decisions, and the Board of Directors provides oversight.

Future Evolutions

OpenAI acknowledges that as it moves toward its next frontier model, it must evolve its practices. This includes increasing its security posture to be resilient against sophisticated state actor attacks and allocating additional time for safety testing before major launches.

Sources