OpenAI AI Governance Commitments

OpenAI has committed to a a set of voluntary commitments to move AI governance forward, focusing on safety, transparency, and the responsible development of frontier AI systems. These commitments center on red-teaming, cross-industry information sharing, the identification of AI-generated content, and public reporting on model limitations.

Comprehensive Red-Teaming and Safety Evaluations

OpenAI commits to internal and external red-teaming for all major public releases of new models within scope. This process involves drawing on independent domain experts to evaluate systems for misuse, societal risks, and national security concerns.

Specific areas of focus for red-teaming include:

  • Biological, Chemical, and Radiological Risks: Evaluating how systems might lower barriers to entry for the development, design, acquisition, or use of weapons.
  • Cyber Capabilities: Assessing how systems could aid in vulnerability discovery, exploitation, or operational use, while noting that these capabilities can also serve defensive purposes.
  • System Interaction and Tool Use: Analyzing the capacity for models to control physical systems.
  • Self-Replication: Evaluating the capacity for models to make copies of themselves.
  • Societal Risks: Addressing bias and discrimination.

To support these efforts, OpenAI is advancing research into AI safety, specifically the interpretability of decision-making processes and the robustness of systems against misuse. These safety procedures will be disclosed in transparency reports.

Information Sharing and Standard Setting

OpenAI will work toward information sharing among companies and governments regarding trust and safety risks, dangerous or emergent capabilities, and attempts to circumvent safeguards. This includes establishing or joining forums to develop and adopt shared standards and best practices, such as the NIST AI Risk Management Framework.

These mechanisms will facilitate the sharing of information on frontier capabilities and emerging threats, and the engagement of governments, civil society, and academia to develop technical working groups on priority areas of concern.

Identification of AI-Generated Content

To ensure users can distinguish between human and AI-generated content, OpenAI commits to developing robust mechanisms for audio and visual content created by its publicly available systems. This includes the use of provenance and watermarking systems.

Key details of this implementation include:

  • Scope: Applies to audio or visual content introduced after the watermarking system is developed.
  • Tools: The development of APIs or tools to determine if content was created by their system.
  • Exclusions: Content that is readily distinguishable from reality (e.g., default AI assistant voices) is outside the scope.
  • Data: Watermarks or provenance data will include a model or service identifier, but will not include identifying user information.

OpenAI also pledges to work with industry peers and standards-setting bodies to develop a technical framework for distinguishing AI-generated content from user-generated content.

Transparency and Public Reporting

OpenAI commits to publishing reports for all new significant model public releases within scope. These reports will detail:

  • Safety Evaluations: Results of safety evaluations, including dangerous capabilities, provided they are responsible to disclose publicly.
  • Performance Limitations: Significant limitations in performance that have implications for domains of appropriate use.
  • Appropriate Use: Domains of appropriate and inappropriate use.
  • Societal Risks: Discussions on the model's effects on fairness and bias.
  • Adversarial Testing: Results of adversarial testing to evaluate the model's fitness for deployment.

Societal Risk Mitigation and Global Challenges

OpenAI prioritizes research on societal risks, including the avoidance of harmful bias and discrimination and the protection of privacy. This involves empowering trust and safety teams and protecting children to proactively manage AI risks.

Additionally, OpenAI agrees to support the research and development of frontier AI systems to address global challenges, such as:

  • Climate change mitigation and adaptation.
  • Early cancer detection and prevention.
  • Cybersecurity: Combating cyber threats.

OpenAI also commits to supporting initiatives for education and training to help students and workers prosper from AI and to help citizens understand the technology's impact.

Sources