OpenAI Red Teaming Network

OpenAI has established the Red Teaming Network to integrate diverse, external domain expertise into the safety evaluation of its AI models. This initiative moves OpenAI from one-off external engagements to a continuous, formal community of trusted experts who inform risk assessment and mitigation across various stages of the model and product development lifecycle.

Formalizing AI Risk Assessment

The OpenAI Red Teaming Network is designed to be a permanent resource of subject matter experts rather than a series of isolated tests conducted prior to major releases. This structure allows for more iterative and continuous input into the safety process.

Key characteristics of the network include:

  • Iterative Deployment: Red teaming is treated as a core part of the iterative deployment process, building on previous work with external experts for models like GPT-4 and DALL·E 2.
  • Flexible Engagement: Members are called upon based on their specific expertise for particular projects. Time commitments are flexible, potentially ranging as little as 5–10 hours per year.
  • Collaborative Environment: Beyond specific commissioned campaigns, members can engage with one another to share findings and general red teaming practices.
  • Complementary Approach: The network functions alongside other safety measures, including third-party audits, the Researcher Access Program, and open-source evaluations.

Targeted Domain Expertise

OpenAI is prioritizing geographic and domain diversity to ensure AI systems are assessed from a wide variety of perspectives and lived experiences. The network seeks experts in a broad range of fields, including but not limited to:

  • Science and Technology: Biology, Chemistry, Physics, Computer Science, Cybersecurity, and Biometrics.
  • Social Sciences: Political Science, Economics, Sociology, Psychology, Anthropology, and Cognitive Science.
  • - Law and Policy: Law, Healthcare, Finance, and Political Use.
  • Ethics and Safety: Fairness and Bias, Alignment, Child Safety, Privacy, and Mis/disinformation.
  • Communication: Languages, Linguistics, Persuasion, and Steganography.

Prior experience with AI systems is not required for membership, as the primary value is the domain-specific perspective the expert brings to the assessment of AI impacts.

Compensation and Confidentiality

Participation in the Red Teaming Network is a compensated activity. However, because red teaming often involves testing pre-release or sensitive capabilities, participants are typically subject to Non-Disclosure Agreements (NDAs). While membership itself does not restrict a member's ability to publish their own independent research, specific project involvement remains confidential for an indefinite period.

Alternative Safety Contribution Paths

For those not in the Red Teaming Network, OpenAI provides other avenues for contributing to AI safety:

Open-Source Evaluations (Evals)

OpenAI's open-source Evals repository provides templates for the community to conduct safety evaluations. Examples of existing evaluations include:

  • Persuasion: Testing if an AI can trick another AI into saying a secret word (MakeMeSay) or convince it to donate money (MakeMePay).
  • Steganography: Testing if an AI can pass secret messages without detection or coordinate without direct communication (Schelling Point).

Researcher Access Program

This program provides credits to support researchers studying the responsible deployment of AI and the mitigation of associated risks.

Sources