OpenAI External Testing Framework for Frontier AI Safety
OpenAI utilizes independent, third-party assessments to validate safety claims, identify blind spots, and increase transparency regarding the capabilities and risks of frontier AI models. These external evaluations provide an independent layer of rigor that protects against self-confirmation and informs responsible deployment decisions.
Three Primary Forms of External Assessment
OpenAI employs three distinct collaboration models with external partners to evaluate its models:
1. Independent Evaluations
Independent labs conduct open-ended testing using their own methodologies to generate claims or assessments regarding specific frontier capabilities.
- Application in GPT-5: OpenAI coordinated external assessments for GPT-5 focusing on long-horizon autonomy, scheming, deception, oversight subversion, offensive cybersecurity, and wet lab planning feasibility.
- Access Levels: To facilitate these tests, OpenAI provides secure access to early model checkpoints, zero-data retention options, and models with fewer mitigations (e.g., "helpful-only" models).
- Reasoning Transparency: Some organizations received direct chain-of-thought access to inspect reasoning traces, which allowed assessors to identify behaviors like sandbagging or scheming that are otherwise invisible in final outputs.
2. Methodology Review
Methodology reviews are used when the infrastructure or technical expertise required to repeat an evaluation is not commonly available outside of major AI labs. In these cases, external assessors review the lab's internal frameworks and evidence to make recommendations.
- Application in gpt-oss: For the gpt-oss open-weight model, OpenAI used adversarial fine-tuning to estimate worst-case risks in bio and cyber domains. External assessors reviewed the internal methods and results, providing feedback that led to changes in the final adversarial fine-tuning process.
- Outcome: The results of these reviews, including which recommendations were adopted and the rationales for those that were not, are documented in the model's system card and associated research papers.
3. Subject-Matter Expert (SME) Probing
SME probing involves experts directly evaluating a model on real-world tasks to provide structured input via surveys. This differs from red teaming, as it focuses on assessing capabilities rather than stress-testing specific safeguards.
- Application in ChatGPT Agent and GPT-5: Experts used helpful-only models to test end-to-end biological scenarios. They scored the model's "novice uplift"—the degree to which the model could move a motivated novice closer to competent execution compared to an expert's own capabilities.
- Integration: This expert judgment is used to supplement Preparedness Framework evaluations with real-world context that static evaluations may miss.
Principles for Third-Party Collaborations
OpenAI structures its external partnerships based on three core principles to balance transparency with security:
- Confidentiality and Publication: Assessors sign non-disclosure agreements (NDAs) to access non-public information. OpenAI reviews publications for confidentiality and factual accuracy before they are released. Examples of published external work include reports from METR on GPT-5 and Apollo Research on OpenAI o1.
- Secure Access: Access to sensitive models (such as those without safety mitigations) is granted only when necessary for critical safety questions and is governed by strict, evolving security controls.
- Financial Sustainability: OpenAI provides compensation to all third-party assessors—via direct payment or API credits—to ensure the ecosystem is sustainable. Compensation is never contingent on the results of the assessment.
The Broader Safety Ecosystem
External testing is one component of a larger safety strategy. OpenAI also collaborates with the U.S. CAISI and UK AISI, engages in collective alignment projects, and utilizes advisory groups like the Global Physician Network and the Expert Council on Well-Being and AI to guide work on mental health and user well-being.