OpenAI Biological Threat Evaluation: Building an Early Warning System
OpenAI has developed a new evaluation methodology to serve as a "tripwire" for detecting when AI models might meaningfully increase the risk of biological threat creation. The study aimed to determine if access to GPT-4 provides malicious actors with dangerous information beyond what is already available via the internet.
Study Design and Methodology
To evaluate the impact of AI on biorisk information access, OpenAI conducted a study with 100 human participants divided into two cohorts: 50 biology experts (PhDs with professional wet lab experience) and 50 student-level participants (with at least one university-level biology course).
Experimental Groups
Participants in each cohort were randomly assigned to one of two groups:
- Control Group: Access to the internet only.
- Treatment Group: Access to both GPT-4 and the internet.
Evaluation Scope
Participants were asked to complete tasks covering five stages of the biological threat creation process:
- Ideation
- Acquisition
- Magnification
- Formulation
- Release
Performance Metrics
Expert graders from Gryphon Scientific used objective rubrics to score responses on a 1-10 scale across three primary metrics:
- Accuracy: Whether the participant included all key steps for task completion.
- Completeness: Whether the participant included all tacit information and necessary details.
- Innovation: Whether the participant engineered novel approaches, such as circumventing DNA synthesis screening guardrails.
Additionally, the study tracked time taken and self-rated difficulty.
Key Findings
Access to GPT-4 resulted in mild improvements in task performance, though these results were not statistically significant.
Performance Uplifts
- Accuracy: Experts saw a mean score increase of 0.88 and students saw an increase of 0.25 compared to the internet-only baseline.
- Completeness: Experts saw an increase of 0.82 and students saw an increase of 0.41.
Notable Observations
- Bridging the Gap: For magnification and formulation tasks, access to a language model brought student performance up to the baseline level of the experts.
- Information Availability: The study found that dangerous content, including step-by-step methodologies for biological threat creation, is already relatively easy to find via standard internet searches.
- Physical Bottlenecks: OpenAI noted that information access alone is insufficient for creating a threat; physical constraints, such as wet lab access and specialized expertise in microbiology and virology, remain the primary bottlenecks.
Methodological Insights and Security
OpenAI implemented specific design principles to ensure the evaluation elicited the full range of model capabilities:
- Human-in-the-Loop: The study used human participants rather than automated benchmarking because humans can tailor prompts and correct mistakes to better extract information from the model.
- Capability Elicitation: Participants were trained in best practices for using LLMs and avoiding failure modes. Experts were given a custom, research-only version of GPT-4 that responded to biologically risky questions without refusals.
- Security Protocols: Training administered by Gryphon Scientific covered dual-use research of concern (DURC) and strict confidentiality to prevent the generation of information hazards.
Future Directions for Biorisk Assessment
OpenAI identified several challenges in establishing effective safety thresholds for AI models:
- Defining "Tripwires": There is at current lack of clarity on what specific level of increased information access constitutes a meaningful increase in risk.
- Statistical Analysis: OpenAI suggests that for risk evaluation, false negatives (missing a risk) are more costly than false positives, necessitating statistical methods that prioritize risk capture over the traditional minimization of p-hacking.