OpenAI and Los Alamos National Laboratory Bioscience Research Partnership
OpenAI and Los Alamos National Laboratory (LANL) have entered into a research partnership to study the safe application of artificial intelligence in laboratory settings for bioscientific research. This collaboration aims to balance the acceleration of scientific discovery with the rigorous mitigation of biological risks.
Evaluation of Multimodal AI in Physical Laboratories
OpenAI and the LANL Bioscience Division are conducting an evaluation study to assess how frontier models, specifically GPT-4o, can assist humans in performing physical laboratory tasks using multimodal capabilities such as vision and voice. This study includes biological safety evaluations for GPT-4o and its unreleased real-time voice systems to determine how they can support bioscience research while maintaining safety standards.
Experimental Design and Task Proxy
The partnership is conducting the first experiment of its kind to test multimodal frontier models in a lab setting. The study assesses the ability of both experts (PhDs) and novices to perform and troubleshoot a safe protocol consisting of standard laboratory experimental tasks. These tasks serve as a proxy for more complex tasks that may pose dual-use concerns.
Specific tasks included in the evaluation include:
- Transformation: Introducing foreign genetic material into a host organism.
- Cell Culture: Maintaining and propagating cells in vitro.
- Cell Separation: Utilizing techniques such as centrifugation.
By measuring the uplift in task completion and accuracy, the researchers aim to quantify how frontier models can upskill professionals and novices in real-world biological tasks.
Technical Advancements in Safety Evaluations
This partnership extends previous AI safety work in two primary dimensions:
- Integration of Wet Lab Techniques: Previous evaluations focused on written tasks and responses for synthesizing compounds, which did not capture the physical skills required for biological benchwork. This study moves beyond theoretical knowledge (e.g., knowing the steps of mass spectrometry) to the actual performance of tasks with real samples.
- Multimodal Reasoning: While previous work focused on GPT-4's written outputs, GPT-4o's ability to process voice and visual inputs allows for real-time troubleshooting. For example, a user can show a wet lab setup to the model via camera and prompt it with questions to resolve scenarios visually rather than relying on written descriptions.
Institutional Framework and Safety Commitments
The collaboration is led by Los Alamos National Laboratory's new AI Risks Technical Assessment Group, which is specifically tasked with assessing and understanding AI-related risks. This partnership aligns with the U.S. Department of Energy's mandate under the White House Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence, which tasks national labs to evaluate the biological capabilities of frontier AI models.
This work builds upon OpenAI's existing biothreat risk research and follows the company's Preparedness Framework for tracking, forecasting, and protecting against model risks. It is also consistent with the commitments to Frontier AI Safety agreed upon at the 2024 AI Seoul Summit.
"AI is a powerful tool that has the potential for great benefits in the field of science, but, as with any new technology, comes with risks," said Nick Generous, deputy group leader for Information Systems and Modeling at Los Alamos.