OpenAI GPT-5 Biological Research Capabilities
GPT-5 Increases Molecular Cloning Efficiency 79-Fold
OpenAI, in collaboration with biosecurity start-up Red Queen Bio, has demonstrated that GPT-5 can autonomously reason about, propose, and iterate on wet lab protocols to significantly accelerate biological research. In a controlled experimental setup, GPT-5 optimized a molecular cloning protocol, resulting in a 79-fold increase in efficiency compared to the baseline HiFi assembly protocol.
This improvement was achieved through a combination of two distinct optimizations:
- Enzymatic Assembly (RAPF): A 2.6-fold improvement.
- Transformation Protocol (T7): A 36-fold improvement.
Combined, these additive improvements resulted in the total 79-fold increase in the recovery of sequence-verified clones for a fixed amount of input DNA.
Discovery of the RAPF-HiFi Enzymatic Mechanism
GPT-5 proposed a novel enzymatic procedure termed "RecA-Assisted Pair-and-Finish HiFi Assembly" (RAPF-HiFi). The model identified a synergistic combination of two proteins to improve the homology search and annealing process of DNA assembly:
- gp32 (phage T4 gene 32 single-stranded DNA-binding protein): Acts to smooth and detangle loose DNA ends by removing secondary structures.
- RecA (recombinase from E. coli): Guides each DNA strand to its correct match, promoting homology search and annealing.
Technical Mechanism of Action
OpenAI hypothesizes the RAPF-HiFi mechanism operates as follows:
- T5 exonuclease creates 3’ overhangs.
- gp32 coats non-annealed single-stranded DNA (ssDNA) tails to suppress secondary structure.
- RecA invades from the 3’ ends, displacing gp32 and driving the annealing of matching sequences.
- Thermal Reset: A return to 50°C displaces both RecA and gp32, allowing the standard polymerase and ligase to complete the reaction.
The model's reasoning was highlighted by its choice of gp32 over the more common E. coli SSB protein; while SSB is a natural partner to RecA, it binds too stably to allow for the efficient RecA displacement required for this protocol.
Optimization of the Transformation Protocol (T7)
While the enzymatic assembly was optimized iteratively over five rounds, the transformation procedure was optimized in a "one-shot" round. GPT-5 identified a modification called "Transformation 7" (T7), which involved pelleting the cells, removing half of the supplied volume, and resuspending the cells at 4°C before adding DNA.
Despite the common laboratory belief that high-efficiency chemically competent cells are fragile and should not be handled this way, the cells tolerated the concentration process. This modification increased molecular collisions and reduced inhibitory buffer, leading to a transformation efficiency increase of over 30-fold.
AI-Driven Evolutionary Framework for Wet Lab Iteration
To test the model's capabilities, OpenAI used an evolutionary framework where GPT-5 learned "online" from experimental feedback:
- Iterative Loop: GPT-5 proposed batches of 8-10 reactions per round.
- Human-in-the-Loop Execution: Human scientists executed the protocols and provided colony counts relative to the baseline.
- Standardized Prompting: To ensure the AI's contributions were independent of human guidance, the prompting was fixed with no human intervention beyond clarifying questions.
OpenAI noted that the current system's fixed prompting limited the balance between exploration and exploitation, suggesting that future advances in planning and task-horizon reasoning could yield even larger gains.
Autonomous Robotic Execution
In collaboration with Robot on Rails, OpenAI developed a robotic system to execute these natural language protocols. The system integrates a human-to-robot LLM, a real-time vision system for labware localization, and a robotic path planner.
In comparative tests, the robot successfully executed the improved R8 protocol, achieving a 2.13-fold improvement over the baseline, which represents 89% of the performance achieved by human researchers (2.39-fold). While absolute colony counts were approximately ten-fold lower than manual execution due to liquid handling and temperature calibration differences, the robot's ability to rank protocols accurately demonstrates the potential for fully autonomous biological experiment optimization.
Biosecurity and Safety Framework
Due to the potential risks associated with biological reasoning capabilities, this research was conducted in a tightly controlled environment using a benign experimental system. OpenAI is utilizing these results to inform biosecurity risk assessments and develop safeguards as part of its Preparedness Framework.