OpenAI AI Chemist: Improving Chan-Lam Coupling with GPT-5.4

OpenAI AI Chemist: Improving Chan-Lam Coupling with GPT-5.4

TL;DR

OpenAI and Molecule.one have developed a near-autonomous AI chemistry system that successfully improved a challenging reaction in medicinal chemistry. By pairing GPT-5.4 with the Maria AI agent and a high-throughput laboratory, the system identified that using mild oxidants like TEMPO significantly increases the yield of carbon-nitrogen bonds when coupling primary sulfonamides with boronic acids.

Improving the Chan-Lam Coupling Reaction

The system achieved a significant increase in the yield of primary sulfonamide Chan-Lam coupling, a reaction critical for creating carbon-nitrogen bonds in medicines used for oncology and infectious diseases. Historically, this specific reaction class has suffered from low yields, creating a bottleneck in drug discovery because scientists can only test molecules they are able to synthesize.

Key performance improvements under optimized conditions include:

  • Mean Yield Increase: The average yield rose from 16.6% to 25.2%.
  • Success Rate: The share of reactions with yields above 30% increased from 15.6% to 37.5%.
  • Substrate Improvement: Yields improved for 88% of the boronic acids and 83% of the sulfonamides tested.

These results were validated by human chemists at bench scale, where 11 of 14 representative substrate pairs showed higher yields, with most seeing a more than twofold increase.

Technical Architecture: GPT-5.4 and Maria AI

The research utilized a "near-autonomous" workflow combining a frontier model, an agentic AI, and physical laboratory infrastructure. The process functioned as follows:

  1. Proposal Generation: GPT-5.4 was used within a harness to generate and rank thousands of research proposals based on an open-ended goal to improve reaction classes.
  2. Human Selection: Human chemists reviewed the top-ranked proposals and selected four for testing.
  3. Execution: Maria AI translated high-level plans into detailed laboratory instructions and executed thousands of high-throughput experiments in the Maria Lab.
  4. Analysis and Iteration: Maria AI analyzed raw data and returned structured results to GPT-5.4, which then proposed follow-up experiments to refine hypotheses.

In the specific case of proposal OAI-M1-03, the system identified primary sulfonamides as a high-value target and suggested the use of TEMPO as a mild oxidant. A subsequent iteration found that a cheaper analog, 4-hydroxy-TEMPO, could be used with little loss in performance.

Scale of Experimentation

The Maria Lab executed 10,080 reactions for the OAI-M1-03 proposal. This volume of data—equivalent to a decade of work for a chemist running three reactions per day—allowed the system to identify TEMPO among ten tested oxidants and ensure the results were consistent across diverse molecular combinations, reducing the risk of artifacts common in small-scale screenings.

Limitations and Human Oversight

OpenAI emphasizes that the system is not fully autonomous. Human chemists remained essential for:

  • Steering and Judgment: Designing prompts and selecting which proposals to test.
  • Experimental Correction: Making critical adjustments, such as removing dimethyl sulfoxide (DMSO) as a solvent to prevent reactions with strong oxidants.
  • Physical Support: Preparing reagents and consumables and performing manual bench-scale validation.

Furthermore, the results do not yet establish that this method generalizes to other coupling reactions or manufacturing conditions, and the reaction mechanism requires further characterization.

Safety and Preparedness

To mitigate risks associated with chemical AI, the project was strictly scoped to a legitimate medicinal chemistry problem. The workflow included multiple layers of control:

  • Domain Restriction: The experiments focused on known coupling reactions for drug-like molecules and did not involve toxins or chemical weapons.
  • Model Safeguards: The model had undergone evaluations with the UK AI Security Institute and was designed to refuse harmful requests.
  • Human Gatekeeping: Human chemists retained full control over the physical infrastructure and the selection of experimental plans.

Sources