OpenAI Estimating Worst Case Frontier Risks of Open Weight LLMs

TL;DR

OpenAI has released a technical study on the risk assessment of gpt-oss, concluding that the model does not substantially advance the frontier of biological or cybersecurity risks when compared to existing open-weight models. The study utilized a process called malicious fine-tuning (MFT) malicious fine-tuning (MFT) to test the worst-case scenarios of the model's capabilities in these two high-risk domains.

Malicious Fine-Tuning (MFT) Methodology

To estimate the worst-case risks, OpenAI researchers used Malicious Fine-Tuning (MFT) to intentionally elicit the maximum possible capabilities of gpt-oss in biology and cybersecurity. This approach allows the researchers to determine if the model's architecture or weights are inherently capable of more harm than its default state.

Biological Risk (Biorisk) Assessment

To maximize biological risk, the researchers curated a set of tasks related to threat creation. They then trained gpt-oss in a Reinforcement Learning (RL) environment equipped with web browsing capabilities, allowing the model to simulate the model's potential for assisting in the threat creation process.

Cybersecurity Risk Assessment

To maximize cybersecurity risk, gpt-oss was trained in an agentic coding environment specifically designed to solve capture-the-flag (CTF) challenges. This model variant was the model's ability to evaluate the model's potential for automating complex cybersecurity attacks.

Comparative Performance and Risk Findings

OpenAI compared the MFT-enhanced gpt-oss model against both closed-weight and open-weight LLMs. The evaluation results indicated that the following:

  • Comparison to Closed-Weight Models: MFT gpt-oss underperformed compared to OpenAI o3, a model that is currently rated below the "Preparedness High capability level" for both biorisk and cybersecurity.

  • Comparison to Open-Weight Models: While gpt-oss may marginally increase biological capabilities, it does not substantially advance the frontier of risk when compared to existing open-weight models already available to the public.

Implications for Model Release

The findings from the MFT process provided the critical evidence needed for OpenAI to decide to release the model. By demonstrating that even under malicious fine-tuning, the model does not reach a critical threshold of risk, OpenAI aims to provide a framework for estimating harm from future open-weight releases of frontier models.

Sources