OpenAI FrontierScience Benchmark Release

OpenAI has released FrontierScience, a new benchmark designed to evaluate expert-level scientific capabilities in AI models. This benchmark addresses the saturation of existing science benchmarks by providing original, difficult tasks that measure both structured scientific reasoning and open-ended research abilities.

FrontierScience Benchmark Structure

FrontierScience consists of over 700 textual questions across physics, chemistry, and biology, with a "gold set" of 160 questions used for primary evaluation. The benchmark is divided into two distinct tracks to measure different types of scientific intelligence:

FrontierScience-Olympiad

This track contains 100 questions designed by international olympiad medalists and national team coaches (representing 109 total medals). It assesses scientific reasoning through constrained, short-answer formats. These theoretical questions are designed to be at least as difficult as those found in international olympiad competitions.

FrontierScience-Research

This track consists of 60 original research subtasks designed by PhD-level scientists, including doctoral candidates, professors, and postdoctoral researchers. These tasks are self-contained, multi-step problems that mirror the difficulty a PhD scientist would encounter during active research. Performance on this track is graded using a 10-point rubric that evaluates both the final answer and intermediate reasoning steps.

Model Performance and Results

In initial evaluations, GPT-5.2 emerged as the top-performing model on both tracks, though results highlight a significant gap between structured reasoning and open-ended research:

  • FrontierScience-Olympiad: GPT-5.2 scored 77%, with Gemini 3 Pro following closely at 76%.
  • FrontierScience-Research: GPT-5.2 scored 25%.

Evaluations included other frontier models such as Claude Opus 4.5, Gemini 3 Pro, GPT-4o, OpenAI o4-mini, and OpenAI o3. The data indicates that longer reasoning effort (thinking time) leads to improved accuracy.

Despite these gains, OpenAI notes that frontier models still struggle with reasoning, logic, and calculation errors, factual inaccuracies, and a lack of understanding of niche scientific concepts.

Grading Methodology

To ensure scalability and accuracy, OpenAI employs different grading mechanisms for the two tracks:

  • Short Answer Verification: The Olympiad set uses numbers, expressions, or fuzzy string matches to verify correctness.
  • Rubric-Based Grading: For the Research set, a model-based grader (GPT-5) evaluates responses against a 10-point rubric. A solution is considered "correct" if it earns at least 7/10 points. This rubric-based approach allows for the analysis of intermediate reasoning steps rather than just the final output.

Implications for Scientific Research

While FrontierScience is a standardized measurement tool, OpenAI reports that models like GPT-5 are already accelerating real-world scientific workflows. According to the paper Early science acceleration experiments with GPT-5 (November 2025), these systems are being used for cross-disciplinary literature searches and complex mathematical proofs, often reducing work that previously took weeks to hours.

OpenAI positions FrontierScience as a "north star" for expert-level reasoning, though it acknowledges the benchmark's limitations. It does not currently measure the ability to generate genuinely novel hypotheses or the ability to interact with physical experimental systems and multi-modal data (such as video).

Sources