OpenAI First Proof Submissions
OpenAI has utilized an internal reasoning model to attempt all ten problems in the First Proof challenge, a research-level math competition designed to test AI's ability to produce checkable, end-to-end arguments in specialized domains. This milestone demonstrates the model's capacity for long-chain reasoning and the ability to tackle problems that were previously open for years.
Model Performance on First Proof Problems
OpenAI believes that at least five of the model's proof attempts—specifically problems 4, 5, 6, 9, and 10—have a high probability of being correct. While several other attempts remain under review, the team has retracted their initial belief that the attempt for problem 2 was correct following official commentary and community analysis.
Researcher James R. Lee notes that the model's capabilities improved tangibly during the training process:
‘We’re currently training a new model for which a primary focus is increasing the level of rigor in its thinking, with the goal that the model can think continuously for many hours and remain highly confident in its conclusions. When the First Proof problems were announced, it seemed like the perfect testbed... Already it was able to solve two of the problems (#9 and #10). As it trained, it became increasingly capable, eventually solving’—in our estimation—at least three more.‘
Methodology and Human Supervision
The proof attempts were generated with limited human supervision. The process involved several specific interventions:
- Strategy Guidance: Humans suggested retrying strategies that appeared fruitful in earlier attempts.
- Refinement: The model was asked to expand or clarify sections of a proof based on expert feedback to aid verification.
- Verification Support: A back-and-forth interaction between the internal model and ChatGPT was used for formatting, style, and verification.
- Selection: In cases where multiple attempts were made, humans selected the best version based on judgment.
OpenAI acknowledges that this was a "fast sprint" and that the process lacked the controlled environment of a formal evaluation, expressing a desire to collaborate with First Proof organizers for more rigorous future frameworks.
The Role of Frontier Research in AI Evaluation
OpenAI posits that novel frontier research is the most effective way to evaluate next-generation AI models. The company argues that standard benchmarks often fail to capture the most difficult aspects of research, such as:
- Sustaining long chains of reasoning.
- Selecting appropriate abstractions.
- Handling ambiguity in problem statements.
- Producing arguments that survive expert scrutiny.
Context of Reasoning Capabilities
These results build upon a series of advancements in AI reasoning within math and science:
- July 2025: A general-purpose reasoning model achieved gold medal-level performance on the International Mathematical Olympiad (35/42 points).
- November 2025: Case studies were shared regarding GPT-5's ability to accelerate science across physics, biology, and mathematics.
- Recent Developments: GPT-5.2 proposed a candidate expression for a gluon-amplitude formula in theoretical physics, which was subsequently formally proved by an internal model and verified by authors.
Sources
- OriginalOur First Proof submissions