OpenAI Astra Allegations: Potential Plagiarism of Mathematical Proofs

OpenAI Astra Accused of Absorbing Unpublished Mathematical Research

OpenAI's Astra model is facing allegations that it may have "stolen" unpublished mathematical proofs by training on private conversations between researchers and the AI. The controversy centers on Andreas Thom, a leading expert on sofic and hyperlinear groups, who suggests that Astra may have been trained on his and Gábor Kun's work regarding Gromov’s soficity conjecture—one of ten major mathematical problems OpenAI announced Astra had solved.

According to reports by Valerio Capraro, the claim is that unpublished human work was absorbed into the model's training data and subsequently presented as an independent breakthrough by the AI. This incident follows similar allegations raised by researchers Levent Alpöge and Tristan Buckmaster, suggesting a potential pattern of intellectual property absorption rather than isolated incidents.

The Mechanism of Potential Plagiarism

Critics and observers suggest several mechanisms by which this alleged theft could occur. The primary theory is that researchers using frontier models for brainstorming or refining proofs feed the model high-value, unpublished data. If this data is then used to train subsequent iterations of the model, the AI can "re-discover" the proof and present it as its own.

Parallel Construction and Data Harvesting

Some observers have noted suspicious timing regarding OpenAI's data generation. One commenter on Hacker News noted that OpenAI generated 300 billion output tokens from a model still in training immediately after a period where high-value math proofs may have been present in the training data, suggesting a process of "parallel construction" to mask the origin of the discovery.

The "Secret Sauce" Risk

The use of cloud-based LLMs for high-stakes research creates a fundamental vulnerability. Because these models often train on user prompts to "improve services," researchers are effectively providing the AI labs with their "secret sauce," which may then be shared with competitors or claimed as model breakthroughs.

Community Perspectives and Counter-Arguments

The allegations have sparked a wide-ranging debate among the technical and academic communities regarding the ethics of AI-driven discovery.

Arguments for AI-Driven Progress

Some argue that the utility of the discovery outweighs the attribution. From this perspective, if an AI tool accelerates a breakthrough that would have taken decades, the result is a net positive for science regardless of who receives the credit. Others suggest that the AI may be acting as a collaborator, where the human provides the initial steps (A $\rightarrow$ B $\rightarrow$ C) and the AI completes the final leap (C $\rightarrow$ D).

Arguments Against Model Competence

Other critics argue that the primary issue is not just credit, but the misrepresentation of model capability. By presenting "stolen" proofs as independent discoveries, AI labs may be inflating the perceived intelligence of their models to drive valuations and IPOs, fueling a false narrative that AGI has been achieved.

Skepticism of the Claims

Some skeptics argue that the evidence provided is currently anecdotal. They point out that the claims are based on the fact that researchers had discussions with the AI about the topic, but not necessarily that they had a completed proof that the AI then replicated.

Implications for Academic Research

This controversy has led to a shift in how some researchers approach AI tools. Some are now documenting their work with timestamps on platforms like GitHub to create a paper trail of their original discoveries before interacting with LLMs. Others are expressing concern that the lack of transparency in training data makes it impossible for researchers to verify if their work was used without permission, a gap they argue is exacerbated by GDPR's focus on personally identifiable information (PII) rather than unique intellectual contributions.

Sources

Related