Mathematicians accuse OpenAI of scraping private math proofs
- Mathematicians accuse OpenAI of scraping private Codex session data to claim a Millennium Prize math breakthrough.
- OpenAI deployed 100 AI agents burning over $1 million in compute after learning about human proof efforts.
- OpenAI denied targeting user data directly but could not rule out model training contamination.
OpenAI claimed a mathematical breakthrough. Independent researchers call it theft.
The conflict began after OpenAI promoted pure math proofs generated by its automated systems. On The a16z Show, researchers Metab Swani and Mark Selke detailed how their internal models solved long-standing conjectures in sphere packing and high-dimensional geometry. Rather than relying on brute force, the models backtracked when stuck, updated internal confidence, and pruned dead search paths to generate concise proofs.
That technical achievement quickly turned into an intellectual property war. On This Week in Startups, NYU mathematician Tristan Buckmaster and Anthropic researcher Levant Alpoge revealed they spent months using OpenAI's Codex tool to draft proofs for the 90-year-old Navier-Stokes existence and smoothness problem. Shortly before the pair published, OpenAI launched a focused internal push using the exact same unconventional mathematical direction.
The timeline raised immediate alarm. On Breaking Points, Buckmaster stated that OpenAI launched its project only after learning about his work in September. When Buckmaster asked whether OpenAI used his private Codex session data for model training, the company denied looking up user data directly but admitted it could not rule out training contamination. Buckmaster added that an OpenAI representative explicitly warned him against ruining his academic career if he went public.
Following OpenAI's deployment of GPT-6 Astra on September 13, details emerged regarding its compute usage. On The Intelligence from The Economist, science correspondent Sam Weigley reported that OpenAI deployed roughly 100 AI agents running on Astra. The agent swarm burned through more than $1 million in compute over four days to claim the solution. Though OpenAI waived the $1 million Millennium Prize money, the published paper lacked the detailed structural proofs demanded by traditional peer review.
The scandal exposes growing vulnerabilities for scientists using commercial AI tools. On This Week in Startups, Canvas Ventures co-founder Rebecca Lynn warned researchers against feeding proprietary research into commercial models. Agent Fund general partner Yohi Nakajima noted that autonomous web-scraping agents now index partial preprints aggressively. OpenAI executive Sebastian Bubeck rejected claims of deliberate IP theft, but the company's own report confirmed its effort began only after hearing industry rumors.
Math research now faces a stark reality. Scientists who use commercial AI assistants risk watching tech giants scrape their insights to claim credit.