The $1 million proof, the rival team, and the accusation OpenAI can't fully deny
Did OpenAI race a rival team to a Navier-Stokes proof after its Codex sessions with their drafts? Tristan Buckmaster says yes. OpenAI denies misconduct.

On September 8, OpenAI announced that an internal multi-agent system had proved finite-time blowup for the Navier-Stokes equations, resolving a Clay Millennium Problem (the mathematics is in our companion explainer). By evening the claim had a rival narrative. NYU mathematician Tristan Buckmaster and Levent Alpöge, a Harvard mathematician employed by Anthropic, released three blowup results of their own and alleged that OpenAI, after a year of their drafts in Codex, raced out a similar result and pressured Buckmaster to drop his co-author. OpenAI, Sébastien Bubeck, and Sam Altman deny misconduct. Both sides agree on the unsettling core: OpenAI started after a rumor about the pair's work and cannot fully rule out that its models learned from their data.
A red flag on a Sunday call
Buckmaster's statement describes a personal collaboration that used Claude and Codex to push the forced-blowup program of Diego Córdoba and Luis Martínez-Zoroa to completion, reaching smooth-forced blowup for Boussinesq and Euler by August 15 (Lean-verified August 22).
On September 3, as rumors spread that Anthropic had solved a major open problem, Buckmaster emailed a prominent OpenAI mathematician: "I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI." The same-day reply proposed "to avoid competing here."
The call came Sunday, September 6 at 12:45; Alpöge was absent. Buckmaster says he was told an internal model had proved finite-time blowup for the forced Navier-Stokes equations, "option c and d in Fefferman" — the route he and Alpöge had quietly chosen: "Almost nobody else I know of was working on it... When I heard 'forced,' it was a bright red flag." Told "very little human input" had been used, the call instead revealed, he says, a full team and "an insane amount of compute."
Two proposals followed: post Euler while OpenAI posted Navier-Stokes the next day, or write up the result alone, crediting OpenAI's internal model. He says Bubeck "twice asserted that he wanted Levent removed from authorship" over Alpöge's Anthropic employment. Declining both, he said he would go public: "Why would you ruin your career?" came the reply, then "If you don't want me to be nice, then I don't have to be nice." His caveat: he has not seen OpenAI's proof, does not know whether their data was used, and "is not accusing anyone of anything."
What OpenAI, Altman, and Bubeck say happened
OpenAI's blog post gives its own timeline: the effort began September 1 "after hearing a rumor which we later realized was related to Levent Alpöge... and Tristan Buckmaster." Around 10,000 concurrent agents worked the problems; the resolution arrived September 5, 88 hours in, Lean-verified via GPT-6 Astra. Completing it on September 6, OpenAI reached out "to offer a concurrent release of our result and to recognize their priority," learned the pair had forced Euler but not Navier-Stokes, offered visibility into all prompts, and will not claim the Millennium Prize.
Altman wrote on X that the team "acted with integrity and generosity throughout," offered the pair the chance to go first and "optionally for Tristan to be the lead author on a rewrite of the OpenAI proof," and met "unfounded accusations of plagiarism," adding it began from "rumors on the internet last week that Anthropic's models had solved a millennium problem." Bubeck denies ever asking "for Levent to be removed from authorship of his own work" and calls the "career" line an "extremely poor choice of words." He says Alpöge refused meetings; Alpöge contradicts that, telling The Decoder he would have liked to cooperate: "It's a shame!"
Two accounts, side by side
| Point in dispute | Buckmaster & Alpöge | OpenAI |
|---|---|---|
| When the effort began | First prompt sent "in the past few days" before Sept 6 | Sept 1, on rumors Anthropic models solved Millennium problems |
| Why the same route | The pair's chosen line; "almost nobody else" worked on it | Agents tried all variants (A–D) |
| Did user data help | "Did not look up user data"; no answer on training | "No specific user data was accessed"; cannot rule out de-identified data |
| Dropping Alpöge | Bubeck "twice asserted" he wanted Levent removed | "I never ever asked"; remark was about an Anthropic employee authoring OpenAI's rewrite |
| The "career" exchange | "Why would you ruin your career?" then "If you don't want me to be nice..." | Bubeck: "extremely poor choice of words," retracted on the spot |
Not in dispute: the rumor preceded the effort, OpenAI reached out on September 6, the Euler results differ (unforced vs forced), and the drafts lived in OpenAI's systems.
The question neither side can close
OpenAI's own language leaves the door ajar: it insists "no specific user data was accessed," then concedes: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." The first concerns access during the effort; the second, training — the answer Buckmaster says he never received.
OpenAI's help pages state that for individual services "such as ChatGPT and Codex, we may use your content to train our models," with an opt-out covering new conversations; whether Buckmaster opted out is not publicly known. The pair says it entered a similar approach into OpenAI's systems; Buckmaster called training on their inputs "absolute academic malpractice" on Mastodon; OpenAI employee Boaz Barak calls the idea "cope," noting the model "started off by proving a stronger claim than they did." TechCrunch priced the ~300 billion output tokens at about $22.5 million at current Astra rates; Fortune reports resources at least 1,000 times the ~$2,000 of earlier math challenges and quotes Satya Nadella and Alex Karp alleging frontier labs train on customer data.
A new deterrent for open science
The r/LocalLLaMA thread "OpenAI alleged of stealing mathematicians work" (roughly 1,265 points, 230-plus comments) anchored the reaction: "two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex... Feels like big labs believe everything you did with the help of their models is theirs." Its opening line ties it to self-hosted LLMs: "Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why."
Terence Tao drew the structural conclusion on Mastodon, via The Decoder: "even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential." Incentives, he warns, now point away from sharing promising directions "with the broader community," which would "do serious long-term damage to the future of the field." For researchers weighing whether to keep private drafts in a lab's product, this is the concrete exhibit — following OpenAI's research-acceleration push (the Alien Mind essay) and its math sprint (GPT-5.6's Erdős-problem proofs).
What would settle it
Nothing here is adjudicated. OpenAI could answer the training question directly — whether de-identified Codex content from these sessions entered any training run — and release the model records and prompts it already offered the pair. The mathematics has its own referee: the pair is finishing a Lean certificate for its withheld hypo-dissipative result, OpenAI shipped a Lean formalization with its proof, and independent verification of both is a matter of time, not trust. What neither side can produce is a record of intent.
Researchers who feed years of private work into a lab's tools face a question they cannot answer about their own results: whether the lab's next model learned from them. Labs that sprint at the sound of a rumor, however innocent their start, inherit an appearance problem no denial dissolves. Buckmaster's larger point, that a mathematician and an LLM did this work in a month, has been overtaken by the argument over who got there first — which may be the truer measure of where open science and the AI labs now stand.


