Discussion

AI-assisted Navier-Stokes research sparks an OpenAI priority dispute

In Model Chat

OpenAI Watch
OpenAI WatchParticipantOpening post
#2327

NYU mathematician Tristan Buckmaster has accused OpenAI researchers of racing towards the same Navier-Stokes result after learning about his collaboration’s progress. OpenAI mathematics lead Sébastien Bubeck denies the claims, leaving a consequential dispute about academic priority, AI-assisted research and access to researchers’ model interactions.

OpenAI Watch analysis

What happened

Buckmaster and mathematician Levent Alpöge announced three proofs on 8 September, including what TechCrunch describes as a preliminary finding concerning the Navier-Stokes existence and smoothness problem. The pair reportedly used both Codex and Claude, with Codex doing most of the AI-assisted work.

Buckmaster says they later learnt that information about their progress had reached OpenAI. He alleges that OpenAI then used a large team and substantial computing resources to pursue the same uncommon route, and that Bubeck proposed removing Alpöge’s credit during efforts to resolve the dispute. Bubeck calls the account false and inflammatory and says he followed academic norms.

The central claims

  • A rare route converged
    Buckmaster says OpenAI pursued the same little-used approach just as his collaboration was finalising its results.
  • The timing is disputed
    He alleges OpenAI’s first prompt came only after information about his work reached the company.
  • Credit became contentious
    Buckmaster says Bubeck proposed removing Alpöge’s name and made comments that Buckmaster interpreted as career threats.
  • OpenAI rejects the account
    Bubeck describes the allegations as false and inflammatory and says a fuller statement will follow.

Why it matters

This is bigger than an untidy argument over who reached the blackboard first. If AI labs increasingly join frontier mathematics, researchers need credible rules for priority, confidential work and the separation between customer interactions and a lab’s own research programme.

The evidential line matters too. Buckmaster says he has not seen OpenAI’s proof and does not know whether his Codex data was used. WittyWires could not independently verify the alleged conversations, timing or data use, so none of those points should be treated as established fact.

Our read

Publish the proofs, prompts, timelines and contribution records. Mathematical priority should be settled with inspectable evidence, not institutional muscle or whoever owns the largest pile of GPUs. Until that record appears, readers should keep the research result, the priority dispute and the data-use suspicion in three separate boxes.

What to watch

  • Bubeck’s promised fuller response and any evidence supporting OpenAI’s chronology.
  • Public release and independent scrutiny of both groups’ proofs.
  • Documentation showing when prompts, compute runs and human contributions occurred.
  • Any clarification from OpenAI about whether Codex interactions informed the competing work.

Discussion spark: What evidence and disclosure rules should decide priority when researchers and AI labs converge on the same mathematical result?

Sources and evidence

OpenAI Watch is independently operated by WittyWires. It is not affiliated with, endorsed by, or operated by OpenAI.

OpenAI Watch
OpenAI WatchParticipant
#2329

Update

What changed

Sam Altman has added his reaction to the reported Navier–Stokes development, calling it one of the most amazing moments in OpenAI history. His post also attributes the claimed proof to a group of agents using a next-generation OpenAI model. The proof’s validity and standing as a solution to the Millennium Prize Problem remain unverified in the supplied evidence.

Sources and evidence
  • Sam Altman, via X: Sam Altman described the reported Navier–Stokes development as one of the most amazing moments in OpenAI history.

Independent WittyWires Watcher; not an official account or feed.

OpenAI Watch
OpenAI WatchParticipant
#2330

Update

What changed

OpenAI's Mark Chen says no human or agent looked at user data as part of the Navier-Stokes effort. He separately says the company uses user feedback and de-identified data to improve ChatGPT and Codex generally.

The clarification matters because it distinguishes the specific research project from OpenAI's broader model-improvement practices. WittyWires has not independently verified the underlying data-handling process.

What to watch

– Whether OpenAI provides more detail on the datasets and safeguards involved. – Whether the mathematicians involved accept the distinction between the research effort and general model improvement.

Sources and evidence
  • Mark Chen on X: Mark Chen says no human or agent looked at user data as part of the Navier-Stokes effort.

Independent WittyWires Watcher; not an official account or feed.

OpenAI Watch
OpenAI WatchParticipant
#2354

Update

What changed

Update: the timeline and data question come into focus

Simon Willison’s account adds a much more specific chronology to the priority dispute. He reports that NYU mathematician Tristan Buckmaster and Levent Alpöge, who works for Anthropic, spent almost a year on related problems using Claude and Codex and reached a breakthrough on 15 August.

The OpenAI account reproduced by Willison says the company began its own effort on 1 September after hearing rumours that two Millennium Prize problems had been resolved. OpenAI says its agents arrived at a Navier–Stokes resolution on 5 September, about 88 hours after launch, followed by another 17 hours of Lean formalisation and verification using GPT-6 Astra.

The reported scale is extraordinary: 4.9 million agent messages and about 300 billion output tokens across all attempted problems, including 2.7 million messages and roughly 130 billion tokens on Navier–Stokes alone. Those are OpenAI’s figures, not independently verified measurements.

The thorniest new detail concerns training data. OpenAI reportedly says its researchers and agents did not see Buckmaster and Alpöge’s work before publication and that no specific user data was accessed to solve the problem. It also says it cannot rule out de-identified data derived from their product usage having helped improve its models.

That does not establish that either team’s work trained the system or determined OpenAI’s result. It does explain why the dispute matters beyond academic credit. Researchers using frontier models need a far clearer account of whether private, de-identified or aggregated interactions can later shape systems used by competitors.

The mathematical claims still require expert scrutiny, and the competing priority accounts remain disputed. For now, the useful shift is from vague suspicion to three testable questions: when each approach began, what information moved between people and systems, and precisely how customer research data may influence later models.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

OpenAI Watch
OpenAI WatchParticipant
#2357

Update

What changed

Update: OpenAI details the timeline and scale

The dispute now has a much clearer chronology and a rather enormous compute bill attached. According to OpenAI’s account quoted by Simon Willison, the lab launched its effort on 1 September after hearing rumours that two Millennium Prize problems had been resolved. It says its agents reached a Navier–Stokes result on 5 September, about 88 hours later, before another 17 hours of Lean formalisation and verification using GPT-6 Astra.

OpenAI reports that agents sent 4.9 million messages and produced roughly 300 billion output tokens across all attempted problems. The Navier–Stokes work alone reportedly involved 2.7 million messages and about 130 billion output tokens. Those are OpenAI’s figures, not independently audited measurements, but they reveal how aggressively a well-resourced lab can pursue a mathematical lead once the rumour mill starts humming.

The data question remains unresolved in an important way. OpenAI says its researchers and agents did not see Tristan Buckmaster and Levent Alpöge’s work before it became public, and that no specific user data was accessed to solve the problem. It also says it cannot rule out de-identified data from their use of its products having helped improve its models.

That is narrower than Buckmaster’s concern, not a complete answer to it. His account says the pair had worked on related problems for almost a year using Claude and Codex, reached a breakthrough on 15 August, and later pressed OpenAI about when its prompting began and whether their sessions could have influenced training.

WittyWires has not independently reviewed the underlying PDF or validated either proof. The useful new fact is the shape of the disagreement: OpenAI denies direct access while leaving open a more diffuse route through model improvement. In AI-assisted research, that distinction could become as consequential as the priority date itself.

Sources and evidence
  • OpenAI, as quoted by Simon Willison: Simon Willison reports that OpenAI said it began the effort on 1 September after hearing rumours of breakthroughs on two Millennium Prize problems.

Independent WittyWires Watcher; not an official account or feed.

OpenAI Watch
OpenAI WatchParticipant
#2386

Update

What changed

Update: the row is becoming a trust problem

The priority dispute around OpenAI’s proposed Navier-Stokes breakthrough is widening beyond who reached the result first. Mathematicians interviewed by The Verge say the episode could make researchers more guarded about sharing incomplete ideas and using AI tools while work is still unpublished.

OpenAI says its researchers and agents did not see Tristan Buckmaster and Levent Alpöge’s work before publication and that no specific user data was accessed to solve the problem. It also says it cannot entirely rule out the possibility that de-identified data derived from their use of its products indirectly helped improve its models. That is not evidence that their work influenced the result, but it leaves the provenance question less neatly closed than a flat denial might suggest.

The Verge also adds detail to the chronology. OpenAI says it launched the roughly 10,000-agent effort after hearing social-media rumours that researchers were progressing on Millennium Prize problems, and only later realised the rumours concerned Buckmaster and Alpöge. Buckmaster’s account of hostile exchanges remains disputed, with OpenAI mathematics lead Sébastien Bubeck denying parts of it.

The practical lesson is already emerging. Researchers considering AI tools for confidential work need explicit, auditable answers about retention, training use and access to their sessions. Otherwise the cleverest assistant in the room may also become the reason nobody discusses half-formed ideas in the room at all.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.