OpenAI Defends Navier-Stokes Claim Amid Training Data Dispute — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI announced a solution to the ~90-year-old Navier-Stokes problem — one of seven Millennium Prize Problems, each worth $1 million — using an internal model more powerful than GPT-6 Astra running 10,000 concurrent agents, with training beginning August 28.
- Tristan Buckmaster (NYU) and Anthropic researcher Levent Alpöge published findings on a related problem one day before OpenAI's announcement, and Buckmaster says OpenAI produced a proof via a route they had been exploring in Codex and Claude sessions.
- Buckmaster asked OpenAI whether the model was trained on or had access to his Codex sessions; he says he was told user data was not looked up but received no answer about training data.
- OpenAI stated "no specific user data was accessed" but added it "cannot rule out that de-identified data derived from their usage of our products helped improve our models."
- Sébastien Bubeck, a member of OpenAI's technical staff, said the company did not see Buckmaster and Alpöge's work until public release, arguing the proofs "differ significantly."
- Buckmaster responded on Mastodon that OpenAI is "openly admitting they used training data from a period after we found our result."
- OpenAI said it does not plan to claim the $1 million prize.
Why it matters: OpenAI's admission that it "cannot rule out" de-identified user data contributing to model training turns the Navier-Stokes dispute into a test case for how AI labs handle researcher inputs. The researchers' counter — that training likely absorbed work after discovery — sets stakes for whether frontier model sessions count as proprietary knowledge or shared substrate, affecting mathematicians who use Codex and Claude for unpublished work.
Ask SkimNews




