Mathematician calls OpenAI 'dishonest' over training data origins — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Andreas Thom posted on Mastodon accusing OpenAI of "dishonesty," saying the company's command of techniques in non-sofic groups — his area of expertise — was "neither the most obvious nor the most promising" routes at the time
- Thom emailed OpenAI researchers Sébastien Bubeck and Mark Sellke asking whether his prior ChatGPT interactions were part of training data or accessible to the reasoning process, but said the answer only addressed direct access, not training data inclusion
- Thom argued OpenAI's distinction between direct data access and de-identified user data is obfuscatory, writing: "De-identification may remove a name; it does not remove the intellectual content of a mathematical idea"
- OpenAI quietly amended its writeup on the non-sofic groups result to acknowledge recent contributions from Thom and Gábor Kun after criticism in mathematical circles, having originally announced ten breakthroughs last month
- OpenAI's blog post on its Millennium Prize Navier-Stokes solution stated it "did not see any of their work through any means" but added it "cannot rule out that de-identified data derived from their usage of our products helped improve our models"
- Tristan Buckmaster, a New York University mathematics professor working on Millennium Prize problems with Anthropic researcher Levent Alpöge, had publicly questioned weeks earlier whether his use of OpenAI's Codex had benefited the company
- Multiple mathematicians told The Verge they fear such behavior will push the field toward greater secrecy, with researchers reluctant to share early-stage ideas with AI assistants that might later race them to publication
Why it matters: The burden of proof now sits squarely with OpenAI: mathematicians argue only the company can verify whether user interactions entered training pipelines, and if it denies this pathway, it must disclose datasets and clarify its data-use terms. If OpenAI cannot conclusively rule out de-identified influence, the mathematicians warn their community will withdraw from sharing early-stage ideas — cutting off the informal exchanges that AI assistants currently benefit from.
Ask SkimNews




