Mathematician calls OpenAI 'dishonest' over training data — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Andreas Thom posted on Mastodon accusing OpenAI of 'dishonesty' after the company announced a result on non-sofic groups that it acknowledged built heavily on Thom's and Gábor Kun's previous work.
- OpenAI initially failed to credit Thom and Kun in its writeup, then quietly amended the document after the omission drew criticism in mathematical circles.
- Thom emailed OpenAI researchers Sébastien Bubeck and Mark Sellke asking whether his ChatGPT conversations were 'part of the training data or accessible to the reasoning process' but said the response only addressed direct access, not absorption into training pools.
- OpenAI told mathematician Tristan Buckmaster it 'cannot rule out' that 'de-identified data derived from their usage of our products helped improve our models' — a distinction Thom called 'obfuscatory,' writing that 'de-identification may remove a name; it does not remove the intellectual content of a mathematical idea.'
- Buckmaster, an NYU professor who had publicly questioned OpenAI earlier over its Navier-Stokes Millennium Prize solution, was working on those problems in a personal capacity with Anthropic researcher Levent Alpöge.
- Mathematicians told The Verge they worry the episode could push the field toward greater secrecy, since researchers could now lose work to a well-resourced AI company even before publishing.
Why it matters: OpenAI's 'we can't rule it out' stance places the burden of proof on mathematicians rather than the company that controls the training pipeline, creating an asymmetry where researchers cannot verify whether their unpublished interactions were absorbed. If the math community retreats into secrecy to protect work-in-progress, that could slow the open collaboration that helped produce the very breakthroughs OpenAI is now racing to claim credit for.
Ask SkimNews




