Mathematician calls OpenAI 'dishonest' on training data — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Mathematician Andreas Thom publicly accused OpenAI of 'dishonesty' on Mastodon, raising concerns that his pre-announcement interactions with ChatGPT may have contributed to the company's breakthrough on non-sofic groups — a result OpenAI acknowledged built heavily on his and Gábor Kun's prior work.
- OpenAI initially failed to credit Thom and Kun in its writeup for the non-sofic groups result — one of 10 mathematical breakthroughs announced last month — and only quietly amended the post after criticism from the mathematics community.
- OpenAI researchers Sébastien Bubeck and Mark Sellke told Thom his ChatGPT conversations were not directly accessible, but Thom said their response addressed only direct access — not whether his exchanges entered OpenAI's broader training data pools.
- OpenAI denied using specific user data from NYU professor Tristan Buckmaster or Anthropic researcher Levent Alpöge in its Navier-Stokes solution post, but would not rule out that 'de-identified data derived from their usage of our products helped improve our models.'
- Thom dismissed that de-identification distinction as obfuscatory, writing that 'de-identification may remove a name; it does not remove the intellectual content of a mathematical idea' and calling Sellke's categorical answer 'plainly dishonest.'
- Multiple mathematicians told The Verge they worry the episode will push the field toward greater secrecy if researchers fear that even rumors of a breakthrough could trigger a race with a well-resourced OpenAI eager to publish first.
Why it matters: Multiple mathematicians told The Verge the episode is already leaving a 'sour taste' and warning it will push the field toward secrecy — cutting off the very informal idea-sharing with ChatGPT that fed into the breakthroughs OpenAI is now publicly celebrating.
Ask SkimNews




