Mathematician demands proof OpenAI didn't train on his work — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Andreas Thom accused OpenAI of 'dishonesty' after the company announced a breakthrough on non-sofic groups that built on his and Gábor Kun's work, yet failed to conclusively rule out whether his ChatGPT conversations entered its training data.
- OpenAI researchers Sébastien Bubeck and Mark Sellke told Thom only that his conversations couldn't be 'accessed directly' — not whether they fed into training pools — a distinction Thom called 'materially misleading.'
- Tristan Buckmaster, an NYU professor, had publicly questioned whether OpenAI's models benefited from his Codex use; OpenAI quietly amended its non-sofic groups writeup after mathematicians criticized it for not crediting Thom and Kun.
- In announcing its Navier-Stokes/Millennium Prize solution, OpenAI said it 'did not see any of their work through any means' but conceded it 'cannot rule out that de-identified data derived from their usage of our products helped improve our models.'
- Thom argued that 'de-identification may remove a name; it does not remove the intellectual content of a mathematical idea,' and said only OpenAI has the data needed to prove a negative — placing the burden of disclosure on the company.
- Multiple mathematicians told The Verge they worry this behavior will push the field into a more secretive state if researchers fear even rumored breakthroughs could ignite a race with a 'well-resourced tech giant eager for glory.'
Why it matters: Multiple mathematicians told The Verge this episode could push their field toward secrecy, with researchers fearing rumored breakthroughs might trigger a race with OpenAI. The deeper problem: OpenAI's standard 'de-identified data' caveat lets it deny specific theft while refusing to prove a negative, leaving mathematicians like Thom unable to verify whether their work powered its headline results.
Ask SkimNews




