Learning more about Claude's mathematical capabilities

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Claude improved a longstanding lower bound for zeros of the Riemann zeta function that satisfy the Riemann hypothesis from 41.6% to 67.2%, synthesizing prior work by Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh with a 2000 paper by Bombieri.
- Anthropic staff member Jarred Sumner, a self-described non-mathematician, prompted Claude to "take a real stab" at the Riemann hypothesis itself; Claude failed at that goal but produced the new bound as a byproduct.
- Claude initially generated and tried 650 unsuccessful ideas, then spent about a day and a half coordinating roughly 60 subagents that ran 2,400 shell commands and wrote hundreds of Python scripts, consuming 31 million output tokens across two Claude Code sessions.
- Sumner's input during the extended session was mostly limited to encouragement messages like "keep going" or "believe in yourself," which seemed to help Claude overcome initial skepticism that it could make meaningful progress.
- Anthropic mathematicians Levent Alpöge and Ralph Furman examined and validated Claude's proof; Claude also produced a Lean formalization that passes a standard validation comparator, re-proved the result from scratch, and downloaded 54 arXiv papers to confirm the finding was novel.
- External experts Brian Conrey and Dan Goldston also reviewed the paper on short notice, though Anthropic notes the techniques used are not expected to lead to a proof of the Riemann hypothesis itself.
- Claude was initially skeptical of its own result, possibly because its training emphasized the difficulty of open math problems and the limitations of AI models, but arrived at the finding after sustained encouragement.
Why it matters: This is a concrete mathematical advance, not a benchmark score: the known minimum proportion of Riemann zeta zeros on the critical line jumps by over 25 percentage points using a result a non-mathematician elicited through encouragement alone. The validation pipeline — subagent peer review, Lean formal verification, novelty check against 54 arXiv papers, and external expert review — offers a template for how AI-generated proofs can be checked before being trusted by the math community.
Ask SkimNews



