Research·4 min read·Anthropic

Claude Raised a Riemann Zeta Bound From 41.6% to 67.2%

An unreleased research version of Claude combined two existing results to lift the proven share of Riemann zeta zeros satisfying the hypothesis from 41.6% to 67.2%. It took 650 failed ideas, about 60 subagents, 31 million output tokens and a Lean proof — and Anthropic says it is not a route to proving the hypothesis itself.

A 160-Year-Old Bound, Moved Share of Riemann zeta zeros proven to satisfy the hypothesis 41.6% previous best 67.2% Claude's result 0% 100% 650 ideas that failed ~60 subagents 31M output tokens Lean formally verified BITSMINDS.COM
Share:

Anthropic asked an unreleased research version of Claude to take a real stab at the Riemann hypothesis. It did not solve it — nobody expected otherwise — but it produced something more unusual than a failed attempt: a genuine improvement to a related bound that number theorists have been inching forward for decades. Claude raised the proven lower bound on the fraction of Riemann zeta zeros that satisfy the hypothesis from 41.6% to 67.2%.

That number needs unpacking, because it is easy to overstate. The Riemann hypothesis asserts that every non-trivial zero of the zeta function lies on a single vertical line in the complex plane. Nobody can prove that holds for all of them, so mathematicians instead prove it holds for a guaranteed proportion. Getting that proportion above a half was a long-standing marker of progress, and the previous state of the art sat at 41.6% after decades of incremental work. Moving it to 67.2% does not make the hypothesis two-thirds proven in any meaningful sense — the remaining third is where all the difficulty lives — but it is a real, citable result rather than a benchmark score.

The route Claude took was synthesis rather than invention. It found that results from a series of papers published between 2023 and 2025 by Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh — work that adapted Montgomery's 1973 techniques so they no longer assume the Riemann hypothesis — could be combined with a 2000 paper by Bombieri to clear the old 41.6% ceiling. Getting there was not elegant. Claude generated and discarded roughly 650 ideas that went nowhere before the productive combination surfaced.

The process is arguably the more interesting disclosure. The work ran across two Claude Code sessions and consumed about 31 million output tokens, 2,400 shell commands and hundreds of Python scripts. Claude coordinated roughly 60 subagents with a rough division of labour: two developed the key mathematical ideas, 13 contributed ideas, 30 attempted new ones, 13 validated arguments and two helped write the paper. The prompt came from Jarred Sumner, an Anthropic staff member who is not a mathematician, and who left the mathematical choices to the model.

Crucially, the result was checked rather than asserted. The proof was formalised in Lean, which makes the argument machine-verifiable and closes off the usual failure mode where a model produces confident prose concealing a broken step. Anthropic mathematicians Levent Alpöge and Ralph Furman examined the work, and Brian Conrey and Dan Goldston reviewed the paper externally — Goldston being a co-author of the very research Claude built on, which is about as informed a referee as the result could get.

Anthropic is also explicit about the ceiling on all this. The company states it does not expect the techniques Claude used to lead to proving the Riemann hypothesis. The gap between improving a bound and settling the conjecture is not a matter of more tokens. The framing is a deliberate contrast with the noisier end of this genre: OpenAI's claim that its Astra family solved ten decade-old math problems drew scrutiny precisely over what counted as a solution and who had checked it.

What makes this result worth attention is not that a model did mathematics, but that the output survived contact with the two things that usually deflate such announcements: a formal verifier and the specialists whose prior work it depends on. That combination is still rare enough to be the story.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles