An unreleased Anthropic model has made verified progress on the Riemann hypothesis, a 150-year-old problem in number theory carrying a $1 million bounty. The model raised the lower bound of solutions for which the hypothesis holds true, testing 650 different ideas across 60 coordinated subagents and burning 31 million output tokens in the process. Anthropic disclosed the result on Monday, and two of its in-house mathematicians confirmed the work, which was then formalized in the open-source proof assistant Lean.
The setup is nearly as striking as the result. An Anthropic staff member without significant mathematical training prompted the model to take a real stab at the hypothesis and let it run for roughly a day and a half. The model orchestrated the entire multi-agent workflow itself — assigning idea generation, validation, and drafting to different subagents — without a mathematician steering the search.
The subagent breakdown is unusually specific. Two subagents produced the key mathematical ideas. Thirteen fed ideas into those two. Thirty more tried to contribute and failed. Thirteen served as validators, checking correctness. The final two drafted the paper.
Key facts
- 01An unreleased [Anthropic](/claude) model raised the known lower bound of solutions for which the Riemann hypothesis holds, a 150-year-old problem carrying a $1M bounty.
- 02The run tested 650 distinct ideas across 60 coordinated subagents and consumed 31 million output tokens over roughly a day and a half.
- 03Two Anthropic in-house mathematicians confirmed the result, which was formalized in the open-source proof assistant Lean.
- 04[OpenAI](/openai) recently disclosed 10 major results proved by its internal Astra model, and a separate Anthropic effort disproved the Jacobian conjecture.
That structure hints at how frontier labs are stretching test-time compute into something closer to a research team's org chart. Anthropic did not disclose which model powered the run, only that it has not yet been released publicly.
The result lands in a year that has already reshaped expectations for AI in mathematics. Several long-standing Erdős problems have fallen to language models in 2026. OpenAI disclosed 10 major results proved by its internal Astra model — which AI Chat Daily covered when the numbers first surfaced — and a separate Anthropic effort disproved the Jacobian conjecture, a problem that had stood since 1939.
The Riemann hypothesis, first posed in 1859, concerns the distribution of prime numbers and underpins large parts of modern number theory. No general proof exists. Anthropic's model did not produce one; it extended the range over which the hypothesis is verified, a legitimate incremental result of the kind mathematicians publish regularly, only produced by software rather than a person.
The mathematical community is split on what to do with results like this. In June, a group of prominent mathematicians signed a public declaration warning that AI-driven proofs could undermine a core value of the field: that theorems should be attributable to specific authors who take credit for their discovery and assume responsibility for their correctness. A machine-generated proof coordinated across 60 subagents fits awkwardly into that framework.
Fields Medal winner Timothy Gowers pushed back on the pessimistic reading in a blog post responding to the declaration. He argued the influence of AI on mathematics might turn out to be more complicated and, on balance, positive.
“If we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won't be any more problematic than the fact that stars aren't named after astronomers and most aren't named at all.”— Timothy Gowers, Fields Medal winner
There are real caveats. Anthropic ran the experiment on a model it has not shipped, so outside researchers cannot yet reproduce it. The 650 ideas tested include 30 subagents that produced nothing usable, meaning the raw hit rate on genuinely novel mathematical ideas remains low. And formal verification in Lean confirms the argument is valid — not that the approach generalizes to a full proof of the hypothesis. The $1 million bounty is safe.
The competitive read is that Anthropic and OpenAI are now visibly racing on frontier mathematics as a proxy for reasoning depth, and both are willing to spend eight-figure token counts on single problems to show it. If a day-and-a-half run producing a Lean-verified result is now the cost of a demo, the ceiling on what unreleased models can do in domains with clean verification — mathematics, formal software proofs, theorem-heavy physics — is climbing faster than most benchmark leaderboards suggest. The interesting question is no longer whether models can contribute to research mathematics. It's who gets credit, and how the field's authorship norms bend to accommodate a coauthor that runs on 60 subagents.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




