Skip to main content
Live
Main content

Anthropic's unreleased model advances the Riemann hypothesis after 31M tokens

Coordinating 60 subagents across a day and a half, the model tested 650 ideas and improved the known bound on a 150-year-old problem.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

An unreleased Anthropic model has made verified progress on the Riemann hypothesis, a 150-year-old problem in number theory carrying a $1 million bounty. The model raised the lower bound of solutions for which the hypothesis holds true, testing 650 different ideas across 60 coordinated subagents and burning 31 million output tokens in the process. Anthropic disclosed the result on Monday, and two of its in-house mathematicians confirmed the work, which was then formalized in the open-source proof assistant Lean.

The setup is nearly as striking as the result. An Anthropic staff member without significant mathematical training prompted the model to take a real stab at the hypothesis and let it run for roughly a day and a half. The model orchestrated the entire multi-agent workflow itself — assigning idea generation, validation, and drafting to different subagents — without a mathematician steering the search.

The subagent breakdown is unusually specific. Two subagents produced the key mathematical ideas. Thirteen fed ideas into those two. Thirty more tried to contribute and failed. Thirteen served as validators, checking correctness. The final two drafted the paper.

Key facts

  • 01An unreleased [Anthropic](/claude) model raised the known lower bound of solutions for which the Riemann hypothesis holds, a 150-year-old problem carrying a $1M bounty.
  • 02The run tested 650 distinct ideas across 60 coordinated subagents and consumed 31 million output tokens over roughly a day and a half.
  • 03Two Anthropic in-house mathematicians confirmed the result, which was formalized in the open-source proof assistant Lean.
  • 04[OpenAI](/openai) recently disclosed 10 major results proved by its internal Astra model, and a separate Anthropic effort disproved the Jacobian conjecture.

That structure hints at how frontier labs are stretching test-time compute into something closer to a research team's org chart. Anthropic did not disclose which model powered the run, only that it has not yet been released publicly.

The result lands in a year that has already reshaped expectations for AI in mathematics. Several long-standing Erdős problems have fallen to language models in 2026. OpenAI disclosed 10 major results proved by its internal Astra model — which AI Chat Daily covered when the numbers first surfaced — and a separate Anthropic effort disproved the Jacobian conjecture, a problem that had stood since 1939.

The Riemann hypothesis, first posed in 1859, concerns the distribution of prime numbers and underpins large parts of modern number theory. No general proof exists. Anthropic's model did not produce one; it extended the range over which the hypothesis is verified, a legitimate incremental result of the kind mathematicians publish regularly, only produced by software rather than a person.

The mathematical community is split on what to do with results like this. In June, a group of prominent mathematicians signed a public declaration warning that AI-driven proofs could undermine a core value of the field: that theorems should be attributable to specific authors who take credit for their discovery and assume responsibility for their correctness. A machine-generated proof coordinated across 60 subagents fits awkwardly into that framework.

Fields Medal winner Timothy Gowers pushed back on the pessimistic reading in a blog post responding to the declaration. He argued the influence of AI on mathematics might turn out to be more complicated and, on balance, positive.

If we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won't be any more problematic than the fact that stars aren't named after astronomers and most aren't named at all.
Timothy Gowers, Fields Medal winner
Related · from this week
OpenAI's Astra model solves 10 long-open math problems for $2,000 in tokens
Jaeden Schafer · 5 min read →

There are real caveats. Anthropic ran the experiment on a model it has not shipped, so outside researchers cannot yet reproduce it. The 650 ideas tested include 30 subagents that produced nothing usable, meaning the raw hit rate on genuinely novel mathematical ideas remains low. And formal verification in Lean confirms the argument is valid — not that the approach generalizes to a full proof of the hypothesis. The $1 million bounty is safe.

The competitive read is that Anthropic and OpenAI are now visibly racing on frontier mathematics as a proxy for reasoning depth, and both are willing to spend eight-figure token counts on single problems to show it. If a day-and-a-half run producing a Lean-verified result is now the cost of a demo, the ceiling on what unreleased models can do in domains with clean verification — mathematics, formal software proofs, theorem-heavy physics — is climbing faster than most benchmark leaderboards suggest. The interesting question is no longer whether models can contribute to research mathematics. It's who gets credit, and how the field's authorship norms bend to accommodate a coauthor that runs on 60 subagents.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

OpenAI logo
Models

OpenAI's Astra model solves 10 long-open math problems for $2,000 in tokens

The unreleased model cracked problems that had eluded mathematicians for decades — and set off a credit dispute with the researchers whose work it built on.

Jaeden Schafer5 min read
Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model
Models

Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model

The London lab, fresh off a $50M seed, says its DeepMind-alumni-built agent matches frontier systems using a fraction of the parameters.

Jaeden Schafer5 min read
China's Z.ai claims GLM-5.2 matches Anthropic's Mythos on bug-finding
Models

China's Z.ai claims GLM-5.2 matches Anthropic's Mythos on bug-finding

The open-weight model lags on general tasks but reportedly closes the gap on cybersecurity — the exact capability US export curbs were meant to contain.

Jaeden Schafer4 min read