Skip to main content
Live
Main content

Google DeepMind's Athletica Solves Six of Ten Novel Math Proofs at Publishable Quality

DeepMind's new autonomous math agent, built on Gemini 3 DeepThink, scored above 91.9% on the IMO proof benchmark.

Jaeden Schafer
Editor in Chief · · 3 min read
Google logo
AICD
AI Chat Podcast

Sam Altman's Dispute with CFO Explored

Google DeepMind has released Athletica, an autonomous math agent built on top of Gemini 3 DeepThink, and the early benchmark results suggest the system is closing in on research-grade mathematics. On a curated set of unsolved and novel proof challenges, human expert evaluators graded six of Athletica's 10 submitted solutions as publishable after only minor revisions.

The launch lands in a field where progress has historically been measured in inches. Proving novel theorems is widely regarded as one of the hardest tests for an AI system because the model cannot pattern-match its way to an answer — it has to construct a chain of reasoning that withstands expert scrutiny. Athletica's evaluators were not grading for plausibility but for whether the proofs could survive peer review.

Alongside the proof challenge, Athletica posted a score above 91.9% on the IMO proof benchmark, a test built around the kind of formal proofs demanded at the International Mathematical Olympiad. That figure puts the agent firmly in the territory previously occupied only by specialist systems trained narrowly for olympiad-style problems, rather than general-purpose reasoning models.

Key facts

  • 01Google DeepMind released Athletica this week, an autonomous math agent built on top of Gemini 3 DeepThink.
  • 02Human expert evaluators graded Athletica's solutions as publishable after minor revisions on six of 10 novel proof challenge problems.
  • 03Athletica scored above 91.9% on the IMO proof benchmark, a test focused on formal mathematical proofs.
  • 04Andrej Mirha of Andreessen Horowitz called the result the moment AI math went from playing the game to writing the rules.

The architecture matters here. Athletica is not a from-scratch math model but an agent layer sitting on top of Gemini 3 DeepThink, Google's extended-reasoning variant of its flagship Gemini 3 family. That suggests DeepMind is betting that frontier reasoning models, paired with the right agent scaffolding, can attack problems that until recently required bespoke architectures like AlphaProof.

AI math went from playing the game to writing the rules.
Jaeden Schafer

Reaction from the venture side has been pointed. Anjay Mirha at Andreessen Horowitz called it the moment "AI math went from playing the game to writing the rules," a framing Jaeden Schafer endorsed on the podcast. "I think that's kind of a fair framing because up until this point with all of the different AI models proving these kind of novel theories is one of the hardest possible tests for these AI agents," Schafer said. "And Athletica just did it."

The stakes extend well beyond olympiad scoreboards. If autonomous agents can reliably draft publishable proofs of novel results, the timeline for AI contributing original work to mathematics — and, by extension, to theoretical physics, cryptography and algorithm design — compresses sharply. Schafer noted on the show that the proofs Athletica handled were "all these really complex math things to a high level, which was basically ready to get published."

DeepMind has not yet detailed how widely Athletica will be available or whether the underlying agent stack will surface in consumer Gemini products. But the result raises the bar for rivals at OpenAI and Anthropic, both of which have leaned heavily on math and reasoning benchmarks to market their latest frontier models. The competitive pressure now shifts from beating contest problems to producing work that mathematicians are willing to put their names on.

Related · from this week
Google DeepMind's WeatherNext 3 delivers hourly forecasts at 5-kilometer resolution
Jaeden Schafer · 5 min read →
ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Google logo
Models

Google DeepMind's WeatherNext 3 delivers hourly forecasts at 5-kilometer resolution

The new model trains on live satellite data instead of physics simulations, cutting the six-hour lag that plagues traditional forecasting.

Jaeden Schafer5 min read
Google logo
Models

Google DeepMind ships Gemini Omni 1.1 Flash with 4K video and scene extension

The Omni update pushes generative video to 4K, extends clips to 40 seconds, and cuts draft costs to a third at 360p.

Jaeden Schafer5 min read
Google logo
Models

Google ships Nano Banana 2 Lite and Gemini Omni Flash to developers

DeepMind's fastest image model generates in 4 seconds at $0.034 per 1K images; Omni Flash matches Veo 3.1 Fast at $0.10 per second of video.

Jaeden Schafer5 min read