Skip to main content
Live
Main content

Gemini 3.1 Pro sets a new reasoning high-water mark

Google's flagship now leads GPT-5.2 and Claude Opus 4.6 on GPQA Diamond, with 94.3% on a benchmark that has broken most prior models.

Jaeden Schafer
· 4 min read

Google DeepMind's Gemini 3.1 Pro, rolled out globally through the Gemini app in early April, scored 94.3% on GPQA Diamond, a graduate-level science reasoning benchmark that has historically eaten frontier models alive. GPT-5.2 trails at 92.4%; Claude Opus 4.6 at 91.3%.

GPQA Diamond is the harder of two GPQA subsets. Questions are drawn from physics, chemistry, and biology at the PhD-qualifying-exam level; a domain expert in an unrelated subfield scores roughly 34% on it. A year ago, the best published frontier model sat at 78%.

Gemini 3.1 Pro is paired with a new side-panel desktop experience on Windows, macOS, and Chromebook Plus in the U.S., which lets the model operate across Google apps and the user's open browser tabs — the agentic computer-use pattern that OpenAI introduced with GPT-5.4 last month.

Key facts

  • 01Google. A key thread of reporting in this story.
  • 02Gemini. A key thread of reporting in this story.
  • 03Benchmarks. A key thread of reporting in this story.

Google also launched Gemma 4 — an open-weights sibling family — and extended free Gemini access to students in Indonesia, Japan, the UK, and Brazil through July. The distribution playbook is aggressive by historical Google standards and explicit about its target: the education market, where ChatGPT currently dominates.

Related · from this week
Google's Gemini 3.5 Transcribe strips filler words across 85+ languages
Jaeden Schafer · 4 min read →
ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Google logo
Models

Google's Gemini 3.5 Transcribe strips filler words across 85+ languages

The new audio model handles specialized jargon, tags up to three speakers, and lands while Gemini 3.5 Pro remains overdue.

Jaeden Schafer4 min read
Google logo
Models

Google makes visible watermark on AI images, video, and audio optional

Nano Banana, Omni, and Lyria outputs can now ship without the visible tag; invisible SynthID and C2PA metadata stay on.

Jaeden Schafer4 min read
Google logo
Tools

Google Photos adds Video Remix, an AI video editor powered by Gemini Omni

The feature applies cinematic relighting, background swaps, and painterly styles to clips in a few taps, and starts rolling out today across 15 countries.

Jaeden Schafer4 min read