Google DeepMind's Gemini 3.1 Pro, rolled out globally through the Gemini app in early April, scored 94.3% on GPQA Diamond, a graduate-level science reasoning benchmark that has historically eaten frontier models alive. GPT-5.2 trails at 92.4%; Claude Opus 4.6 at 91.3%.
GPQA Diamond is the harder of two GPQA subsets. Questions are drawn from physics, chemistry, and biology at the PhD-qualifying-exam level; a domain expert in an unrelated subfield scores roughly 34% on it. A year ago, the best published frontier model sat at 78%.
Gemini 3.1 Pro is paired with a new side-panel desktop experience on Windows, macOS, and Chromebook Plus in the U.S., which lets the model operate across Google apps and the user's open browser tabs — the agentic computer-use pattern that OpenAI introduced with GPT-5.4 last month.
Key facts
- 01Google. A key thread of reporting in this story.
- 02Gemini. A key thread of reporting in this story.
- 03Benchmarks. A key thread of reporting in this story.
Google also launched Gemma 4 — an open-weights sibling family — and extended free Gemini access to students in Indonesia, Japan, the UK, and Brazil through July. The distribution playbook is aggressive by historical Google standards and explicit about its target: the education market, where ChatGPT currently dominates.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



