ai-models
ModelsLead story
Gemini 3.1 Pro sets a new reasoning high-water mark
Google's flagship now leads GPT-5.2 and Claude Opus 4.6 on GPQA Diamond, with 94.3% on a benchmark that has broken most prior models.
Jaeden Schafer4 min read
Tag · 1 story
Every story tagged Benchmarks on AI Chat Daily.
Google's flagship now leads GPT-5.2 and Claude Opus 4.6 on GPQA Diamond, with 94.3% on a benchmark that has broken most prior models.
The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.
The briefing read inside teams at