Models — page 2

DiG-bench and Faraday mark AI's slow climb toward recursive self-improvement
New benchmarks show frontier models cracking discovery games and replicating research papers, with Opus 5 and a 27B scientist model leading the way.

Google makes visible watermark on AI images, video, and audio optional
Nano Banana, Omni, and Lyria outputs can now ship without the visible tag; invisible SynthID and C2PA metadata stay on.

OpenAI launches Ultrafast mode, pushing GPT-5.6 Sol to 750 tokens per second
The new preview mode runs at 14x standard speed, powered by a Cerebras partnership, and targets enterprise workflows where latency is the constraint.

Google ships Gemini 3.7 Flash at half the price of 3.6, three weeks later
The new workhorse model posts 43.6% on FrontierCode 1.1 and 65.3% on DeepSWE v1.1, with input tokens at $0.75 per million.

Google DeepMind ships sign language AI to Pixel 11 with new SL2T model
SL2T translates ASL to English on-device in Gboard and Live Transcribe, trained on 100,000+ hours across 50+ sign languages.

Anthropic's unreleased model advances the Riemann hypothesis after 31M tokens
Coordinating 60 subagents across a day and a half, the model tested 650 ideas and improved the known bound on a 150-year-old problem.

OpenAI's Astra model solves 10 long-open math problems for $2,000 in tokens
The unreleased model cracked problems that had eluded mathematicians for decades — and set off a credit dispute with the researchers whose work it built on.

Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard for agent workloads
The 30B mixture-of-experts model runs 4x faster on output, while the routing library cuts task cost to a third of Opus 4.8.

Nvidia ships Nemotron 3.5 Lightning, a 30B open MoE for local agents
The new open-weights model runs 4x faster than class rivals and slots into RTX PCs, DGX Spark, and Jetson for always-on agentic workloads.

Meta pivots back to open weights with Muse Glimmer and a 6,000-word Zuckerberg essay
A 30B open-weight model, a promise to open Muse Spark 1.2, and a manifesto against 'singular superintelligence' mark Meta's latest AI reset.

Meta releases Muse Glimmer, a 30B open-weight model for on-device AI agents
The Apache 2.0 model runs locally on a single consumer GPU and sketches Zuckerberg's 'personal superintelligence' pitch.

Intology's Locus beats human baseline on PostTrainBench, signaling AI R&D automation
Locus hits 51.6% on PostTrainBench+ with 4,000+ H100 hours, surpassing the 51.1% human baseline and every frontier-agent competitor.

DeepMind's WeatherNext gives hurricane forecasters an extra day of lead time
The open-sourced AI model predicted Hurricane Melissa's Category 5 landfall in Jamaica five days out with 80% confidence.

ByteDance trains 10 trillion parameter model to rival Anthropic
The TikTok parent's next model would be three times the size of Moonshot's Kimi K3 and larger than estimates for Anthropic's Mythos 5.

Nvidia releases Cosmos 3, an open world model family for physical AI
The three-tier release spans 4B to 64B parameters and tops five benchmarks for open-weights image, video, world and robot generation.

Nvidia opens Alpamayo 2 Super for commercial robotaxi deployment
The 30B-parameter driving model tops LingoQA against Qwen, Gemini and GPT-4o, and ships under a permissive Linux Foundation license.

Apple ships working Siri AI in iOS 27 beta, trained on Google's Gemini
Siri finally understands personal context and holds a conversation, but arrives years after the AI assistant race moved on.

Alibaba releases Qwen3.8-Max, a 2.4-trillion-parameter model rivaling Claude
Alibaba's largest model to date ranks second only to Anthropic's Fable 5 on Arena.AI, with open weights due next week.

Google DeepMind's Gemini Robotics 2 takes control of a humanoid from feet to fingertips
The update extends the model from upper-body manipulation to full-body motion, letting Apptronik's Apollo 2 walk, crouch and pick items off shelves.

Google DeepMind ships Gemini Robotics ER 2 as a video-native brain for robots
The new embodied reasoning model tracks task progress from live video, orchestrates other robots, and runs at 4x the speed of larger rivals.

Google DeepMind's Gemini Robotics 2 controls humanoids across dexterous tasks
The updated model fuses a vision-language brain with two action models to run robots like Apptronik's Apollo 2 on real chores.

Google DeepMind ships Lyria 3.5 in Flow Music with sharper vocals and lyrics
The new music model adds tempo and duration controls, better prompt adherence on lyrics, and more expressive vocals, rolling out today in Flow Music.

Claude Opus 4.7 finishes a two-week coding job in 14 hours for $251
Epoch and METR's new MirrorCode benchmark shows frontier models reimplementing 61k-line codebases from scratch through CLI access alone.

Nvidia turns its own Vera CPU on the job of designing future Nvidia chips
Early tests with Cadence and Synopsys show 1.5x speedups on formal verification and functional simulation workloads.
Stay ahead of everyone in AI.
The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.
The briefing read inside teams at