Skip to main content
Live
Main content

Category

Models

Foundation models, benchmarks, and the research race.

Models — page 2

DiG-bench and Faraday mark AI's slow climb toward recursive self-improvement
Models

DiG-bench and Faraday mark AI's slow climb toward recursive self-improvement

New benchmarks show frontier models cracking discovery games and replicating research papers, with Opus 5 and a 27B scientist model leading the way.

Jaeden Schafer5 min read
Google logo
Models

Google makes visible watermark on AI images, video, and audio optional

Nano Banana, Omni, and Lyria outputs can now ship without the visible tag; invisible SynthID and C2PA metadata stay on.

Jaeden Schafer4 min read
OpenAI logo
Models

OpenAI launches Ultrafast mode, pushing GPT-5.6 Sol to 750 tokens per second

The new preview mode runs at 14x standard speed, powered by a Cerebras partnership, and targets enterprise workflows where latency is the constraint.

Jaeden Schafer4 min read
Google logo
Models

Google ships Gemini 3.7 Flash at half the price of 3.6, three weeks later

The new workhorse model posts 43.6% on FrontierCode 1.1 and 65.3% on DeepSWE v1.1, with input tokens at $0.75 per million.

Jaeden Schafer5 min read
Google logo
Models

Google DeepMind ships sign language AI to Pixel 11 with new SL2T model

SL2T translates ASL to English on-device in Gboard and Live Transcribe, trained on 100,000+ hours across 50+ sign languages.

Jaeden Schafer5 min read
Anthropic logo
Models

Anthropic's unreleased model advances the Riemann hypothesis after 31M tokens

Coordinating 60 subagents across a day and a half, the model tested 650 ideas and improved the known bound on a 150-year-old problem.

Jaeden Schafer5 min read
OpenAI logo
Models

OpenAI's Astra model solves 10 long-open math problems for $2,000 in tokens

The unreleased model cracked problems that had eluded mathematicians for decades — and set off a credit dispute with the researchers whose work it built on.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard for agent workloads

The 30B mixture-of-experts model runs 4x faster on output, while the routing library cuts task cost to a third of Opus 4.8.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia ships Nemotron 3.5 Lightning, a 30B open MoE for local agents

The new open-weights model runs 4x faster than class rivals and slots into RTX PCs, DGX Spark, and Jetson for always-on agentic workloads.

Jaeden Schafer5 min read
Meta logo
Models

Meta pivots back to open weights with Muse Glimmer and a 6,000-word Zuckerberg essay

A 30B open-weight model, a promise to open Muse Spark 1.2, and a manifesto against 'singular superintelligence' mark Meta's latest AI reset.

Jaeden Schafer5 min read
Meta logo
Models

Meta releases Muse Glimmer, a 30B open-weight model for on-device AI agents

The Apache 2.0 model runs locally on a single consumer GPU and sketches Zuckerberg's 'personal superintelligence' pitch.

Jaeden Schafer5 min read
Anthropic logo
Models

Intology's Locus beats human baseline on PostTrainBench, signaling AI R&D automation

Locus hits 51.6% on PostTrainBench+ with 4,000+ H100 hours, surpassing the 51.1% human baseline and every frontier-agent competitor.

Jaeden Schafer5 min read
Google logo
Models

DeepMind's WeatherNext gives hurricane forecasters an extra day of lead time

The open-sourced AI model predicted Hurricane Melissa's Category 5 landfall in Jamaica five days out with 80% confidence.

Jaeden Schafer5 min read
ByteDance trains 10 trillion parameter model to rival Anthropic
Models

ByteDance trains 10 trillion parameter model to rival Anthropic

The TikTok parent's next model would be three times the size of Moonshot's Kimi K3 and larger than estimates for Anthropic's Mythos 5.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia releases Cosmos 3, an open world model family for physical AI

The three-tier release spans 4B to 64B parameters and tops five benchmarks for open-weights image, video, world and robot generation.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia opens Alpamayo 2 Super for commercial robotaxi deployment

The 30B-parameter driving model tops LingoQA against Qwen, Gemini and GPT-4o, and ships under a permissive Linux Foundation license.

Jaeden Schafer5 min read
Apple ships working Siri AI in iOS 27 beta, trained on Google's Gemini
Models

Apple ships working Siri AI in iOS 27 beta, trained on Google's Gemini

Siri finally understands personal context and holds a conversation, but arrives years after the AI assistant race moved on.

Jaeden Schafer5 min read
Alibaba releases Qwen3.8-Max, a 2.4-trillion-parameter model rivaling Claude
Models

Alibaba releases Qwen3.8-Max, a 2.4-trillion-parameter model rivaling Claude

Alibaba's largest model to date ranks second only to Anthropic's Fable 5 on Arena.AI, with open weights due next week.

Jaeden Schafer5 min read
Google logo
Models

Google DeepMind's Gemini Robotics 2 takes control of a humanoid from feet to fingertips

The update extends the model from upper-body manipulation to full-body motion, letting Apptronik's Apollo 2 walk, crouch and pick items off shelves.

Jaeden Schafer4 min read
Google logo
Models

Google DeepMind ships Gemini Robotics ER 2 as a video-native brain for robots

The new embodied reasoning model tracks task progress from live video, orchestrates other robots, and runs at 4x the speed of larger rivals.

Jaeden Schafer5 min read
Google logo
Models

Google DeepMind's Gemini Robotics 2 controls humanoids across dexterous tasks

The updated model fuses a vision-language brain with two action models to run robots like Apptronik's Apollo 2 on real chores.

Jaeden Schafer5 min read
Google logo
Models

Google DeepMind ships Lyria 3.5 in Flow Music with sharper vocals and lyrics

The new music model adds tempo and duration controls, better prompt adherence on lyrics, and more expressive vocals, rolling out today in Flow Music.

Jaeden Schafer4 min read
Anthropic logo
Models

Claude Opus 4.7 finishes a two-week coding job in 14 hours for $251

Epoch and METR's new MirrorCode benchmark shows frontier models reimplementing 61k-line codebases from scratch through CLI access alone.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia turns its own Vera CPU on the job of designing future Nvidia chips

Early tests with Cadence and Synopsys show 1.5x speedups on formal verification and functional simulation workloads.

Jaeden Schafer4 min read
AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at