Skip to main content
Live
Main content

Category

Models

Foundation models, benchmarks, and the research race.

Models — page 7

Google logo
Models

Google connects Street View to Genie 3 world model for interactive street simulation

The integration draws on 20 years of Street View data — 280 billion images across 110 countries — to let users simulate real-world streets with adjustable conditions.

Jaeden Schafer5 min read
Mira Murati previews Thinking Machines' interaction models with humans in the loop
Models

Mira Murati previews Thinking Machines' interaction models with humans in the loop

The ex-OpenAI CTO's lab shows AI that natively reads pauses, interruptions, and tone through camera and mic — its second product since launch.

Jaeden Schafer4 min read
Adaption launches AutoScientist, an AI tool that trains models on themselves
Models

Adaption launches AutoScientist, an AI tool that trains models on themselves

Co-founder Sara Hooker says the system doubled win-rates across models; Adaption is offering 30 days free to prove it.

Jaeden Schafer4 min read
Thinking Machines unveils 'full duplex' AI that listens while it talks
Models

Thinking Machines unveils 'full duplex' AI that listens while it talks

Mira Murati's startup says TML-Interaction-Small responds in 0.40 seconds, matching the cadence of natural human conversation.

Jaeden Schafer4 min read
Mira Murati's Thinking Machines unveils real-time 'interaction models'
Models

Mira Murati's Thinking Machines unveils real-time 'interaction models'

The startup says its new approach lets AI continuously process audio, video, and text instead of waiting turn-by-turn for users to finish.

Jaeden Schafer4 min read
Anthropic logo
Models

Anthropic blames sci-fi tropes for Claude's blackmail behavior in tests

The company says training on stories of AI behaving admirably cut blackmail attempts from up to 96% to zero in Claude Haiku 4.5.

Jaeden Schafer4 min read
OpenAI logo
Models

OpenAI adds GPT-Realtime-2, Translate and Whisper to its Realtime API

The new voice stack handles 70 input languages and 13 output languages, pushing the API beyond simple call-and-response.

Jaeden Schafer4 min read
Anthropic logo
Models

Anthropic Teaches Claude Code to 'Dream' Between Sessions

A new idle-time memory consolidation feature lets Claude Code review up to 100 prior sessions and lifts task success by up to 10 points.

Jaeden Schafer4 min read
OpenAI logo
Models

OpenAI's GPT-5.5 Instant Cuts Hallucinations 52% as Default ChatGPT Model

OpenAI replaces GPT-5.3 Instant with a new default model touting sharper accuracy in medicine, law, and finance, plus user-visible memory controls.

Jaeden Schafer4 min read
Genesis AI unveils GENE-26.5 model and in-house robotic hands after $105M seed
Models

Genesis AI unveils GENE-26.5 model and in-house robotic hands after $105M seed

The Khosla- and Eclipse-backed startup is building its own human-shaped hands and a sensor glove to close the embodiment gap in robotics data.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia opens Spectrum-X's MRC protocol via OCP after OpenAI, Microsoft, Oracle deployments

The Multipath Reliable Connection transport, proven on Blackwell training runs, is now an open spec through the Open Compute Project.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia opens Spectrum-X's MRC protocol to the industry after OpenAI, Microsoft, Oracle deployments

The Multipath Reliable Connection protocol, co-developed with OpenAI and Microsoft, reroutes failed network paths in microseconds across gigascale GPU fabrics.

Jaeden Schafer5 min read
OpenAI logo
Models

OpenAI's o1 Beats ER Doctors at Triage Diagnosis in Harvard Study

A Harvard Medical School study found OpenAI's o1 preview diagnosed ER triage cases correctly 67% of the time, ahead of attending physicians.

Jaeden Schafer4 min read
OpenAI logo
Models

OpenAI makes GPT-5.5 Instant the default ChatGPT model, claims 52.5% fewer hallucinations

The new default cuts hallucinated claims by more than half on medical, legal, and financial prompts and scores 81.2 on AIME 2025 math.

Jaeden Schafer4 min read
Harvard study: OpenAI's o1 beat ER doctors on triage diagnoses, 67% to 55%
Models

Harvard study: OpenAI's o1 beat ER doctors on triage diagnoses, 67% to 55%

In a 76-patient Beth Israel trial, o1 matched or outperformed two attending physicians and GPT-4o at the first diagnostic touchpoint.

Jaeden Schafer5 min read
Oxford Internet Institute
Models

Oxford study: warmer AI models are 60% more likely to be wrong

Fine-tuning five models including GPT-4o for empathy raised error rates 7.43 points and worsened sharply when users said they felt sad.

Jaeden Schafer5 min read
Synthetic biology
Models

Columbia and Harvard researchers cut E. coli's genetic code from 20 amino acids to 19

Using AI protein-design tools, the team rebuilt 20 of 21 ribosomal small-subunit genes to work without isoleucine, with cells growing at 60% normal speed.

Jaeden Schafer5 min read
Google logo
Models

Google DeepMind unveils AI co-clinician, pitching a new healthcare model

DeepMind frames the system as a working partner for doctors, building on healthcare AI research it began in July 2023.

Jaeden Schafer4 min read
OpenAI logo
Models

OpenAI Strikes Back With GPT-5.5 as Anthropic Hits $1tn Secondary Valuation

OpenAI's new model targets agentic coding work as Anthropic pulls ahead in enterprise and trades at a higher private-market valuation.

Jaeden Schafer4 min read
OpenAI's Codex system prompt tells GPT-5.5 to never mention goblins
Models

OpenAI's Codex system prompt tells GPT-5.5 to never mention goblins

A 3,500-word base instruction set leaked on GitHub bans talk of goblins, gremlins, and raccoons unless the user asks first.

Jaeden Schafer4 min read
SenseTime
Models

SenseTime open-sources SenseNova U1, an image model tuned for Chinese chips

Ten Chinese chip designers including Cambricon and Biren backed the model on day one as the sanctioned firm bets on open source to catch up.

Jaeden Schafer5 min read
Google logo
Models

Google DeepMind's Athletica Solves Six of Ten Novel Math Proofs at Publishable Quality

DeepMind's new autonomous math agent, built on Gemini 3 DeepThink, scored above 91.9% on the IMO proof benchmark.

Jaeden Schafer3 min read
Nvidia logo
Models

NVIDIA's Nemotron 3 Nano Omni claims 9x throughput edge over rival open multimodal models

The 30B-A3B mixture-of-experts model fuses vision, audio and text encoders, with a 256K context window and a 1920×1080 native input resolution.

Jaeden Schafer5 min read
DeepSeek's V4 lands with 1M-token context and a pivot to Huawei chips
Models

DeepSeek's V4 lands with 1M-token context and a pivot to Huawei chips

The open-source flagship matches GPT-5.4 and Claude-Opus-4.6 on benchmarks while pricing V4-Pro at $1.74 per million input tokens.

Jaeden Schafer5 min read
AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at