Skip to main content
Live
Main content

Google DeepMind ships Gemini 3.8 Live with parallel reasoning voice mode

The new Live Extended Thinking model tops Artificial Analysis's Speech-to-Speech index at 82.6 and handles 97 languages mid-conversation.

Jaeden Schafer
Editor in Chief · · 5 min read
Google logo

Google DeepMind launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, two voice-first models that reason and speak in parallel and topped Artificial Analysis's Speech to Speech Quality Index at 82.6. The Extended Thinking variant also leads agentic voice benchmarks with 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark, positioning it as the enterprise-grade tier while the base 3.8 Live model handles high-volume, cost-sensitive workloads.

Both models are shipping today into the Gemini API and Google AI Studio for developers, private preview in Gemini Enterprise for businesses, and Search Live, the Gemini app, Docs Live, Gmail Live, and Keep Live for consumers. Google AI Pro and Ultra subscribers get Extended Thinking inside Google Workspace immediately.

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet.
Tom Ouyang, Principal Engineer, Google DeepMind

The technical pitch is that voice agents no longer have to pause while they think. Extended Thinking narrates its own progress with cues like "Let me check that…" while running multi-step background tasks, and executes tool calls and API requests without dropping the conversation. Gemini 3.8 Live processes visual inputs in near real-time and automatically switches between 97 supported languages mid-sentence.

Key facts

  • 01Gemini 3.8 Live Extended Thinking took the #1 spot on Artificial Analysis's Speech to Speech Quality Index with a score of 82.6.
  • 02The model hit 68.6% on τ-Voice agentic task completion and 35.1% on Sierra's τ-Voice-banking benchmark.
  • 03Gemini 3.8 Live automatically detects and transitions between 97 supported languages mid-conversation.
  • 04The base 3.8 Live model scored 97.7% on Big Bench Audio and took second place in the Speech Agent Arena.
  • 05Both models launched September 15, 2026, via the Gemini API, Google AI Studio, and private preview in Gemini Enterprise.

On raw reasoning, 3.8 Live scored 97.7% on Big Bench Audio and took second place in the Speech Agent Arena, a user-preference leaderboard. On ServiceNow's EVA-Bench, run through the Gemini Enterprise Agent Platform, Google claims both models push the accuracy-versus-conversational-quality Pareto frontier for complex workflows.

The pricing story is unusually direct for a frontier release. Google is holding Extended Thinking at what it describes as a highly competitive price point against other frontier voice models, and pitching base 3.8 Live explicitly on cost efficiency at scale. That matters because voice agents in production get expensive fast — a customer-service deployment answering millions of calls a month lives or dies on per-minute economics.

These models handle complex reasoning, real-time visual context, and background task execution without interrupting your conversation.
Malini Jaganathan, Member of Technical Staff, Gemini Audio Team

Malini Jaganathan of the Gemini Audio Team framed the design around not breaking flow, and the demos back that up: one shows the model playing chess using live visual context, another turns whiteboard sketches into working React components through voice feedback alone, and a third builds full business plans on the fly.

On the ecosystem side, the Gemini Live API is now integrated with voice infrastructure platforms Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents, which handle the real-time media streaming so developers can focus on UX. Google also named Salesforce, Genspark, and Lumeris as launch enterprise partners citing latency, fluidity, and tool-calling. All audio output carries SynthID watermarking.

The obvious counterweight: benchmark leadership on Artificial Analysis and τ-Voice doesn't automatically translate into production reliability, and 35.1% on the banking-specific τ-Voice variant is a reminder that regulated verticals remain unsolved territory. Even the best agentic voice models today fail more than half the time on realistic financial-services task completion. Enterprises evaluating this against existing voice stacks will want their own domain evals before committing.

Related · from this week
Google DeepMind's WeatherNext 3 delivers hourly forecasts at 5-kilometer resolution
Jaeden Schafer · 5 min read →

This launch slots into a broader shift Google DeepMind has been telegraphing: voice as the primary interface for AI agents, not as a bolt-on. The company is now selling three surfaces at once — the API for builders, Gemini Enterprise for corporate deployments, and Live inside Search and Workspace for hundreds of millions of consumer users. Each reinforces the others' data flywheel.

The competitive read is that Google is trying to make voice the layer where it wins the agent race, even in categories where OpenAI and Anthropic lead on text benchmarks. Parallel reasoning that speaks while it thinks is a genuine product-experience wedge, and shipping it simultaneously into consumer, enterprise, and developer channels — at a price point Google is willing to defend — signals that DeepMind now views voice agents as the near-term battleground where model quality directly converts into paid-seat expansion inside Workspace and Gemini Enterprise.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Google logo
Models

Google DeepMind's WeatherNext 3 delivers hourly forecasts at 5-kilometer resolution

The new model trains on live satellite data instead of physics simulations, cutting the six-hour lag that plagues traditional forecasting.

Jaeden Schafer5 min read
Google logo
Models

Google ships three new Gemini models but delays Pro update

Gemini 3.6 Flash cuts token usage by up to 17%, but the flagship Pro refresh is still stuck as OpenAI and Anthropic pull ahead.

Jaeden Schafer5 min read
Google logo
Models

Google ships Nano Banana 2 Lite and Gemini Omni Flash to developers

DeepMind's fastest image model generates in 4 seconds at $0.034 per 1K images; Omni Flash matches Veo 3.1 Fast at $0.10 per second of video.

Jaeden Schafer5 min read