Skip to main content
Live
Main content

Google launches Gemini 3.5 Transcribe with 2.6% word error rate

The new speech-to-text model hits a 4.0% streaming WER, cuts time to final transcription by 70%, and supports over 85 languages.

Jaeden Schafer
Editor in Chief · · 5 min read
Google logo

Google DeepMind released Gemini 3.5 Transcribe on August 26, 2026, a speech-to-text model that posts a 2.6% Word Error Rate on pre-recorded audio and 4.0% on real-time streaming, according to benchmarks by Artificial Analysis. The model cuts time to final transcription by 70% versus Google's previous Chirp 3 system and is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.

The launch pushes Google into a transcription market where accuracy differences of a single percentage point translate directly into fewer downstream edits for developers building voice agents, captioning tools, and call-analytics pipelines. Gemini 3.5 Transcribe auto-detects over 85 languages, handles up to three speakers with word-level timestamps, and offers 3-plus speaker attribution in experimental form.

Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
Diego Melendo Casado, Senior Director, Engineering, Gemini Audio at Google DeepMind

The model splits into two API surfaces. Real-time streaming runs via the Live API under the identifier gemini-3.5-transcribe-live, with sub-second latency for interactive voice apps. Pre-recorded processing runs through the Interactions API as gemini-3.5-transcribe, aimed at meetings, call logs, and archival audio.

Key facts

  • 01Gemini 3.5 Transcribe posts a 4.0% Word Error Rate for streaming and 2.6% for non-streaming, as measured by Artificial Analysis.
  • 02Time to final transcription improves 70% over Google's prior Chirp 3 model.
  • 03On the FLEURS multilingual benchmark, the model hits 5.50% WER streaming and 5.04% non-streaming across top languages.
  • 04The model auto-detects over 85 languages and supports speaker attribution for up to three speakers in pre-recorded audio.
  • 05Available via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform starting August 26, 2026.

On the FLEURS multilingual benchmark, Gemini 3.5 Transcribe posts a 5.50% WER in streaming mode and 5.04% in non-streaming across a set of top languages and locales. Both figures improve on Chirp 3, though they trail the English-heavy averages Google is publishing for the flagship benchmarks — a reminder that multilingual accuracy still lags the best-case English numbers by roughly 2 to 3 percentage points.

Google is pitching the model less as raw transcription and more as intent capture. It cleans self-corrections like "let's meet Tuesday — no, Wednesday," strips filler words such as "ums" and "ahs," and auto-formats output. The model can also delegate downstream work — image generation, file analysis — to other Gemini models through function calls, a capability currently live in the Gemini app on macOS.

Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.
Luke Leonhard, Chief of Staff, Gemini Audio at Google DeepMind

The consumer surfaces went live in parallel. Rambler, the voice-to-text feature in Gboard on Android, now runs on 3.5 Transcribe in select countries. The Gemini app on macOS uses the model to pair voice commands with screen context, letting users summarize local files or generate images at the cursor. Chrome will get voice-to-type in any web field "coming soon," per Google.

For developer distribution, Google is leaning on the platform ecosystem around the Live API. Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents are all shipping integrations that handle the real-time media streaming layer so developers can focus on the interface. Early enterprise customers named by Google include Vivo, Intellitek Health, and Lingopal.

Custom vocabulary handling is the feature most likely to matter to enterprise buyers. The model adapts to provided jargon and unique spellings — the kind of domain-specific accuracy that has kept vertical transcription vendors in business against general-purpose speech-to-text. Alphanumeric strings like postal codes and order IDs also get specific treatment, an area where Chirp 3 struggled in noisy real-world audio.

Related · from this week
Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model
Jaeden Schafer · 5 min read →

The gap to close is against OpenAI's Whisper family and the specialized speech vendors — Deepgram, AssemblyAI, and Rev — that have dominated the enterprise transcription market. Whisper's open weights set the accuracy floor for self-hosted deployment; the specialized vendors compete on domain tuning and pricing. Google has not published API pricing for 3.5 Transcribe in the launch materials, which will decide whether developers migrate off existing pipelines.

The 70% latency improvement over Chirp 3 is the number that matters for voice-agent builders, where user-perceived responsiveness determines whether a product feels usable. Sub-second turnaround from raw audio to formatted, disfluency-cleaned text is what closes the gap between transcription-as-plumbing and voice-as-interface. Google's bet is that transcription is no longer a discrete product category but a substrate for agent workflows — which is why 3.5 Transcribe ships with function calling baked in rather than as a bolt-on. If that framing holds, the specialized transcription vendors have a harder pitch on their hands than a benchmark spreadsheet would suggest.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model
Models

Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model

The London lab, fresh off a $50M seed, says its DeepMind-alumni-built agent matches frontier systems using a fraction of the parameters.

Jaeden Schafer5 min read
Google logo
Models

DeepMind's SIMA 2 agent plays, reasons and learns in 3D game worlds

Google DeepMind extends 15 years of games research from Atari to EVE Online with a new agent that collaborates with human players.

Jaeden Schafer4 min read
Google logo
Models

Google bakes computer use into Gemini 3.5 Flash as a native tool

DeepMind folds its standalone agent model into Flash, letting developers build agents that drive browsers, mobile apps and desktops via one API call.

Jaeden Schafer4 min read