Google released Gemini 3.5 Transcribe on August 26, 2026, a new audio model that automatically detects more than 85 languages, strips filler words like 'um' and 'uh,' and attributes speech for up to three speakers in pre-recorded audio. The launch lands while Google's overdue Gemini 3.5 Pro, promised in June, still has no release date. It follows the recent debut of Gemini 3.5 Live Translate as Google fills out its audio stack piece by piece.
The pitch is a transcription model that behaves less like a captioner and more like an editor. Users can dictate and edit with their voice, feed the model a custom vocabulary for specialized jargon, and get word-level timestamps back with speaker attribution. Google says the accuracy gains show up most clearly in multilingual performance and word error rates, the two metrics that historically separate usable transcription from unusable transcription.
Google is positioning 3.5 Transcribe as a direct successor to Chirp 3, the company's previous transcription model.
Key facts
- 01Gemini 3.5 Transcribe supports more than 85 languages and auto-removes filler words like 'um' and 'uh.'
- 02The model attributes speech for up to three speakers in pre-recorded audio with word-level timestamps.
- 03Rolling out August 26, 2026 in English on macOS, on Android via Rambler, and in public preview through the Gemini API.
- 04Google says the model is a step up from Chirp 3 on multilingual performance and word error rates.
- 05Gemini 3.5 Pro, promised in June, has still not shipped.
Custom vocabulary is the underrated feature. Legal, medical, and engineering transcription workflows fail on proper nouns, acronyms, and domain-specific terms; letting users prime the model with a term list means those words survive the auto-formatting pass rather than getting silently rewritten. Combined with automatic filler-word removal, the output is closer to a first draft than a raw transcript, which is the actual bar for professional use.
The rollout is staggered. 3.5 Transcribe starts today in English for all macOS Gemini app users and inside the Rambler dictation feature on Android in select countries and languages. Developers get public preview access through the Gemini API via AI Studio and Antigravity. Chrome support is coming soon, per Google, without a firm date.
There was a stumble on launch-day messaging. Google initially told reporters that Gemini 3.5 Live and Gemini 3.5 Live Experimental would ship alongside 3.5 Transcribe, then walked that back after publication, saying only 3.5 Transcribe is being announced today and declining to provide a new launch date for the Live models. Gemini 3.5 Live was pitched as better at handling mid-sentence interruptions, language recognition, and live visual processing; Gemini 3.5 Live Experimental was described as narrating its reasoning step by step in real time.
The confusion feeds an ongoing critique of Gemini's product naming and release cadence. AI Chat Daily covered the branding sprawl across Spark, Daily Brief, and Gemini's chat surface last week, and the launch-day retraction of two of three announced models compounds the problem. Reporter Jess Weatherbed of The Verge put it plainly: 'We got a new Gemini Audio model while we're still waiting for the overdue Gemini 3.5 Pro launch.'
Competition in transcription is thick. OpenAI's Whisper family remains the default open-weights choice, and a growing set of dedicated players ship purpose-built dictation and meeting-transcription products. Google's advantage is distribution: the Gemini app on macOS and Android, plus API access through AI Studio, puts 3.5 Transcribe in front of developers and end users on day one without a separate signup.
The unknowns are the numbers Google didn't publish. There is no headline word error rate, no benchmark comparison against Chirp 3 with a specific delta, and no latency figure for the API path. 'Major advancement' is Google's phrasing; the field will judge it against Whisper large-v3 and against dedicated transcription vendors once developers run their own evals. Speaker attribution capped at three is also a real limit for meeting and podcast workflows that routinely involve four or more voices.
For Google, shipping a solid transcription model without Gemini 3.5 Pro underneath it is a tell. The audio stack is progressing independently of the flagship reasoning model, which suggests the Pro delay is a training or safety issue specific to that model rather than an infrastructure problem. If 3.5 Transcribe holds up under third-party benchmarks, it becomes the sharpest piece of Google's audio story and the easiest one for developers to adopt today — Pro or no Pro.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



