Google DeepMind released Gemini 3.5 Live Translate today, a speech-to-speech model that detects 70+ languages automatically and streams translated audio with a lag of just a few seconds. The model rolls out across three surfaces simultaneously: the Gemini Live API in public preview for developers, Google Meet in private preview for select Workspace customers this month, and the Google Translate app globally on Android and iOS.
The headline jump is inside Google Meet. The prior translation feature supported 5 languages and only paired them to and from English. Live Translate lifts that ceiling to 70+ languages and 2,000+ language combinations within a single meeting, which turns Meet from an English-pivot tool into a genuine multilingual conferencing product.
The architectural change is continuous generation. Earlier turn-by-turn systems waited for a speaker to finish before producing output, which created the stilted pauses familiar to anyone who has used interpretation software. Gemini 3.5 Live Translate generates speech as the source audio streams in, trading a small amount of context for sync, and preserving the speaker's intonation, pacing, and pitch in the output voice.
“Unlike turn by turn systems that wait for the speaker to finish speaking before responding, 3.5 Live Translate generates speech continuously, balancing the trade-off between waiting for context to improve quality and translating immediately to stay in sync with the speaker.”— Anuda Weerasinghe, Google DeepMind Product Manager
Key facts
- 01Gemini 3.5 Live Translate detects and translates 70+ languages in near real-time speech-to-speech, staying just a few seconds behind the speaker.
- 02Google Meet jumps from 5 supported languages to 70+, and from English-only pairs to 2,000+ language combinations in a single meeting.
- 03Grab is testing the model on the 10 million voice calls per month between drivers and riders on its platform.
- 04Google's translation systems now process over 1 trillion words per month across products, a milestone from a project started 20 years ago.
- 05Rolling out today to developers via the Gemini Live API and AI Studio, this month to Meet in private preview, and globally on Google Translate for Android and iOS.
Google says its translation systems now process over 1 trillion words per month across products, a scale reached 20 years after the company's first machine-learning translation experiments. Live Translate is the audio-native extension of that stack — built to handle noisy environments and mixed-language input without manual configuration.
The early enterprise tester named in the launch is Grab, the Southeast Asian ride-hailing and delivery company, which routes more than 10 million voice calls per month between drivers and travelers at pickup. Grab is piloting the model to bridge language gaps in those calls in near real-time. CJ ENM and infrastructure partners including LiveKit have also given the model early feedback citing translation quality and low latency.
On the developer side, Google is leaning on a partner ecosystem to handle the plumbing. Agora, Fishjam, LiveKit, Pipecat, and Vision Agents have integrations with the Gemini Live API that handle real-time media streaming, leaving developers to focus on product logic. Sample code lives in the Gemini Cookbook and Google AI Studio.
“Twenty years ago, translation at Google began as one of our pioneering machine learning experiments to turn the science of language into the magic of human connection.”— Tony Lu, Google DeepMind Senior Staff Software Engineer
For consumers, the Google Translate app gets the same model behind its Live translate feature on both Android and iOS. Plug in any headphones and the app mirrors the speaker's tone in the target language. Android users also get a new "listening mode" that routes the translated audio through the phone's earpiece, so a user can hold the device like a regular phone call and hear a translation of, for instance, a Spanish-language museum tour without others nearby hearing it.
All audio generated by the model is watermarked with SynthID, Google's imperceptible audio watermark designed to flag AI-generated speech downstream. That matters more for a translation model than for most generative audio products: the output is a synthetic voice speaking in someone else's name, and the watermark is the audit trail.
The caveats sit in the gap between demo and production. Continuous generation means the model commits to a translation before the speaker finishes the sentence, which trades the occasional correction or backtrack for fluency. Noise robustness is claimed but not benchmarked publicly, and the broader Meet rollout is gated to "later this year" after the private-preview phase wraps. Real-world latency at scale, across 2,000+ language pairs and unpredictable network conditions, is what the Grab pilot and the Meet preview will test.
Live, low-latency speech translation has been the holy grail since the original Google Translate demo, and shipping it across Meet, Translate, and an API on the same day is how Google converts research lead into product distribution. The bigger competitive signal is what this does to standalone translation startups and to the AI-meeting-assistant category: when 70+ languages and 2,000+ pairings ship inside Meet as a default feature, the bar for a separate product gets meaningfully higher, and the value migrates to whoever owns the meeting surface.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




