Skip to main content
Live
Main content

Google DeepMind ships Gemini 3.5 Live Translate across 70+ languages

The new audio model streams speech-to-speech translation continuously, lagging the speaker by just a few seconds across Meet, Translate, and the Live API.

Jaeden Schafer
Editor in Chief · · 5 min read
Google logo

Google DeepMind released Gemini 3.5 Live Translate today, a speech-to-speech model that detects 70+ languages automatically and streams translated audio with a lag of just a few seconds. The model rolls out across three surfaces simultaneously: the Gemini Live API in public preview for developers, Google Meet in private preview for select Workspace customers this month, and the Google Translate app globally on Android and iOS.

The headline jump is inside Google Meet. The prior translation feature supported 5 languages and only paired them to and from English. Live Translate lifts that ceiling to 70+ languages and 2,000+ language combinations within a single meeting, which turns Meet from an English-pivot tool into a genuine multilingual conferencing product.

The architectural change is continuous generation. Earlier turn-by-turn systems waited for a speaker to finish before producing output, which created the stilted pauses familiar to anyone who has used interpretation software. Gemini 3.5 Live Translate generates speech as the source audio streams in, trading a small amount of context for sync, and preserving the speaker's intonation, pacing, and pitch in the output voice.

Unlike turn by turn systems that wait for the speaker to finish speaking before responding, 3.5 Live Translate generates speech continuously, balancing the trade-off between waiting for context to improve quality and translating immediately to stay in sync with the speaker.
Anuda Weerasinghe, Google DeepMind Product Manager

Key facts

  • 01Gemini 3.5 Live Translate detects and translates 70+ languages in near real-time speech-to-speech, staying just a few seconds behind the speaker.
  • 02Google Meet jumps from 5 supported languages to 70+, and from English-only pairs to 2,000+ language combinations in a single meeting.
  • 03Grab is testing the model on the 10 million voice calls per month between drivers and riders on its platform.
  • 04Google's translation systems now process over 1 trillion words per month across products, a milestone from a project started 20 years ago.
  • 05Rolling out today to developers via the Gemini Live API and AI Studio, this month to Meet in private preview, and globally on Google Translate for Android and iOS.

Google says its translation systems now process over 1 trillion words per month across products, a scale reached 20 years after the company's first machine-learning translation experiments. Live Translate is the audio-native extension of that stack — built to handle noisy environments and mixed-language input without manual configuration.

The early enterprise tester named in the launch is Grab, the Southeast Asian ride-hailing and delivery company, which routes more than 10 million voice calls per month between drivers and travelers at pickup. Grab is piloting the model to bridge language gaps in those calls in near real-time. CJ ENM and infrastructure partners including LiveKit have also given the model early feedback citing translation quality and low latency.

On the developer side, Google is leaning on a partner ecosystem to handle the plumbing. Agora, Fishjam, LiveKit, Pipecat, and Vision Agents have integrations with the Gemini Live API that handle real-time media streaming, leaving developers to focus on product logic. Sample code lives in the Gemini Cookbook and Google AI Studio.

Twenty years ago, translation at Google began as one of our pioneering machine learning experiments to turn the science of language into the magic of human connection.
Tony Lu, Google DeepMind Senior Staff Software Engineer

For consumers, the Google Translate app gets the same model behind its Live translate feature on both Android and iOS. Plug in any headphones and the app mirrors the speaker's tone in the target language. Android users also get a new "listening mode" that routes the translated audio through the phone's earpiece, so a user can hold the device like a regular phone call and hear a translation of, for instance, a Spanish-language museum tour without others nearby hearing it.

All audio generated by the model is watermarked with SynthID, Google's imperceptible audio watermark designed to flag AI-generated speech downstream. That matters more for a translation model than for most generative audio products: the output is a synthetic voice speaking in someone else's name, and the watermark is the audit trail.

Related · from this week
Google ships three new Gemini models but delays Pro update
Jaeden Schafer · 5 min read →

The caveats sit in the gap between demo and production. Continuous generation means the model commits to a translation before the speaker finishes the sentence, which trades the occasional correction or backtrack for fluency. Noise robustness is claimed but not benchmarked publicly, and the broader Meet rollout is gated to "later this year" after the private-preview phase wraps. Real-world latency at scale, across 2,000+ language pairs and unpredictable network conditions, is what the Grab pilot and the Meet preview will test.

Live, low-latency speech translation has been the holy grail since the original Google Translate demo, and shipping it across Meet, Translate, and an API on the same day is how Google converts research lead into product distribution. The bigger competitive signal is what this does to standalone translation startups and to the AI-meeting-assistant category: when 70+ languages and 2,000+ pairings ship inside Meet as a default feature, the bar for a separate product gets meaningfully higher, and the value migrates to whoever owns the meeting surface.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Google logo
Models

Google ships three new Gemini models but delays Pro update

Gemini 3.6 Flash cuts token usage by up to 17%, but the flagship Pro refresh is still stuck as OpenAI and Anthropic pull ahead.

Jaeden Schafer5 min read
Google logo
Models

Google ships Nano Banana 2 Lite and Gemini Omni Flash to developers

DeepMind's fastest image model generates in 4 seconds at $0.034 per 1K images; Omni Flash matches Veo 3.1 Fast at $0.10 per second of video.

Jaeden Schafer5 min read
Google logo
Models

Google bakes computer use into Gemini 3.5 Flash as a native tool

DeepMind folds its standalone agent model into Flash, letting developers build agents that drive browsers, mobile apps and desktops via one API call.

Jaeden Schafer4 min read