Gemini 3.5 Live Translate Release Notes

Gemini 3.5 Live Translate Release Notes

Google DeepMind has introduced Gemini 3.5 Live Translate, a new audio model designed for near real-time speech-to-speech translation across more than 70 languages. This model shifts away from traditional turn-by-turn translation systems to a continuous generation approach, reducing awkward pauses and maintaining a fluid conversation flow that stays only a few seconds behind the speaker.

Continuous Speech-to-Speech Translation

Gemini 3.5 Live Translate enables fluid, natural-sounding translated speech by balancing the trade-off between waiting for sufficient context to ensure quality and translating immediately to maintain synchronization with the speaker. Key technical capabilities include:

  • Prosody Preservation: The model preserves the original speaker's intonation, pacing, and pitch during translation.
  • Automatic Language Detection: It automatically detects 70+ languages without requiring manual configuration.
  • Noise Robustness: The model is engineered to function in loud and unpredictable environments, making it suitable for real-world applications.

Integration and Availability

Gemini 3.5 Live Translate is being deployed across several platforms for different user segments:

  • Developers: Available in public preview via the Gemini Live API and Google AI Studio.
  • Enterprises: Entering private preview this month in Google Meet for select business Google Workspace customers.
  • General Users: Rolling out globally via the Google Translate app on Android and iOS.

Google Meet Enhancements

The integration of Gemini 3.5 Live Translate into Google Meet significantly expands the platform's translation capabilities. The update increases the supported language count from five to over 70, and expands translation combinations from English-only pairs to over 2,000 language combinations within a single meeting.

Mobile Experience and Listening Mode

On the Google Translate mobile app, users can utilize headphones for a seamless experience that mirrors the speaker's tone. Android users specifically receive a new "listening mode," which allows translated audio to stream directly through the phone's earpiece, enabling private translation of live speech (such as guided tours) without the need for headphones.

Developer Ecosystem and Real-World Use Cases

Developers can build voice translation apps using the Gemini Live API, supported by infrastructure partners including Agora, Fishjam, LiveKit, Pipecat, and Vision Agents.

Real-world testing is already underway with Grab, which is testing the model to facilitate near real-time communication between drivers and travelers during pickups, impacting a service that handles over 10 million voice calls per month.

Safety and Watermarking

To prevent misinformation, all audio generated by Gemini 3.5 Live Translate is watermarked using SynthID. This imperceptible watermark is woven directly into the audio output to ensure that AI-generated content remains detectable.

Sources