Gemini 3.1 Flash Live release notes / what's new

Google DeepMind has announced the release of Gemini 3.1 Flash Live, its highest-quality audio and voice model to date. This model is designed to enable more fluid, natural, and precise real-time dialogue by reducing latency and improving the model's ability to understand acoustic nuances.

Technical Performance and Benchmarks

Gemini 3.1 Flash Live demonstrates significant improvements in reasoning and task execution for voice-first agents. The model leads in several key audio benchmarks:

  • ComplexFuncBench Audio: The model achieved a score of 90.8%, leading in multi-step function calling with various constraints.
  • Audio MultiChallenge (Scale AI): With "thinking" enabled, the model achieved a score of 36.1%, leading in tests of complex instruction following and long-horizon reasoning during real-world audio interruptions and hesitations.

Enhanced Audio Capabilities

The model features improved tonal understanding and acoustic sensitivity, allowing for more natural interactions:

  • Tonal Recognition: Gemini 3.1 Flash Live is more effective at recognizing pitch and pace than the previous 2.5 Flash Native Audio model.
  • Dynamic Response: The model can dynamically adjust its responses based on user expressions of frustration or confusion.
  • Environmental Robustness: The model is designed to handle complex tasks even in noisy environments.

Product Integration and Availability

Gemini 3.1 Flash Live is integrated across several Google platforms for different user segments:

  • Developers: Available in preview via the Gemini Live API in Google AI Studio.
  • Enterprises: Available through Gemini Enterprise for Customer Experience.
  • General Users: Integrated into Search Live and Gemini Live.

User Experience Improvements

For end-users, the 3.1 Flash Live model provides a more intuitive interaction experience:

  • Faster Responses: Gemini Live now delivers faster responses compared to previous iterations.
  • Extended Context: The model can follow the thread of a conversation for twice as long, improving the continuity of longer brainstorming sessions.
  • Global Expansion: Due to the model's inherent multilinguality, Search Live has expanded to more than 200 countries and territories, enabling real-time multimodal conversations in preferred languages.

Safety and Responsibility

To prevent the spread of misinformation, all audio generated by Gemini 3.1 Flash Live is watermarked using SynthID. This imperceptible watermark is interwoven directly into the audio output to allow for the reliable detection of AI-generated content.

Sources