Gemini 3.1 Flash TTS release notes / what's new

Gemini 3.1 Flash TTS release notes / what's new

Google DeepMind has introduced Gemini 3.1 Flash TTS, a next-generation text-to-speech (TTS) model designed for higher expressivity, improved naturalness, and granular controllability. The model is now available in preview for developers via the Gemini API and Google AI Studio, for enterprises through Vertex AI, and for Workspace users via Google Vids.

High-Fidelity Speech Quality and Performance

Gemini 3.1 Flash TTS is positioned as Google's most natural and expressive TTS model to date. According to the Artificial Analysis TTS leaderboard, which utilizes blind human preference testing, the model achieved an Elo score of 1,211.

Artificial Analysis further categorizes Gemini 3.1 Flash TTS within its "most attractive quadrant," citing an optimal balance between high-quality speech generation and low operational costs. Key capabilities include native multi-speaker dialogue and support for over 70 languages.

Granular Control via Natural Language Audio Tags

The defining feature of Gemini 3.1 Flash TTS is the introduction of audio tags. These allow users to embed natural language commands directly into text inputs to steer the vocal style, pace, and delivery of the AI speech.

To provide developers with a "director's chair" experience, Google AI Studio includes several configurable controls:

  • Scene Direction: Users can define the environment and provide dialogue instructions to ensure characters remain consistent and react naturally across multiple turns of conversation.
  • Speaker-Level Specificity: Developers can assign unique Audio Profiles to characters and use "Director's Notes" to adjust tone, accent, and pace. Inline tags enable speakers to change their expression mid-sentence.
  • Seamless Export: Once a vocal performance is finalized in the studio, the parameters can be exported as Gemini API code to maintain voice consistency across different platforms and projects.

Global Scalability and Language Support

Gemini 3.1 Flash TTS is optimized for global deployment, supporting more than 70 languages. This allows developers to implement advanced style, accent, and pacing controls in major international markets, enabling the creation of localized and expressive audio experiences at scale.

AI Safety and SynthID Watermarking

To combat misinformation and ensure the reliable detection of AI-generated content, all audio produced by Gemini 3.1 Flash TTS is watermarked using SynthID. This watermark is imperceptible to the human ear but is interwoven directly into the audio output for verification purposes.

Sources