Gemini 3.8 Flash TTS and Flash-Lite TTS launch

TL;DR

Google DeepMind released Gemini 3.8 Flash TTS and Gemini 3.8 Flash‑Lite TTS, two new text‑to‑speech models that enable creators to generate custom, expressive voices and direct line‑by‑line performance, while offering high‑volume, cost‑efficient scaling for enterprise use.

Gemini 3.8 Flash TTS: Creative Voice Studio

  • Purpose: Designed for deep creative direction and character design.
  • Capabilities: Generate entirely new voices from natural‑language prompts, control acting cues, pacing, dialect shifts, and back‑channeling.
  • Use Cases: Gaming, immersive audiobooks, podcasts, interactive media, and any scenario requiring bespoke character voices.
  • Voice Library: Access to 2,000+ production‑ready voices covering more than 100 languages and dialects, including regional variants such as Mexican Spanish, Quebec French, and Scots English.
  • Voice Replication: Create a consistent vocal profile from a 30‑second sample, with built‑in consent verification, SynthID watermarking, and C2PA credentials to protect voice talent.
  • Future Feature – Voice Remixing: Planned ability to fine‑tune timbre, pitch, pace, and accent of existing library voices via prompts.

Gemini 3.8 Flash‑Lite TTS: Scalable High‑Volume Generation

  • Purpose: Optimized for high‑volume dubbing, audio content creation, and expressive voice agents.
  • Efficiency: Provides fine‑grained control over tone, pacing, and nuance while minimizing cost and latency.
  • Target Scenarios: Large‑scale media localization, voice‑assistant deployments, and any application needing rapid, reliable speech synthesis.

Precise Line‑by‑Line Performance Control

  • Stage Directions: Users can write explicit script cues or let Gemini interpret natural language directions for each line.
  • Long‑Form Generation: Maintains high voice quality and minimal speaker drift across hours of continuous audio, ideal for podcasts and audiobooks.
  • Two‑Speaker Scene Staging: Supports native multi‑turn conversations with distinct voice separation for dialogue‑driven content.
  • Vocal Bursts & Back‑channeling: Allows insertion of non‑verbal cues (e.g., <laughs>, <sigh>) and active‑listening interjections (|mhm|, |yeah|) for realistic conversational texture.

Benchmark Performance and Global Quality

  • Hume AI Voice Design Benchmark: Gemini 3.8 Flash TTS ranks #1 overall (score 71.4) and leads in accent modeling (score 60.8).
  • Overall Quality Index: Flash TTS holds the #1 spot, Flash‑Lite TTS the #2 spot.
  • Human Preference Evaluations (Voice Arena): Both models achieve top positions across key languages—including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.
  • Multilingual Support: Over 100 languages enable worldwide deployment of high‑quality voice experiences.

Trust, Consent, and Transparency Safeguards

  • Consent Verification: Voice replication requires a verbal consent recording from the voice owner that matches the reference speaker.
  • Synthetic Watermarking: Every audio output is embedded with an imperceptible SynthID watermark, ensuring AI‑generated speech is detectable and helping prevent misinformation.
  • Model Card: Detailed safety and responsibility information is available in the Gemini 3.8 Audio model card.

Access and Integration Options

  • Google AI Studio Playground: Immediate hands‑on experience via the speech‑generation workspace, supporting voice design, replication, and dual‑speaker screenplay editing.
  • Gemini API: Enables integration with platforms such as Agora, LiveKit, Pipecat, and Vercel for building high‑performance speech interfaces.
  • Enterprise Rollout: Flash TTS will be available through Gemini Enterprise API; Flash‑Lite TTS will be accessible via Google Vids for end‑users.
  • Partner Deployments: Early adopters include Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang, leveraging the models for dubbing, regional accent localization, and conversational agents.

Getting Started

  • Developers: Access both models via the Gemini API and Google AI Studio.
  • Enterprises: Upcoming availability through Gemini Enterprise API.
  • General Users: Flash TTS appears in Gemini Notebook; Flash‑Lite TTS will be available in Google Vids.

All information is drawn directly from the DeepMind blog post dated September 23 2026.

Sources