Gemini 3.8 Live and 3.8 Live Extended Thinking Release

Google DeepMind has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new models designed to enable more intuitive, intelligent, and near real-time voice interactions. These models provide the foundational infrastructure for production-ready voice agents, integrating conversational intelligence with parallel reasoning and visual grounding.

Model Variants and Core Capabilities

Google has introduced two distinct versions of the Live audio models to balance cost, scale, and intelligence:

  • Gemini 3.8 Live: Optimized for scale and cost-efficiency. It combines fluid dialogue with visual grounding and the ability to process visual inputs in near real-time.
  • Gemini 3.8 Live Extended Thinking: Designed for high-complexity tasks. This model features increased intelligence and multi-step reasoning, allowing it to reason and speak simultaneously.

Multimodal and Linguistic Flexibility

Gemini 3.8 Live supports 97 languages, with the ability to automatically detect and transition between these languages mid-conversation. To maintain conversational flow, the model can execute tool and API calls in the background while continuing to interact with the user.

Advanced Reasoning in Extended Thinking

Gemini 3.8 Live Extended Thinking is built for complex workflows. It utilizes early verbal cues (e.g., "Let me check that...") and live progress narration to keep users informed during multi-step background tasks without interrupting the conversational flow.

Performance Benchmarks

Gemini 3.8 Live Extended Thinking leads in several key speech-to-speech and agentic benchmarks:

  • Speech to Speech Quality Index: Ranked #1 overall by Artificial Analysis with a score of 82.6.
  • Agentic Task Completion: Achieved 68.6% on $\tau$-Voice and 35.1% on Sierra's $\tau$-Voice-banking benchmark.
  • Reasoning: Scored 97.7% on Big Bench Audio.

Gemini 3.8 Live is positioned as a cost-effective alternative for scale, securing second place in the Speech Agent Arena.

On ServiceNow's EVA-Bench, both models are reported to push the Pareto Frontier for complex workflows by balancing accuracy with conversational quality.

Developer and Enterprise Integration

The models are accessible via the Gemini Live API, which is integrated into several developer platforms to simplify the deployment of voice-driven interfaces. These include:

  • Infrastructure Platforms: Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents.
  • Enterprise Partners: Salesforce, Genspark, and Lumeris.

Safety and Transparency

To prevent misinformation, all audio generated by these models is watermarked using SynthID. This imperceptible watermark is woven directly into the audio output to ensure AI-generated content remains detectable.

Availability

Both Gemini 3.8 Live and 3.8 Live Extended Thinking are rolling out starting September 15, 2026, across the following channels:

  • Developers: Available via the Gemini API and Google AI Studio.
  • Enterprises: Available in private preview in Gemini Enterprise and coming soon to Gemini Enterprise for Customer Experience (and Google Workspace business customers for the Extended Thinking model).
  • General Users: Gemini 3.8 Live is available in Search Live. Gemini 3.8 Live Extended Thinking is available in Gemini Live, for Google AI Pro and Ultra subscribers in Workspace Docs, and for all Google AI subscribers in Gmail and Keep.

Sources