Gemini 3.8 Live and 3.8 Live Extended Thinking Release
Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new models optimized for near real-time reasoning and intuitive voice interaction. These models are designed to enable more fluid, collaborative dialogue and the execution of complex tasks via voice, targeting both developers and enterprise users.
High-Performance Voice Reasoning and Benchmarks
Gemini 3.8 Live Extended Thinking is positioned as the high-intelligence model for complex workflows. It currently holds the top spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. In terms of agentic task completion, it achieves 68.6% on $\tau$-Voice and 35.1% on Sierra's $\tau$-Voice-banking benchmark. Additionally, it scores 97.7% on Big Bench Audio.
Gemini 3.8 Live is designed for scale and cost-efficiency, securing second place in the Speech Agent Arena. Both models are noted for their ability to balance accuracy with conversational quality, pushing the Pareto Frontier on ServiceNow's EVA-Bench for complex workflows.
Key Technical Capabilities
Real-Time Multimodal Integration
Gemini 3.8 Live processes visual inputs in near real-time, allowing the model to ground conversations in visual context. It supports the automatic detection and transition between 97 supported languages mid-conversation.
Parallel Execution and Reasoning
Gemini 3.8 Live can execute tools and API calls in the background while maintaining a continuous conversation. Gemini 3.8 Live Extended Thinking takes this further by reasoning and speaking simultaneously. It uses natural verbal cues (e.g., "Let me check that...") and live progress narration to keep users informed during multi-step background tasks without interrupting the conversational flow.
Developer and Enterprise Ecosystem
Google is providing these models via the Gemini Live API, which is integrated with developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms handle the real-time media streaming infrastructure, allowing developers to focus on the user experience.
Enterprise partners such as Salesforce, Genspark, and Lumeris have already begun utilizing the models, citing low latency and strong tool-calling capabilities.
Safety and Availability
To prevent misinformation, all audio generated by these models is watermarked using SynthID, an imperceptible watermark woven directly into the audio output.
Availability rollout:
- Gemini 3.8 Live: Available for developers via the Gemini API and Google AI Studio; in private preview for Gemini Enterprise; and available for all users in Search Live.
- Gemini 3.8 Live Extended Thinking: Available for developers via the Gemini API and Google AI Studio; in private preview for Gemini Enterprise and Google Workspace business customers; and available for Google AI Pro and Ultra subscribers in Workspace (Docs) and all Google AI subscribers in Gmail and Keep.
Community Feedback and User Insights
User reactions to the release have been mixed, highlighting both the strengths and weaknesses of the models in real-world application.
Strengths
- Language Support: Users have praised the model's ability to handle niche languages, with one user noting it is "phenomenal at speaking [Afrikaans]" and providing a valuable conversational partner for rare languages.
- User Experience: Some users reported low latency and pleasant, realistic voices that cope well with thick accents.
Weaknesses and Concerns
- Hallucinations and Reliability: Some developers reported that the model can be "stupider" in agentic scenarios, inventing additional requirements and implementation details that do not exist.
- Reasoning Depth: There is skepticism regarding the "Extended Thinking" branding, with one user suggesting it be renamed to "Slightly Extended Thinking" due to an uncomfortably high number of incorrect replies in some cases.
- Technical Issues: Reports of the model entering infinite loops of replying to itself and randomly switching languages have emerged.
- Product Integration: Users have expressed frustration over the inconsistent rollout of versions (e.g., some users still seeing 3.6 Flash) and the lack of support for SIP (Session Initiation Protocol) for phone call integrations.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch