Google DeepMind announces Gemini 3.8 Live with Live Avatar

TL;DR

Google DeepMind launched Gemini 3.8 Live with Live Avatar, a feature that couples Gemini 3.8 Live’s real‑time dialogue capabilities with low‑latency, lip‑synced video avatars, enabling enterprises to deliver multilingual, visually expressive conversational agents.


Real‑time visual presence for conversational AI

Gemini 3.8 Live with Live Avatar adds near‑real‑time video generation to the existing live dialogue model. The system processes audio and visual inputs simultaneously and produces synchronized speech, facial expressions, and precise lip‑sync. This creates a dynamic visual persona that can listen, see, and speak, making interactions feel more natural for end users.

"By pairing near real‑time video generation with speech, the Live Avatar feature creates an experience that listens, sees, and speaks with a dynamic visual persona." – DeepMind blog

Multimodal reasoning with asynchronous tool execution

The avatar is backed by Gemini’s advanced reasoning engine. It can invoke tools asynchronously—fetching data or performing actions in the background—while maintaining uninterrupted dialogue. This enables complex workflows such as hotel check‑ins without breaking conversational flow.

"Live Avatar can trigger tool calls and fetch data in the background while continuing active dialogue, handling complex tasks while ensuring an uninterrupted conversational flow." – DeepMind blog

Native multilingual lip‑sync across 97 languages

Live Avatar supports seamless language switching. The model dynamically adapts lip‑sync and facial expressions when transitioning among 97 supported languages, preserving video fidelity and avoiding visual drift.

"The feature dynamically adapts its lip‑sync and expressions and can seamlessly transition across 97 languages without degrading video fidelity or introducing visual drift." – DeepMind blog

Customizable avatars for brand alignment

Enterprises can choose from a library of preset avatars or generate a custom avatar from a high‑quality reference image. Custom avatars retain the reference likeness and brand styling, but are currently limited to enterprise allow‑listed customers.

Safety, transparency, and watermarking

All audio‑video output is embedded with DeepMind’s SynthID watermark, an imperceptible signal woven into the media to make AI‑generated content detectable. The model card details additional safeguards and responsible deployment practices.

"All output generated by our AI products is watermarked with SynthID. This imperceptible watermark is woven directly into the audio and video output, helping to ensure AI‑generated content remains detectable." – DeepMind blog

Availability and getting started

Gemini 3.8 Live with Live Avatar is generally available through Gemini Enterprise. Developers can access the feature via the Gemini Enterprise Agent Platform, with API documentation and a console UI for rapid integration.


Implications for enterprise AI

The introduction of real‑time visual avatars expands the scope of conversational agents beyond text and voice, enabling richer customer‑facing experiences such as interactive walkthroughs, virtual assistants, and multilingual support. By integrating asynchronous tool calls, the system can handle sophisticated tasks without sacrificing conversational continuity, a key requirement for high‑volume enterprise deployments.


Limitations and next steps

Custom avatar creation is currently restricted to enterprise allow‑listed users, and the feature’s performance at extreme network latencies has not been disclosed. Future updates are expected to broaden avatar customization, improve latency, and extend language coverage.


Conclusion

Gemini 3.8 Live with Live Avatar represents a significant step toward fully multimodal, brandable AI agents that can see, speak, and be seen in real time, positioning Google DeepMind’s Gemini platform as a leading solution for enterprise‑grade conversational experiences.

Sources