Gemini 2.5 Flash Native Audio release: enhanced live voice agents and real‑time speech translation

TL;DR

Google DeepMind released Gemini 2.5 Flash Native Audio, an upgraded live‑voice model that improves function calling, instruction following, and multi‑turn conversation quality, and adds real‑time speech‑to‑speech translation for over 70 languages.


Overview of Gemini 2.5 Flash Native Audio

Gemini 2.5 Flash Native Audio is the latest iteration of DeepMind’s audio‑centric Gemini models. It is designed for live voice agents and is now generally available through Google AI Studio, Vertex AI, Gemini Live, and Search Live. The release expands the model’s capabilities beyond expressive text‑to‑speech to include robust conversational handling and live speech translation.


Technical Improvements

Sharper Function Calling

  • The model now identifies when to invoke external functions with higher reliability.
  • In the ComplexFuncBench Audio evaluation, Gemini 2.5 Flash Native Audio achieved a score of 71.5 %, leading the benchmark.

Robust Instruction Following

  • Adherence to developer‑provided instructions rose to 90 %, up from 84 % in the prior version.
  • This translates to more complete and predictable audio outputs for complex prompts.

Smoother Multi‑Turn Conversations

  • Enhanced context retrieval across dialogue turns yields more cohesive and natural interactions.
  • Users experience fewer breaks in flow when the model integrates real‑time information into its responses.

Live Voice Agent Use Cases

The model powers a spectrum of conversational experiences, from brainstorming sessions in Gemini Live to real‑time assistance in Search Live. Enterprises can embed the model in customer‑service pipelines, enabling agents that sound natural, handle noisy environments, and switch languages on the fly.

Customer Testimonials

“Users often forget they’re talking to AI within a minute of using Sidekick…New Live API AI capabilities offered through Gemini 2.5 Flash Native Audio empower our merchants to win.” – David Wurtz, VP of Product, Shopify

“By integrating the Gemini 2.5 Flash Native Audio model…we've significantly enhanced Mia's capabilities…generated over 14,000 loans for our broker partners.” – Jason Bressler, CTO, United Wholesale Mortgage

“Working with the Gemini 2.5 Flash Native Audio model through Vertex AI allows Newo.ai AI Receptionists to achieve unmatched conversational intelligence…identify the main speaker even in noisy settings, switch languages mid‑conversation, and sound remarkably natural and emotionally expressive.” – David Yang, Co‑founder, Newo.ai


Live Speech Translation

Gemini now offers native, real‑time speech‑to‑speech translation that works in two modes:

  • Continuous Listening – Translates ambient speech from any language into a single target language, enabling users to wear headphones and hear the world in their preferred language.
  • Two‑Way Conversation – Dynamically switches output language based on the speaker, allowing seamless bilingual dialogues (e.g., English ↔ Hindi).

Core Translation Features

Feature Description
Language coverage Supports >70 languages and 2,000 language pairs, leveraging Gemini’s multilingual knowledge.
Style transfer Preserves intonation, pacing, and pitch so translated speech sounds natural.
Multilingual input Handles multiple source languages within a single session.
Auto detection Detects spoken language automatically, requiring no manual selection.
Noise robustness Filters ambient noise for clear translation in loud environments.

The beta experience is available in the Google Translate app (Android rollout in the US, Mexico, and India; iOS and additional regions forthcoming). Feedback will guide further integration, including a planned Gemini API release in 2026.


Getting Started

Developers can start building voice agents with Gemini 2.5 Flash Native Audio via:

  • Vertex AI – General availability for production workloads.
  • Gemini API preview – Accessible through the Gemini API documentation.
  • Google AI Studio – Interactive playground for rapid prototyping.

The Gemini 2.5 Flash and Gemini 2.5 Pro text‑to‑speech models are also available through the Gemini API, with comprehensive documentation, prompting guides, and a cookbook of example notebooks.


Implications

The release marks a shift from static text‑to‑speech toward fully interactive, audio‑first AI experiences. By combining reliable function calling, high instruction fidelity, and real‑time multilingual translation, Gemini 2.5 Flash Native Audio enables more natural, enterprise‑grade voice agents and opens new avenues for global communication, especially in scenarios where visual interfaces are impractical.


References

Sources