Grok Voice Think Fast 2.0 Release Notes
xAI has announced Grok Voice Think Fast 2.0, a next-generation voice model designed to improve speech reasoning, transcription accuracy, and conversational fluidity. The model is built to outperform previous iterations and competing speech-to-speech models in both intelligence and real-world performance.
Intelligence and Benchmarks
Grok Voice Think Fast 2.0 demonstrates significant gains in speech reasoning and agentic performance. According to data from Artificial Analysis, the model achieves an Overall AA Speech-to-Speech Quality Index of 82.9%, surpassing GPT-Realtime-2.1 (High) at 79.1% and Gemini 3.1 Flash (High) at 69.5%.
Key performance metrics include:
- Speech Reasoning (Big Bench Audio): 97.2%
- Conversational Dynamics (Full Duplex Bench): 95.1%
- Agentic Performance (τ-voice Bench): 56.5%, the highest among the compared models, including GPT-Realtime-2.1 (45.7%) and Gemini 3.1 Flash (37.7%)
- Latency: The Time to First Audio is 0.70s, significantly faster than Grok Voice Think Fast 1.0 (1.25s) and Gemini 3.1 Flash (2.98s)
Transcription Accuracy in Real-World Settings
Grok Voice Think Fast 2.0 provides transcription accuracy that outperforms dedicated state-of-the-art transcription models. In evaluations across 24 languages and thousands of short phrases, the model showed a 1.5–2.0× improvement over Deepgram Nova 3 and ElevenLabs Scribe v2, and a 1.4× improvement over its predecessor.
Performance gains are most pronounced in challenging environments. The accuracy gap between Grok Voice Think Fast 2.0 and dedicated speech-to-text models increases to approximately 10× in noisy settings, specifically those involving background noise and telephony compression.
Parallel Reasoning Efficiency
Unlike traditional speech-to-speech models, Grok Voice Think Fast models reason through queries while speaking. This parallel reasoning allows the model to be more intelligent without increasing latency.
Grok Voice Think Fast 2.0 is significantly more efficient with its reasoning tokens. It uses 0.4× the reasoning tokens per response (P50) compared to Grok Voice Think Fast 1.0. In production, this efficiency results in faster tool calls, which typically execute before the agent finishes its first sentence.
Conversational Capabilities and RL Training
To improve conversational fluidity, xAI used extensive reinforcement learning to align the model with real human conversation patterns. This training results in a model that speaks in shorter sentences, asks only one question at a time, and avoids unnecessary fluff. This allows the model to be capable of guiding users through complex workflows while maintaining a simple and fluid user experience.
Migration, Pricing, and Deployment
On August 5, 2026, the grok-voice-latest alias will automatically transition from version 1.0 to 2.0. Users wishing to remain on version 1.0 must pin their implementation to grok-voice-think-fast-1.0 before that date.
Pricing: Grok Voice Think Fast 2.0 is priced at $0.08 per minute of audio.
Real-world Impact: A/B testing on Starlink services showed that the deployment of this model resulted in a significant increase in sales conversion rates and support containment rates.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch