Mistral AI Voxtral Transcribe 2 Release

Mistral AI has released Voxtral Transcribe 2, a family of speech-to-text models designed for state-of-the-art transcription quality, speaker diarization, and ultra-low latency. The release consists of two primary models: Voxtral Mini Transcribe V2 for batch processing and Voxtral Realtime for live applications.

Voxtral Realtime: Ultra-Low Latency Streaming

Voxtral Realtime is engineered for live applications requiring minimal delay, utilizing a novel streaming architecture that transcribes audio as it arrives rather than processing audio in chunks.

Key Performance Metrics

  • Latency: Delay is configurable down to sub-200ms, enabling responsive voice agents.
  • Accuracy vs. Delay: At a 480ms delay, the model maintains a word error rate (WER) within 1-2%. At a 2.4-second delay, it matches the accuracy of the Voxtral Mini Transcribe V2 batch model.
  • Model Size and Deployment: With a 4B parameter footprint, the model can be deployed on edge devices for privacy-first applications.
  • Licensing: Voxtral Realtime is released as open-weights under the Apache 2.0 license.

Language Support

Voxtral Realtime is natively multilingual, supporting 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch.

Voxtral Mini Transcribe V2: High-Efficiency Batch Transcription

Voxtral Mini Transcribe V2 is optimized for high-accuracy batch transcription and diarization. It achieves approximately 4% word error rate on the FLEURS benchmark and is priced at $0.003 per minute.

Technical Capabilities

  • Speaker Diarization: The model generates transcriptions with speaker labels and precise start/end times. In cases of overlapping speech, the model typically transcribes one speaker.
  • Context Biasing: Users can provide up to 100 words or phrases to guide the model toward correct spellings of technical terms, names, or domain-specific vocabulary. This feature is optimized for English, with experimental support for other languages.
  • Word-Level Timestamps: The model provides precise start and end timestamps for every word, facilitating subtitle generation and audio search.
  • Audio Capacity: The model can process recordings up to 3 hours in a single request.
  • Noise Robustness: The system maintains accuracy in challenging acoustic environments, including field recordings, busy call centers, and factory floors.

Benchmarks and Competitive Positioning

Voxtral Mini Transcribe V2 supports 13 languages (the same set as Voxtral Realtime). Mistral AI reports that the model outperforms GPT-4o mini Transcribe, Gemini 2.5 Flash, Assembly Universal, and Deepgram Nova on accuracy. Additionally, it processes audio approximately 3x faster than ElevenLabs’ Scribe v2 while matching quality at one-fifth of the cost.

Implementation and Tooling

Mistral Studio Audio Playground

An audio playground is available in Mistral Studio for instant testing. It supports .mp3, .wav, .m4a, .flac, and .ogg files up to 1GB each. Users can toggle diarization, adjust timestamp granularity, and add context bias terms.

Pricing and Availability

  • Voxtral Mini Transcribe V2: Available via API at $0.003 per minute.
  • Voxtral Realtime: Available via API at $0.006 per minute and as open weights on Hugging Face.
  • Compliance: Both models support GDPR and HIPAA-compliant deployments via private cloud or secure on-premise setups.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch