Qwen3-Omni-Flash-2025-12-01 Release Notes
Qwen3-Omni-Flash-2025-12-01 is a comprehensively upgraded iteration of the Qwen3-Omni native multimodal large model. It enables the seamless processing of text, images, audio, and video inputs and generates simultaneous text and natural-sounding speech outputs via real-time streaming responses.
Enhanced Audio-Visual Interaction and Control
Qwen3-Omni-Flash-2025-12-01 improves the stability and coherence of multi-turn audio-visual conversations, resolving the "intelligence drop" often encountered in casual spoken scenarios. The model now provides a more natural interaction experience through improved understanding and execution of audio-visual instructions.
System prompt control has been strengthened to allow full customization of model behavior. Users can now finely tune persona styles (such as anime-inspired, cool, or sweet), colloquial tone preferences, and output length constraints.
Multilingual Capabilities and Speech Synthesis
The model ensures more reliable multilingual compliance, addressing language-following instability from previous versions. It supports:
- Text-based interaction: 119 languages
- Speech recognition: 19 languages
- Speech synthesis: 10 languages
Speech synthesis has been upgraded to eliminate robotic or sluggish delivery. By enhancing adaptive control over prosody, the model intelligently adjusts intonation, pauses, and speaking rate based on the textual context to mimic human speech.
Performance Benchmarks
Qwen3-Omni-Flash-2025-12-01 shows substantial improvements over Qwen3-Omni-Flash across all modalities:
Text Understanding and Generation
- Logical Reasoning: +5.6 on ZebraLogic
- Code Generation: +9.3 on LiveCodeBench-v6 and +2.7 on MultiPL-E
- Writing Quality: +2.2 on WritingBench
Speech Understanding and Synthesis
- Speech Understanding: Improved VoiceBench (+3.2) and a significantly lower word error rate on Fleurs-zh.
- Speech Synthesis: Higher quality, human-like voice generation with improved pacing and prosody, particularly in Chinese and multilingual contexts.
Visual and Video Understanding
- Image Understanding: Gains in visual reasoning tasks, including +4.7 on MMMU, +4.8 on MMMU-Pro, and +2.2 on MathVision_full.
- Video Understanding: Improved video semantic comprehension (+1.6 on MLVU) and tighter audio-visual synchronization for real-time conversations.
Future Development Roadmap
The Qwen team plans to advance the model along several axes, including:
- Multi-speaker ASR (Automatic Speech Recognition)
- Video OCR
- Audio-video proactive learning
- Enhanced support for agent-based workflows and function calling