QwenAudio/CosyVoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
What it solves
CosyVoice addresses the challenge of generating natural, consistent, and controllable speech from text. It specifically targets the need for high-quality, zero-shot multilingual speech synthesis that can handle various dialects, emotions, and complex text formats in real-world scenarios.
How it works
Based on large language models (LLM), the system uses supervised semantic tokens to synthesize speech. It supports bi-streaming (text-in and audio-out) to achieve low latency as low as 150ms. The framework includes features like pronunciation inpainting for phonemes, text normalization for special symbols, and instruction support to control language, dialect, emotion, speed, and volume.
Who it’s for
It is designed for developers and researchers working on speech synthesis, production-level audio applications requiring low latency, and users needing highly controllable multilingual voice cloning.
Highlights
- Multilingual & Dialect Support: Covers 9 common languages and over 18 Chinese dialects.
- Zero-Shot Voice Cloning: Enables cross-lingual and multi-lingual voice cloning without prior training on specific speakers.
- High Performance: Achieves state-of-the-art content consistency, speaker similarity, and prosody naturalness.
- Low Latency: Supports streaming inference with latencies as low as 150ms.
- Controllability: Supports pronunciation inpainting and various instructions for emotion and speed.