FunAudioLLM/CosyVoice

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

What it solves

CosyVoice は、大規模言語モデル(LLM)を活用したテキスト・トゥ・スピーチ(TTS)システムで、高品質なゼロショット多言語音声合成を実現します。特に「イン・ザ・ワイルド」シナリオにおいて、自然な音声を生成しつつ、異なる言語や方言間で話者の類似性と内容の一貫性を保つ課題に取り組みます。

How it works

システムは LLM を用いて音声を生成します。最新バージョン(Fun‑CosyVoice 3.0)はスケールアップとポストトレーニングに注力し、プロソディの自然さと話者類似度を向上させています。ゼロショット音声クローンをサポートしており、短いサンプルだけで話者の声を模倣でき、広範な再学習は不要です。また、従来のフロントエンドモジュールを必要とせず、数字や記号を処理するテキスト正規化プロセスも備えています。

Who it’s for

本プロジェクトは、スケーラブルで高性能な TTS システムを必要とする開発者や研究者、特に多言語・跨言語合成が求められる場面や、低遅延ストリーミングを必要とする本番環境のオーディオアプリケーションを構築する方々を対象としています。

Highlights

  • Multilingual & Dialect Support: Covers 9 common languages and over 18 Chinese dialects/accents.
  • Zero-Shot Voice Cloning: Supports multi-lingual and cross-lingual voice cloning.
  • Low Latency: Achieves audio-out streaming latency as low as 150ms.
  • Controllability: Supports pronunciation inpainting for Pinyin and English phonemes, and instructions for emotion, speed, and volume.
  • Deployment Options: Compatible with vLLM and Nvidia TensorRT-LLM for accelerated inference.