FireRedTeam/FireRedTTS3

FireRedTTS3: Multilingual and Multi-Dialect Voice Cloning with Instruction-Guided Voice Design and Speech Editing

What it solves

FireRedTTS3 is a unified speech generation and editing system designed to handle high-fidelity voice cloning across a vast array of languages and dialects, as well as complex speech modifications through natural language instructions.

How it works

The system is built on semantically enriched continuous speech representations and is available in two specialized variants:

  • FireRedTTS3-Base: Focuses on zero-shot voice cloning, allowing the model to mimic a speaker's voice using only a short reference audio clip.
  • FireRedTTS3-Instruct: An instruction-driven model that enables "voice design" (creating new voices from text descriptions) and "speech editing" (modifying existing audio semantically or acoustically).

To ensure high quality, it utilizes a text-normalization frontend that can be powered by local tools or LLMs to convert written text into spoken forms across multiple languages.

Who it’s for

This project is for developers and researchers working in text-to-speech (TTS), voice cloning, and audio editing who need a multilingual system capable of handling 24 languages and 21 Chinese dialects.

Highlights

  • Extensive Language Support: Zero-shot cloning across 24 languages and 21 Chinese dialects.
  • Instruction-Controlled Voice Design: Ability to generate entirely new voices based on descriptions of age, gender, timbre, and emotion without reference audio.
  • Free-Form Speech Editing: Supports semantic edits (inserting, deleting, or substituting words) and acoustic edits (adjusting speed, pitch, and volume) via text instructions.
  • High Performance: Demonstrates state-of-the-art results in Word Error Rate (WER) and speaker similarity on benchmarks like Seed-TTS-eval and MiniMax-MLS-Test.

Related

  • Project
  • Project
  • Project
  • Project