OpenMOSS/MOSS-TTSD
MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references.
What it solves
MOSS-TTSD addresses the limitation of traditional text-to-speech (TTS) systems that are optimized for single-speaker monologues. It transforms static dialogue scripts into cohesive, continuous multi-party conversations, maintaining speaker identity and emotional nuance across long-form audio generation.
How it works
The model uses a continuation workflow where users provide reference audio and transcripts for each speaker (up to 5 speakers). It then generates spoken dialogue based on a provided script, ensuring natural turn-taking and consistent personas. For high-performance deployment, it can be fused with the MOSS-Audio-Tokenizer and run via the SGLang inference engine to accelerate generation speeds.
Who it’s for
It is designed for content creators and developers building long-form audio experiences, such as AI-generated podcasts, audiobooks, sports and esports commentary, dubbing, and entertainment media.
Highlights
- Multi-party Dialogue: Supports 1 to 5 speakers with flexible control over overlapping speech and natural conversational rhythms.
- Extreme Long-Context: Capable of generating up to 60 minutes of coherent audio in a single session while maintaining consistent speaker identity.
- Zero-Shot Voice Cloning: High-fidelity voice cloning using only short reference audio samples.
- Multilingual Support: Robust performance across 20 languages, including Chinese, English, Japanese, and various European languages.
- High-Speed Inference: Integration with SGLang allows for inference acceleration of up to 16x.
Related
- Project
- Project
- Project
- Project
- Project