OpenMOSS/MOSS-TTSD

MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references.

What it solves

MOSS-TTSD addresses the limitation of traditional text-to-speech (TTS) systems that are optimized for single-speaker monologues. It transforms static dialogue scripts into cohesive, continuous multi-party conversations, maintaining speaker identity and emotional nuance across long-form audio generation.

How it works

The model uses a continuation workflow where users provide reference audio and transcripts for each speaker (up to 5 speakers). It then generates spoken dialogue based on a provided script, ensuring natural turn-taking and consistent personas. For high-performance deployment, it can be fused with the MOSS-Audio-Tokenizer and run via the SGLang inference engine to accelerate generation speeds.

Who it’s for

It is designed for content creators and developers building long-form audio experiences, such as AI-generated podcasts, audiobooks, sports and esports commentary, dubbing, and entertainment media.

Highlights

  • Multi-party Dialogue: Supports 1 to 5 speakers with flexible control over overlapping speech and natural conversational rhythms.
  • Extreme Long-Context: Capable of generating up to 60 minutes of coherent audio in a single session while maintaining consistent speaker identity.
  • Zero-Shot Voice Cloning: High-fidelity voice cloning using only short reference audio samples.
  • Multilingual Support: Robust performance across 20 languages, including Chinese, English, Japanese, and various European languages.
  • High-Speed Inference: Integration with SGLang allows for inference acceleration of up to 16x.

Related

  • Project
  • Project
  • Project
  • Project
  • Project