nu-dialogue/j-moshi

J-Moshi: A Japanese Full-duplex Spoken Dialogue System

What it solves

J-Moshi is a full-duplex spoken dialogue system designed for the Japanese language. It addresses the challenge of creating natural, real-time voice interactions that mimic human-to-human conversation, specifically enabling features like overlapping speech and backchanneling (aizuchi) in real-time.

How it works

The system is based on the 7B parameter Moshi model from Kyutai Labs. It was developed through additional training on large-scale Japanese spoken dialogue data, including corpora such as J-CHAT, Japanese Callhome, and internal chat and consultation dialogue corpora. A specialized version, J-Moshi-ext, further incorporates synthetic data generated via Multi-stream TTS to enhance performance.

Who it’s for

Researchers and developers working on Japanese voice AI, spoken dialogue systems, and natural human-computer interaction (HCI).

Highlights

  • Full-duplex communication: Enables real-time, bidirectional voice interaction with natural turn-taking.
  • Curation of Japanese data: Trained on a diverse set of spoken and text-based dialogue corpora.
  • Two model variants: Offers a standard version and an extended version (J-Moshi-ext) using synthetic data.
  • HuggingFace integration: Models are available for easy deployment via the PyTorch implementation of Moshi.

Related

  • Project
  • Project
  • Project
  • Project