FlashLabs-AI-Corp/FlashLabs-Chroma

Worlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.

What it solves

Chroma 1.0 is designed to enable natural, real-time spoken dialogue between humans and AI. It eliminates the need for separate speech-to-text and text-to-speech systems by providing an end-to-end multimodal model that can understand audio inputs and generate both text and synthesized speech responses simultaneously.

How it works

Chroma is a multimodal causal language model that integrates a reasoner (based on Qwen2.5-Omni-3B), a backbone (Llama3), and a decoder (Llama3), utilizing the Mimi codec for audio processing at a 24kHz sampling rate. It processes auditory inputs directly and can use reference audio prompts to guide the style of the generated speech, enabling personalized voice cloning.

Who it’s for

Developers and researchers looking to build virtual human models or voice-driven AI agents that require low-latency, natural-sounding voice interactions and personalized voice styles.

Highlights

  • End-to-End Spoken Dialogue: Processes audio input and generates audio/text output in a single model.
  • Personalized Voice Cloning: Uses reference audio prompts to mimic specific voice styles.
  • Multimodal Generation: Simultaneously produces coherent text and speech responses.
  • Open-Source: Released under the Apache-2.0 license for community development.

Related

  • Project
  • Project
  • Project
  • Project