HeartMuLa/heartlib

HeartMuLa Official Repo: The Most Powerful Open-Source Music Generation Model of 2026

What it solves

HeartMuLa provides a suite of open-source music foundation models designed to generate high-quality music conditioned on lyrics and tags. It addresses the need for controllable, multilingual music generation and high-fidelity audio reconstruction, offering tools for both creation and analysis of music audio.

How it works

The project consists of four primary components:

  • HeartMuLa: A music language model that generates audio based on provided lyrics and tags, supporting a wide range of languages.
  • HeartCodec: A 12.5 Hz music codec that converts audio into discrete tokens and reconstructs 48 kHz stereo audio with high fidelity.
  • HeartTranscriptor: A Whisper-based model specifically tuned for transcribing lyrics from audio.
  • HeartCLAP: An audio-text alignment model that creates a unified embedding space for music descriptions and cross-modal retrieval.

Additionally, the project includes MuLaCover, a model for controllable cover-song and music-remixing that uses symbolic melody and harmony to preserve musical identity while allowing changes to genre, mood, and instrumentation.

Who it’s for

This toolkit is for music producers, AI researchers in audio/music generation, and developers building music-related applications who require high-fidelity, controllable music synthesis and transcription.

Highlights

  • Multilingual Support: Generates music based on lyrics in almost all languages.
  • High Fidelity: HeartCodec enables reconstruction of 48 kHz stereo audio.
  • Controllability: Precise control over musical styles and tags via Reinforcement Learning (RL) refinements.
  • Comprehensive Ecosystem: Includes tools for generation, tokenization, reconstruction, and transcription.

Related

  • Project
  • Project
  • Project
  • Project