wildminder/ComfyUI-VoxCPM
ComfyUI node for highly expressive speech and realistic zero-shot voice cloning
What it solves
ComfyUI-VoxCPM provides a visual interface for the VoxCPM system, enabling users to generate expressive, high-quality speech and perform advanced voice cloning without needing to write code. It simplifies the process of managing models, handling audio processing, and fine-tuning voices via LoRAs.
How it works
The project integrates the tokenizer-free VoxCPM models (based on a MiniCPM-4 backbone) into ComfyUI as a set of custom nodes. It supports two main model versions:
- VoxCPM2: Supports voice design (creating voices from text descriptions), reference cloning (using audio for identity), and "ultimate cloning" (combining identity and prosody from different sources).
- VoxCPM1.5: Focuses on zero-shot TTS and supports native LoRA training and inference.
Users can connect nodes for TTS generation, voice cloning configuration, and advanced diffusion parameters (like temperature and CFG) to control the output audio quality and expressiveness.
Who it’s for
Content creators, audio engineers, and AI researchers who use ComfyUI and want to generate professional-grade synthetic speech, clone specific voices, or design custom voices from natural language descriptions.
Highlights
- Advanced Cloning Modes: Offers zero-shot, reference-based, and "ultimate cloning" (identity + prosody).
- Voice Design: Generate new voices using natural language prompts (e.g., "warm female voice").
- Multilingual Support: VoxCPM2 supports over 30 languages with 48kHz output.
- LoRA Integration: Includes built-in nodes for both training custom LoRA models and performing inference.
- Automated Management: Handles model downloading and memory management end-to-end within the ComfyUI environment.
Related
- Project
- Project
- Project
- Project
- Project