diodiogod/TTS-Audio-Suite
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools
What it solves
TTS Audio Suite is a comprehensive ComfyUI extension that unifies a vast array of Text-to-Speech (TTS), Voice Conversion (VC), and Automatic Speech Recognition (ASR) engines into a single modular interface. It solves the fragmentation of audio AI tools by providing a centralized hub for high-quality voice synthesis, voice cloning, and subtitle-synchronized audio generation, while managing the complex dependency requirements of different AI models through runtime isolation.
How it works
The suite integrates 19 different engines (including F5-TTS, Higgs Audio, MOSS-TTS, and RVC) into ComfyUI nodes. It uses a modular architecture that allows modern engines to run in a primary Transformers 5 environment while isolating fragile legacy stacks in dedicated runtimes to prevent dependency conflicts. For subtitle work, it processes SRT files to generate synchronized audio, using "smart_natural" timing logic to prevent overlaps and ensure natural speech flow.
Who it’s for
It is designed for ComfyUI users, content creators, and animators who need professional-grade voiceovers, precise subtitle timing, voice cloning, or automated transcription and audio editing within a node-based workflow.
Highlights
- Massive Engine Support: Integrates 19 engines covering TTS, Voice Conversion, and ASR across hundreds of languages.
- SRT Integration: Specialized nodes for generating TTS from SRT files with precise timing and segment-level caching.
- Runtime Isolation: Prevents dependency deadlocks by separating modern and legacy model environments.
- Advanced Audio Tools: Includes a Visual Tag Builder for prompt-based attributes, a Silent Speech Analyzer for mouth-movement-based timing, and integrated RVC model training.
- Voice Design: Tools to create reusable voices using compatible engines like Qwen3-TTS and OmniVoice.
Related
- Project
- Project
- Project
- Project
- Project