diodiogod/TTS-Audio-Suite

A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools

What it solves

TTS Audio Suite is a comprehensive ComfyUI extension that unifies a vast array of Text-to-Speech (TTS), Voice Conversion (VC), and Automatic Speech Recognition (ASR) engines into a single modular interface. It solves the fragmentation of audio AI tools by providing a centralized hub for high-quality voice synthesis, voice cloning, and subtitle-synchronized audio generation, while managing the complex dependency requirements of different AI models through runtime isolation.

How it works

The suite integrates 19 different engines (including F5-TTS, Higgs Audio, MOSS-TTS, and RVC) into ComfyUI nodes. It uses a modular architecture that allows modern engines to run in a primary Transformers 5 environment while isolating fragile legacy stacks in dedicated runtimes to prevent dependency conflicts. For subtitle work, it processes SRT files to generate synchronized audio, using "smart_natural" timing logic to prevent overlaps and ensure natural speech flow.

Who it’s for

It is designed for ComfyUI users, content creators, and animators who need professional-grade voiceovers, precise subtitle timing, voice cloning, or automated transcription and audio editing within a node-based workflow.

Highlights

  • Massive Engine Support: Integrates 19 engines covering TTS, Voice Conversion, and ASR across hundreds of languages.
  • SRT Integration: Specialized nodes for generating TTS from SRT files with precise timing and segment-level caching.
  • Runtime Isolation: Prevents dependency deadlocks by separating modern and legacy model environments.
  • Advanced Audio Tools: Includes a Visual Tag Builder for prompt-based attributes, a Silent Speech Analyzer for mouth-movement-based timing, and integrated RVC model training.
  • Voice Design: Tools to create reusable voices using compatible engines like Qwen3-TTS and OmniVoice.

Related

  • Project
  • Project
  • Project
  • Project
  • Project