rsxdalv/TTS-WebUI

A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark!

What it solves

TTS WebUI provides a unified, user-friendly interface for interacting with a vast array of text-to-speech (TTS), audio generation, and audio conversion tools. It eliminates the need to install and manage multiple separate AI audio projects, which often have conflicting dependencies and complex setup processes.

How it works

The project acts as a wrapper and manager for numerous open-source AI audio models. It offers two interface options: a Gradio-based backend and a modern React-based frontend. The system supports an extensible architecture where additional models and tools can be installed as extensions (Python packages) via an internal marketplace or external JSON configurations.

Who it’s for

It is designed for creators, developers, and AI enthusiasts who want to experiment with different high-quality voice synthesis and audio generation models without dealing with the manual installation of each individual project.

Highlights

  • Extensive Model Support: Integrates a wide range of models including Bark, Tortoise, StyleTTS2, MusicGen, and RVC.
  • Multimodal Audio Tools: Includes tools for audio conversion, music generation, and audio separation (e.g., Demucs, Whisper).
  • Flexible Deployment: Supports installation via a dedicated installer (Ignition), Docker containers, or manual setup.
  • API Integrations: Provides OpenAI-compatible APIs to integrate with third-party applications like Silly Tavern and OpenWebUI.
  • Extension System: Features a built-in marketplace for adding new community-created audio capabilities.

Related

  • Project
  • Project
  • Project
  • Project
  • Project