lenML/Speech-AI-Forge

🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.

What it solves

Speech-AI-Forge provides a unified interface and deployment framework for a wide variety of state-of-the-art Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models. It eliminates the need to install and manage multiple separate repositories for different speech AI models, offering a centralized hub for voice generation, cloning, and transcription.

How it works

The project implements an API server and a Gradio-based WebUI that wraps multiple speech models. It supports a broad range of models including ChatTTS, CosyVoice, FishSpeech, GPT-SoVITS, F5-TTS, and Qwen3-TTS, as well as ASR models like Whisper and SenseVoice. Users can deploy it via Windows standalone packages, Google Colab, Docker, or local installation.

Who it’s for

It is designed for developers and creators who need high-quality AI voice synthesis and transcription services without the complexity of managing multiple individual model environments.

Highlights

  • Multi-model Support: Integrates numerous TTS engines (e.g., ChatTTS, CosyVoice, F5-TTS) and ASR engines (Whisper, SenseVoice).
  • Advanced Voice Control: Features speaker switching, custom voice uploads, reference-based cloning, and style control.
  • SSML Support: Includes tools for fine-grained control over long-text synthesis, podcast-style multi-role audio, and subtitle-to-speech generation.
  • Voice Management: Tools for creating, blending, and testing custom voices, including a dedicated hub for voice resources.
  • Comprehensive Tooling: Includes a voice enhancer (ResembleEnhance) and post-processing tools for audio clipping and adjustment.
  • Flexible Deployment: Offers both a full WebUI for interactive use and a standalone API server for high-throughput applications.

Related

  • Project
  • Project
  • Project
  • Project
  • Project