abus-aikorea/voice-pro

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

What it solves

Voice-Pro is an all-in-one multimedia processing tool designed to simplify the complex workflow of creating multilingual content. It eliminates the need for multiple separate tools by combining YouTube downloading, voice separation, speech recognition, translation, and text-to-speech (TTS) into a single interface.

How it works

The application operates as a web-based UI (built with Gradio) that integrates several state-of-the-art AI models:

  • Speech-to-Text (ASR): Uses Whisper, Faster-Whisper, and Whisper-Timestamped for high-accuracy transcription.
  • Voice Cloning & TTS: Employs F5-TTS, E2-TTS, CosyVoice, Edge-TTS, and kokoro for zero-shot voice cloning and multilingual speech generation.
  • Audio Processing: Integrates Demucs for voice separation and yt-dlp for YouTube content extraction.
  • Translation: Uses Deep-Translator and optionally Azure Translator for supporting over 100 languages.

Who it’s for

It is primarily aimed at podcasters, content creators, researchers, and multilingual professionals who need to dub videos, generate subtitles, or clone voices for high-quality audio production.

Highlights

  • Comprehensive Dubbing Studio: A centralized hub for downloading, denoising, translating, and generating TTS for videos.
  • Zero-Shot Voice Cloning: Ability to clone voices using models like F5-TTS and CosyVoice without extensive training.
  • Multilingual Support: Transcription and translation capabilities for over 100 languages.
  • Integrated Subtitle Tools: Dedicated tab for generating and displaying subtitles with word-level highlighting.
  • Simplified Installation: Uses the uv installer for fast, reproducible setups on Windows with NVIDIA GPUs.

Related

  • Project
  • Project
  • Project
  • Project