debpalash/VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

What it solves

VoiceStudio is a local-first audio workstation that eliminates the need for cloud-based subscriptions and API keys for high-quality voice cloning, dubbing, and transcription. It provides a unified interface for multiple AI audio engines, allowing users to generate speech and transcribe audio while keeping data on their own hardware.

How it works

The project uses a Tauri-based desktop shell with a React frontend and a FastAPI backend. It acts as a registry for various Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) engines (such as OmniVoice, WhisperX, and CosyVoice), routing requests to the local GPU (CUDA, MPS, ROCm) or CPU. It also provides an OpenAI-compatible API and an MCP server for integration with other tools and agents.

Who it’s for

It is designed for creators, developers, and privacy-conscious users who need professional audio tools for voice cloning, video dubbing, audiobook production, and system-wide dictation without relying on hosted services.

Highlights

  • Local-first workflow: Core audio generation and transcription stay on the machine by default.
  • Extensive engine support: Includes 16 TTS and 11 ASR engines covering over 600 languages.
  • Advanced audio features: Zero-shot voice cloning, voice design based on attributes, and video dubbing with speaker preservation.
  • Productivity tools: System-wide dictation widget with optional local-LLM cleanup and vocal isolation using Demucs.
  • Developer-friendly: Offers an OpenAI-compatible audio API and support for MCP clients.

Related

  • Project
  • Project
  • Project
  • Project
  • Project