netease-youdao/EmotiVoice

EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine

What it solves

EmotiVoice addresses the lack of emotional expression and voice variety in standard text-to-speech (TTS) systems. It allows users to generate natural-sounding speech in English and Chinese that can be specifically controlled for emotion (such as happiness, sadness, or anger) and speaker identity.

How it works

The engine uses a prompt-controlled architecture based on PromptTTS, allowing users to specify a speaker and an emotion/style prompt to guide the synthesis. It supports over 2,000 different voices and can be deployed via a web interface, a scripting interface for batch processing, or an OpenAI-compatible HTTP API. It also supports voice cloning using personal data.

Who it’s for

This tool is designed for developers and creators who need high-quality, emotionally expressive synthetic speech for applications, as well as users who want to customize voice identities or clone their own voice.

Highlights

  • Emotional Synthesis: Ability to generate speech with specific emotions like happy, excited, sad, and angry.
  • Massive Voice Library: Access to over 2,000 distinct voices.
  • Bilingual Support: Full support for both English and Chinese.
  • Flexible Deployment: Available as a Docker image, a standalone Mac app, or an OpenAI-compatible API.
  • Voice Cloning: Capability to train the model on personal data to clone specific voices.

Related

  • Project
  • Project
  • Project
  • Project