VRCWizard/TTS-Voice-Wizard
Speech to Text to Speech. Song now playing. Sends text as OSC messages to VRChat to display on avatar. (STTTS) (Speech to TTS) (VRC STT System) (VTuber TTS)
What it solves
TTS Voice Wizard is an accessibility tool designed to help users communicate in VRChat and other environments by converting speech to text and back to speech using various TTS and speech recognition methods. It allows users who cannot or prefer not to use their own voice to express themselves with a wide variety of customizable voices and translation capabilities.
How it works
The application integrates with various speech recognition and text-to-speech (TTS) engines. It can convert spoken words into text and then synthesize speech from that text. It also uses OSC (Open Sound Control) messages to send text to VRChat avatars (via tools like KillFrenzyAvatarText or Frosty's Billboard) or the VRChat chatbox. It supports translation into over 50 languages and integrates with external services like Spotify, Pulsoid (for heartrate), and XSOverlay (for battery life).
Who it’s for
Users of VRChat and other virtual environments who need accessibility tools for speech, those who want to translate their speech in real-time for international communication, and users who want to integrate external data (like heartrate or current song) into their virtual avatar's display.
Highlights
- Speech-to-Text and Text-to-Speech: Converts speech to text and and then back to speech using multiple methods.
- Real-time Translation: Translates speech into over 50 supported languages.
- VRChat Integration: Sends text to avatar displays or the chatbox via OSC messages.
- Extensive Voice Library: Offers 100+ different voices with customization options.
- VRChat Avatar Control: Allows controlling avatar parameters via voice commands.
- External Data Integration: Displays current Spotify songs, tracker/controller battery life, and heartrate in VRChat.
- Pro Version: Provides access to premium cloud voices (Azure, Amazon Polly, Google Cloud, IBM Watson) and high-accuracy transcription via DeepGram's Nova-2 model.