Sharrnah/whispering-ui

Native UI for the Whispering Tiger project - https://github.com/Sharrnah/whispering (live transcription / translation)

What it solves

Whispering Tiger UI provides a native Windows interface for the Whispering Tiger application, enabling users to easily transcribe and translate audio streams or in-game images in real-time. It removes the complexity of managing the backend platform, allowing users to configure audio devices, AI models, and output methods without needing todeeply technical knowledge.

How it works

The UI acts as a control panel for the Whispering Tiger backend. It allows users to select audio input/output devices, choose AI model sizes and precisions (CUDA for GPU or CPU), and set up output channels such as web browsers or VRChat via Websockets or OSC. It also supports loopback audio devices to capture PC audio directly.

Who it’s for

Gamers, streamers, and users of virtual reality platforms like VRChat who need real-time translation and transcription of audio or on-screen text (OCR) from their machine.

Highlights

  • Multi-modal AI capabilities: Supports speech-to-text, text translation, text-to-speech, and image-to-text (OCR) for in-game images.
  • Extensible via Plugins: Includes support for plugins such as Emotion Prediction, RVC (Retrieval-based Voice Conversion), and LLM integration.
  • Hardware Acceleration: Supports CUDA for NVIDIA GPUs to improve transcription speed and reduce latency.
  • Flexible Output: Sends results to web browsers or VRChat using Websockets or OSC.

Related

  • Project
  • Project
  • Project
  • Project
  • Project