moonshine-ai/moonshine

Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces

What it solves

Moonshine Voice provides a toolkit for developers to build real-time voice agents and applications that run entirely on-device. It eliminates the need for cloud APIs, accounts, or API keys, ensuring privacy and low latency for live streaming voice interactions.

How it works

The toolkit uses speech-to-text models trained from scratch, ranging from high-accuracy models that outperform Whisper Large V3 to ultra-tiny 1MB models. It is designed for live streaming by processing audio while the user is still speaking to minimize delay.

Who it’s for

Developers building voice-enabled applications across a wide range of platforms, including Python, JavaScript/WASM, iOS, Android, macOS, Linux, Windows, and Raspberry Pi.

Highlights

  • On-device execution for speed and privacy.
  • Low-latency streaming optimization.
  • Scalable model sizes, from high-performance to 1MB tiny models.
  • Broad cross-platform support across mobile, desktop, and embedded systems.

Related

  • Project
  • Project
  • Project
  • Project
  • Project