moonshine-ai/moonshine
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
What it solves
Moonshine Voice provides a toolkit for developers to build real-time voice agents and applications that run entirely on-device. It eliminates the need for cloud APIs, accounts, or API keys, ensuring privacy and low latency for live streaming voice interactions.
How it works
The toolkit uses speech-to-text models trained from scratch, ranging from high-accuracy models that outperform Whisper Large V3 to ultra-tiny 1MB models. It is designed for live streaming by processing audio while the user is still speaking to minimize delay.
Who it’s for
Developers building voice-enabled applications across a wide range of platforms, including Python, JavaScript/WASM, iOS, Android, macOS, Linux, Windows, and Raspberry Pi.
Highlights
- On-device execution for speed and privacy.
- Low-latency streaming optimization.
- Scalable model sizes, from high-performance to 1MB tiny models.
- Broad cross-platform support across mobile, desktop, and embedded systems.
Related
- Project
- Project
- Project
- Project
- Project