Picovoice/leopard

On-device speech-to-text engine powered by deep learning

What it solves

Leopard provides a way to perform high-accuracy speech-to-text transcription directly on a device, eliminating the need to send voice data to a cloud server. This ensures user privacy and allows for offline functionality.

How it works

It is an on-device speech-to-text engine that processes audio files or microphone input locally. It uses a model-based approach where users can use a default model or a custom-trained model. The engine requires an AccessKey for authentication and authorization, which is validated via Picovoice license servers, although the actual voice recognition processing remains 100% offline.

Who it’s for

Developers building cross-platform applications that require private, efficient, and accurate speech-to-text capabilities across desktop (Windows, macOS, Linux), mobile (iOS, Android), web (Chrome, Safari, Firefox, Edge), and embedded devices (Raspberry Pi).

Highlights

  • Privacy-First: All voice processing runs locally on the device.
  • Broad Platform Support: Compatible with a wide range of operating systems, browsers, and hardware.
  • Computationally Efficient: Designed to be compact and efficient in its resource usage.
  • Multi-language Support: Supports English, French, German, Italian, Japanese, Korean, Portuguese, and Spanish.