Picovoice/leopard

On-device speech-to-text engine powered by deep learning

What it solves

Leopard 提供了一種直接在裝置上進行高準確度語音轉文字轉錄的方法,無需將語音數據傳送到雲端伺服器。這確保了使用者隱私並允許離線功能。

How it works

它是一款在裝置端處理音訊檔案或麥克風輸入的語音轉文字引擎。它採用基於模型的做法,使用者可以使用預設模型或自定義訓練的模型。該引擎需要一個 AccessKey 用於身份驗證與授權,該金鑰會透過 Picovoice 授權伺服器進行驗證,但實際的語音辨識處理過程仍然是 100% 離線的。

Who it’s for

開發人員正在構建需要跨平台(包括桌面端(Windows, macOS, Linux)、行動端(iOS, Android)、網頁端(Chrome, Safari, Firefox, Edge)以及嵌入式裝置(Raspberry Pi))提供私密、高效且準確的語音轉文字功能的跨平台應用程式。

Highlights

  • Privacy-First: All voice processing runs locally on the device.
  • Broad Platform Support: Compatible with a wide range of operating systems, browsers, and hardware.
  • Computationally Efficient: Designed to be compact and efficient in its resource usage.
  • Multi-language Support: Supports English, French, German, Italian, Japanese, Korean, Portuguese, and Spanish.