QuentinFuxa/WhisperLiveKit
Simultaneous speech-to-text models
What it solves
WhisperLiveKit は、超低遅延、セルフホスト型の音声からテキストへの変換 (STT) パイプラインを提供します。標準的な Whisper は完全な発話内容を処理するために設計されており、リアルタイムの小さな音声クリップを処理する際、精度が低下したり単語が途切れたりすることがよくあります。この問題を解決します。
How it works
このプロジェクトは、AlignAtt SimulStreaming や LocalAgreement ポリシーを含む、最先端の同時音声研究に基づき、
Who it’s for
リアルタイム文字起こしサービス、200 ヶ国語の同時翻訳アプリ、または話者分離 (speaker diarization) 機能付きのセルフホスト型低遅延 ASR (Automatic Speech Recognition) を必要とする開発者向けに設計されています。
Highlights
- Ultra-nlow latency: 高度なストリーミングポリシーを使用して、リアルタイムの文字起こしを実現します。
- Multiple Backends: Whisper, Voxtral Mini (4B parameters), および Qwen3-ASR を含む幅広いモデルをサポートします。
- Broad API Compatibility: OpenAI と Deepgram API の即用的な代替手段として利用可能です。 extinclude_error: textinclude_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error:
\n
**
_**
Multiple Backends: Whisper, Voxtral Mini (4B parameters), および Qwen3-ASR を含む幅広いモデルをサポートします。
Broad API Compatibility: OpenAI と Deepgram API の即用的な代替手段として利用可能です。
Multilingual Translation: NLLW を介して 200 ヶ国語への同時翻訳を実現します。
Speaker Diarization: Sortformer または Diart を使用したリアルタイム話者識別。
Hardware Optimized: NVIDIA GPU (CUDA) と Apple Silicon (MLX/Metal) に特化したサポート。