QuentinFuxa/WhisperLiveKit

Simultaneous speech-to-text models

What it solves

WhisperLiveKit は、超低遅延、セルフホスト型の音声からテキストへの変換 (STT) パイプラインを提供します。標準的な Whisper は完全な発話内容を処理するために設計されており、リアルタイムの小さな音声クリップを処理する際、精度が低下したり単語が途切れたりすることがよくあります。この問題を解決します。

How it works

このプロジェクトは、AlignAtt SimulStreaming や LocalAgreement ポリシーを含む、最先端の同時音声研究に基づき、

Who it’s for

リアルタイム文字起こしサービス、200 ヶ国語の同時翻訳アプリ、または話者分離 (speaker diarization) 機能付きのセルフホスト型低遅延 ASR (Automatic Speech Recognition) を必要とする開発者向けに設計されています。

Highlights

  • Ultra-nlow latency: 高度なストリーミングポリシーを使用して、リアルタイムの文字起こしを実現します。
  • Multiple Backends: Whisper, Voxtral Mini (4B parameters), および Qwen3-ASR を含む幅広いモデルをサポートします。
  • Broad API Compatibility: OpenAI と Deepgram API の即用的な代替手段として利用可能です。 extinclude_error: textinclude_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_self_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error: textinclude_error:

\n

**

_**

  • Multiple Backends: Whisper, Voxtral Mini (4B parameters), および Qwen3-ASR を含む幅広いモデルをサポートします。

  • Broad API Compatibility: OpenAI と Deepgram API の即用的な代替手段として利用可能です。

  • Multilingual Translation: NLLW を介して 200 ヶ国語への同時翻訳を実現します。

  • Speaker Diarization: Sortformer または Diart を使用したリアルタイム話者識別。

  • Hardware Optimized: NVIDIA GPU (CUDA) と Apple Silicon (MLX/Metal) に特化したサポート。