WEIFENG2333/VideoCaptioner

🎬 卡卡字幕助手 | VideoCaptioner - 基于 LLM 的智能字幕助手 - 视频字幕生成、断句、校正、字幕翻译全流程处理!- A powered tool for easy and efficient video subtitling.

What it solves

VideoCaptioner provides a one-stop solution for processing video subtitles. It eliminates the need to manually handle multiple separate tools for speech-to-text transcription, subtitle refinement, translation, and final video synthesis.

How it works

The tool follows a sequential pipeline: Audio/Video Input $\rightarrow$ Speech Recognition $\rightarrow$ Subtitle Segmentation $\rightarrow$ LLM Optimization $\rightarrow$ Translation $\rightarrow$ Video Synthesis. It uses word-level timestamps and Voice Activity Detection (VAD) for accuracy, and leverages Large Language Models (LLMs) for semantic segmentation and context-aware translation to ensure subtitles are natural and readable.

Who it’s for

Content creators and video editors who need to automate the transcription and translation of videos, as well as those who want to use LLMs to improve the quality of subtitles through semantic optimization.

Highlights

  • Comprehensive Pipeline: Handles everything from downloading videos (YouTube/Bilibili) to transcription, translation, and burning subtitles into the video.
  • Flexible Engine Support: Supports multiple ASR engines including faster-whisper, whisper-api, bijian, and jianying.
  • LLM Integration: Uses OpenAI-compatible APIs for high-quality subtitle optimization and translation.
  • Multiple Interfaces: Available as a CLI tool and a GUI desktop application for different user preferences.
  • Claude Code Skill: Includes a dedicated skill for integration with the Claude Code AI programming assistant.

Related

  • Project
  • Project
  • Project
  • Project
  • Project