MatteoFasulo/Whisper-TikTok

From AI tools to TikTok video creation using FFMPEG, Microsoft Edge read aloud and OpenAI Whisper model

What it solves

Whisper-TikTok automates the creation of engaging TikTok videos by combining transcription, voiceover, and visual assembly. It eliminates the need for manual subtitling and the use of robotic-sounding AI voices, providing a more natural auditory experience.

How it works

The tool integrates three core technologies:

  1. OpenAI-Whisper (via stable-ts) to generate accurate transcriptions from audio files.
  2. Microsoft Edge Cloud TTS API to create high-quality, natural-sounding voiceovers.
  3. FFMPEG to assemble the final video, including the application of subtitles and background visuals.

Who it’s for

Content creators who want to quickly generate TikTok-style videos with accurate captions and high-quality AI voiceovers without manually editing each clip.

Highlights

  • Natural Voiceovers: Uses Edge TTS for a more authentic sound than typical TikTok AI voices.
  • Customizable Subtitles: Allows users to adjust font size, color, and position.
  • Harnesses YouTube: Supports using YouTube URLs as background video sources.
  • Flexible CLI: Provides a wide range of command-line options for voice selection, language, and styling.

Related

  • Project
  • Project
  • Project
  • Project
  • Project