denizsafak/abogen

Generate audiobooks from EPUBs, PDFs and text with synchronized captions.

What it solves

Abogen is a text-to-speech (TTS) tool designed to convert long-form documents—such as ePubs, PDFs, text files, and markdown—into high-quality audiobooks or voiceovers with perfectly synced subtitles. It eliminates the manual effort of timing subtitles and managing long texts for TTS generation.

How it works

The tool leverages the Kokoro-82M model for natural-sounding speech synthesis. It provides two distinct interfaces: a PyQt6 desktop application for stable core features and a Flask-based Web UI for advanced features. It supports a wide range of input formats (ePub, PDF, TXT, MD, SRT, ASS, VTT) and can output audio in various formats like WAV, MP3, and M4B (with chapters).

Who it’s for

It is intended for creators making audiobooks, voiceovers for social media (Instagram, YouTube, TikTok), and users who want to convert their digital reading materials into audio format.

Highlights

  • Multi-format Support: Handles ePub, PDF, and markdown files, including chapter-aware processing for e-books.
  • Custom Voice Mixing: A voice mixer allows users to create unique voices by blending different voice models with adjustable weights.
  • Flexible Subtitle Generation: Offers multiple subtitle styles (by sentence, word, or highlighting) and formats (SRT, ASS).
  • Cuda/ROCm Support: Optimized for NVIDIA and AMD GPUs to ensure fast processing speeds.
  • Advanced Web UI: Includes additional capabilities like LLM-based text normalization and Audiobookshelf integration.

Related

  • Project
  • Project
  • Project
  • Project
  • Project