DrewThomasson/ebook2audiobook

Generate audiobooks from e-books, voice cloning & 1158+ languages!

What it solves

Ebook2audiobook (E2A) converts non-DRM eBooks into high-quality audiobooks with chapters and metadata. It provides a way to turn digital books into spoken audio using various AI-powered text-to-speech (TTS) engines, supporting a vast array of languages and formats.

How it works

The tool processes eBook files (such as .epub or .mobi) and uses a selected TTS engine to synthesize speech. It can be run via a Gradio-based web interface or in a headless CLI mode. Users can choose from multiple TTS engines (like XTTSv2, Bark, and Tortoise) and optionally provide a voice cloning file to mimic a specific voice. It also supports OCR scanning for image-based text pages and SML tags for fine-grained control over pauses and voice switching.

Who it’s for

Readers who want to convert their legal, non-DRM eBook collections into audiobooks, as well as users who want to customize the voice and pacing of their audiobooks using advanced TTS models.

Highlights

  • Broad Engine Support: Supports multiple TTS engines including XTTSv2, Bark, Fairseq, VITS, Tacotron2, Tortoise, GlowTTS, and YourTTS.
  • Extensive Language Support: Compatible with 1158 languages and dialects.
  • Flexible Formats: Converts a wide variety of input formats (including .epub, .pdf, .mobi, and .docx) and outputs to multiple audio formats (such as .m4b, .mp3, and .flac).
  • Voice Cloning: Ability to use a custom voice file for cloning.
  • Low Resource Requirements: Can run on as little as 2GB RAM and 1GB VRAM.
  • SML Tagging: Supports tags like [break], [pause], and [voice] for manual audio control.

Related

  • Project
  • Project
  • Project
  • Project
  • Project