Xerophayze/TTS-Story
TTS-Story is a web-based multi‑voice TTS studio for turning tagged scripts into audiobooks—featuring full speaker management, chunk review/regeneration, a job queue and library system, and local GPU or API backends including Kokoro, Chatterbox, VOX CPM, Pocket-TTS, Kitten-TTS, IndexTTS-2, QWEN3 TTS and Omnivoice engines
What it solves
TTS-Story simplifies the creation of narrated stories and audiobooks by providing a unified web interface for multiple text-to-speech (TTS) engines. It eliminates the need to manually manage disparate TTS tools, handling complex tasks like multi-speaker assignment, voice cloning, and audiobook formatting (MP3/M4B) in one workflow.
How it works
The application acts as a management layer over eighteen different TTS engines, including local CPU/GPU models (like Kokoro, Qwen3, and OmniVoice) and cloud providers (like ElevenLabs and Azure). Users can tag text with speaker labels (e.g., [alice-female]...[/alice-female]) to assign different voices to different characters. It also integrates LLMs (via Gemini, Ollama, etc.) to help prepare and clean up manuscripts before synthesis. Local engines are installed in isolated virtual environments to prevent dependency conflicts.
Who it’s for
It is designed for authors, content creators, and audiobook producers who need high-quality, multi-voice narration with precise control over speaker casting, timing, and audio export.
Highlights
- Extensive Engine Support: Access to 18 TTS options across local hardware and cloud services.
- Multi-Speaker Workflow: Use of speaker tags for seamless character-based narration.
- Voice Cloning: Built-in tools for uploading reference audio to clone specific voices.
- Audiobook Tooling: Automatic chapter detection and professional export formats including M4B.
- LLM Integration: Optional text preparation using various LLM providers to refine manuscripts.
- Isolated Environments: Local engines are managed in separate environments to ensure stability.
相关
- 项目
- 项目
- 项目
- 项目
- 项目