prathoshap/vagdhenu
Vāgdhenu — metered Sanskrit/Vedic chant text-to-speech (DiT + BigVGAN). Apache-2.0.
What it solves
Vágdhenu is a production-grade text-to-speech (TTS) system specifically designed for Sanskrit chanting (páráyaána). Unlike standard TTS, which often reads text flatly, this system generates audio with metrically-aware durations and melodic contours faithful to traditional chanting styles.
How it works
The system uses a flow-matching Diffusion Transformer (DiT) backbone based on IndicF5/F5-TTS. To avoid issues like Hindi schwa-deletion, Sanskrit text is routed through the Kannada script. It employs a fine-tuned NVIDIA BigVGAN-v2 vocoder to ensure stability on long vowels.
Prosody is managed through a combination of a reference clip (controlling voice, swara, and pace) and voice-steering fine-tuning. The system also includes a specialized text frontend that handles complex Sanskrit linguistic rules, including visarga sandhi, homorganic anusvára, and meter/gaána detection.
Who it’s for
It is designed for researchers, developers, and practitioners interested in the synthesis of sacred Sanskrit recitation for study, accessibility, and the production of chanted audio content.
Highlights
- High Fidelity: Achieves a Mean Opinion Score (MOS) of ~4.6 from expert listeners.
- Linguistic Precision: Correctly renders complex conjuncts, including retroflex aspirates.
- Production Proven: Used to generate over 17 hours of audio for the Mahábhárata Tátparya Niráya and over 16,000 verses for the ŜśƱmad Bhágavatam.
- Advanced Frontend: Includes a custom text-processing pipeline for Devanagari to Kannada routing and meter detection.
Related
- Project
- Project
- Project
- Project
- Project