Hugging Face and IISc Partner to Open-Source the Vaani Dataset

Hugging Face has partnered with the Indian Institute of Science (IISc) and ARTPARK to provide global access to Vaani, a massive multi-modal, multi-lingual dataset designed to represent India's linguistic diversity. This collaboration aims to increase the accessibility and usability of the dataset to foster the development of AI systems that better serve the digital needs of India's diverse linguistic population.

The Vaani Dataset: Scope and Methodology

The Vaani dataset is an open-source initiative launched in 2022 by IISc/ARTPARK and Google. It employs a geo-centric approach to collect dialects and languages from remote regions, ensuring that the data represents more than just mainstream languages.

Data Collection Goals

Vaani targets the collection of of the following data points from 1 million people across all 773 districts of India:

  • Total Speech Data: Over 150,000 hours.
  • Transcribed Text Data: 15,000 hours.

Implementation Phases

  • Phase 1: Covered 80 districts and has already been open-sourced.
  • Phase 2: Currently underway, expanding the dataset to an additional 100 districts, extending coverage to all Indian states.

Specialized Data Subsets

To facilitate specific AI tasks, a transcribed subset of the larger Vaani dataset has been open-sourced. This subset contains 790 hours of transcribed audio from approximately 700,000 speakers covering 70,000 images. It is specifically designed for:

  • Speech Recognition: Training models to accurately transcribe spoken language.
  • Language Modeling: Developing more refined language models.
  • Segmentation Tasks: Identifying distinct speech units to improve transcription accuracy.

AI Applications and Utility

The Vaani dataset covers 54 languages and includes spontaneous speech data from diverse geographical, educational, and socio-economic backgrounds. This breadth of data enables the development of several types of AI models:

  • Speech-to-Text (STT) and Text-to-Speech (TTS): Fine-tuning models for both LLM and non-LLM applications, including code-switching ASR models (Indic and English).
  • Foundational Speech Models: Creating robust models for Indic languages based on significant linguistic and geographical coverage.
  • Speaker and Language Identification: Developing models for speaker verification and identification, utilizing data from over 80,000 speakers.
  • Speech Enhancement: Utilizing the dataset's tagging system to build advanced speech enhancement technologies.
  • Multimodal LLMs: Improving the multimodal capabilities of LLMs when combined with other datasets.
  • Performance Benchmarking: Providing a real-world, diverse dataset for benchmarking speech models.

These capabilities can be applied to real-world conversational AI applications in telemedicine, healthcare, educational tools, voter helplines, and media localization.

Sources