Source-linked AI summary

VAANI: Capturing the language landscape for an inclusive digital India

Sujith Pulikodan, Abhayjeet Singh, Agneedh Basu, Nihar Desai, Pavan Kumar J, Pranav D Bhat, Raghu Dharmaraju, Ritika Gupta, Sathvik Udupa, Saurabh Kumar, Sumit Sharma, Visruth Sanka, Dinesh Tewari, Harsh Dhand, Amrita Kamat, Sukhwinder Singh, Shikhar Vashishth, Partha Talukdar, Raj Acharya, Prasanta Kumar Ghosh

arXiv:2603.28714v3eess.AS

TL;DR

Indic speech resources often underrepresent languages, regional variation, and visual context. VAANI constructs a district-centered multimodal dataset using image-elicited spontaneous speech and extensive quality control, releasing broad coverage across languages, districts, and modalities while retaining important coverage limitations.

  • Problem

    Existing Indic speech corpora have narrow language coverage, language-centered sampling, and speech–text-only structures despite substantial linguistic and regional variation.

  • Method

    VAANI collects spontaneous speech elicited by images across districts, aligns recordings with images and available transcriptions, and applies automated and manual quality control.

  • Results

    VAANI releases 31,255 hours of audio, 2,043 hours of transcriptions, and 289,838 images across 165 districts and 105 languages.

  • Takeaways & Limitations

    VAANI provides a foundational open resource for multilingual speech and multimodal research involving underrepresented Indic languages and regional variation.

  • Takeaways & Limitations

    Audio and transcription coverage is small for many languages, and the 165-district scope leaves many regional dialects and minority languages unrepresented.

Abstract

from arXiv · show

Voice based technologies have the potential to bridge digital accessibility gaps; however, existing datasets fail to capture the linguistic and regional diversity of Indic languages. We present Project VAANI, a large scale multimodal dataset designed to represent India's linguistic landscape across 165 districts. Speech data is collected using image based prompts to elicit spontaneous responses, while images are curated through a separate pipeline covering diverse themes across regions. The dataset undergoes a rigorous multi stage quality control process, combining automated and manual evaluation to ensure high audio quality and transcription accuracy. We release approximately 289K images, 31,255 hours of speech, and 2,043 hours of transcribed audio spanning 105 languages from 28 states and 3 union territories. Many of these languages are represented at this scale for the first time, making VAANI a foundational resource for inclusive speech technology. The dataset enables the development of robust, multilingual, and multimodal models, and supports research in speech recognition, language understanding, and cross-modal learning for underrepresented languages.

1 Introduction

India’s linguistic and regional diversity creates a need for inclusive voice technologies and representative multimodal data. Existing linguistic variation spans many languages, dialects, and sociocultural contexts.

  • India’s linguistic diversity reflects historical, cultural, geographic, ethnic, social, political, and religious variation.
  • 22 languages are officially recognized under India’s Eighth Schedule.
  • Voice technologies can reduce dependence on text literacy and support inclusion in education, governance, health care, and digital communication.
  • Inclusive multimodal models require training data representing variation across languages, dialects, accents, and sociocultural contexts.

2 Related Work

Existing Indic speech resources provide limited language coverage and often treat languages as monolithic categories. VAANI addresses these gaps through district-centered, multimodal data collection spanning languages and regional variation.

  • Read speech offers controlled content and simpler annotation, whereas spontaneous speech captures more authentic and linguistically diverse language use.
  • No existing open Indic corpus covers more than 23 languages, leaving most of India’s 100+ mother tongues unrepresented.
  • Language varies across regions, communities, education levels, and genders, so language-only labels miss important sociolinguistic variation.
  • VAANI combines geo-centric sampling across 165 districts, 105 languages, and aligned image–speech–text triplets.
  • Most public Indic speech corpora pair audio with text but lack visual grounding for Indic multimodal model development.

3 The Dataset

VAANI is a large, India-representative multimodal dataset built across districts, languages, and speaker populations. It pairs image-elicited audio with transcriptions and metadata, while balancing geographic coverage during transcription.

  • VAANI contains 24,009,427 audio segments from 158,441 speakers responding to 289,838 images across 165 districts, 105 languages, 28 states, and 3 union territories.The collection totals approximately 31,255 hours of audio, including 2,043 manually transcribed hours.
  • Transcription coverage is distributed nearly evenly across the 165 districts to support balanced geographic representation.
  • The dataset spans Indo-Aryan, Dravidian, Tibeto-Burman, and Austroasiatic language families, with some varieties represented at this scale in an open speech dataset for the first time.
  • Districts were selected for census-language coverage and geographic representation, with recordings gathered from multiple locations to capture intra-district variation.
  • Speaker and recording metadata include age, education, socioeconomic status, and device information for quality control and inclusive representation.
  • The dataset is open-sourced under CC-BY-4.0 with district-level downloads and language-level access to transcribed data.

4 Data Collection

VAANI uses district-centered collection with image prompts to elicit spontaneous speech from demographically diverse speakers. Vendor-operated collection and multi-stage validation support image, audio, metadata, and transcription quality.

  • District-centered collection prioritizes people and demographic diversity across gender, age, education, and socioeconomic background.
  • Image prompts encourage natural, varied responses while remaining accessible to non-literate speakers and speakers of unscripted languages.
  • Vendors coordinated district-level collection, audio recording, transcription, and image acquisition across the operationally demanding pipeline.
  • Each district received 1,700–2,000 newly captured images combining district-specific and general topics.
  • Image acceptance required automated and human checks for quality, uniqueness, relevance, and compliance with project guidelines.
  • Mobile and web applications supported speaker onboarding, recording, metadata checks, and geographic validation using pincodes and GPS.
  • Transcription segments were selected using metadata, random sampling, speaker diversity, language balance, and district-level coverage analysis.

5 Quality Control

VAANI applies layered automated and manual quality assurance across audio, metadata, transcriptions, and vendor submissions. The workflow combines structural, technical, linguistic, contextual, and human validation before data integration.

  • Quality assurance pipeline: Multi-stage quality assurance evaluates audio, metadata, transcriptions, and formatting through automated and manual checks.Each batch receives a final compliance review before integration.
  • Automated validation: Level 1 validates metadata structure, file naming, paths, and consistency with recorded entries, including missing or duplicate identifiers.
  • Automated validation: Level 2 checks 16 kHz mono 16-bit audio, file integrity, naming conventions, and segment-level duration problems.
  • Automated validation: Level 3 compares each segment’s Signal-to-Noise Ratio against a threshold and flags failures for manual review.
  • Transcription quality control: Transcription checks assess metadata alignment, word-count thresholds, linguistic plausibility, WER, and geographic consistency using pincode comparisons.A significant pincode mismatch triggers a warning; the checks target linguistic, structural, and contextual accuracy.
  • Manual validation: Manual review samples 10% of data passing automated checks, includes at least one segment per speaker, and verifies audio, transcription, naturalness, and PII.Unexpected responses can receive full review, with centralized accept/reject decisions.

6 Experiments: Multilingual ASR Fine-Tuning

VAANI is used to fine-tune three multilingual ASR models on four well-represented and two low-resource Indic languages. Evaluation across external benchmarks shows that VAANI fine-tuning substantially improves performance across the resource spectrum.

  • Models and data: Three open-source ASR models—Gemma-3n-E2B, Whisper-large-v3-turbo, and Parakeet-tdt-0.6b-v2—are fine-tuned in a multilingual setting.
  • Models and data: Training covers Hindi, Bengali, Kannada, Telugu, Chakma, and Bhojpuri, with 80% train, 10% validation, and 10% test splits within each district.The splits prevent speaker overlap across partitions.
  • Evaluation: Evaluation uses multiple external benchmarks for Hindi, Bengali, Telugu, and Kannada, while Chakma and Bhojpuri are evaluated on VAANI because public benchmarks are unavailable.
  • Results: Fine-tuning on VAANI substantially improves ASR performance across major and low-resource languages.Non-finetuned baselines frequently produce incorrect scripts or repetitive hallucinations on under-resourced Indic languages.

7 Applications and Limitations

VAANI supports low-resource speech, translation, benchmarking, voice, and multimodal research through its scale, linguistic breadth, and aligned image–speech–text structure. Its coverage remains uneven across languages and geography, and speaker-diverse audio presents voice-cloning risks.

  • Applications: VAANI supports ASR, speech translation, culturally grounded benchmarking, voice conversion, speaker-adaptive TTS, and multimodal audio-visual modeling.Aligned image–speech–text triplets enable multimodal applications.
  • Limitations: Audio and transcription coverage is small for many of the 105 languages, with most data concentrated in a few majority languages.
  • Limitations: VAANI spans 165 of roughly 800 districts, leaving some regional dialects and minority languages unrepresented.
  • Limitations: Speaker-diverse audio may be misused for voice cloning or impersonation, prompting anonymization, PII removal, and consent-based collection.

8 Conclusions and Future work

VAANI is a rigorously curated multimodal dataset spanning 31,255 hours of audio, 2,043 hours of transcriptions, and 289,838 images across 165 districts and 105 languages. The authors report preliminary utility across speech and multimodal tasks and plan broader low-resource and geographic coverage.

  • Contribution: VAANI comprises 31,255 hours of audio, 2,043 hours of transcriptions, and 289,838 images across 165 districts and 105 languages.
  • Contribution: Preliminary experiments confirm utility across speech and multimodal tasks, while the dataset’s diversity supports accent-diverse and low-resource ASR research.
  • Impact and future work: Open-sourcing VAANI under a permissive license is intended to lower barriers to Indic speech research.
  • Impact and future work: Future phases aim to improve audio and transcription coverage for low-resource languages and expand the dataset to additional districts.

A Experiments

VAANI’s experiments evaluate whether its multimodal, geographically diverse data improves language-specific and region-specific speech recognition, while also testing image–text retrieval. Fine-tuning produces consistent ASR gains, geographically patterned Hindi performance, and improved image retrieval.

  • Dataset and experimental scope: 2,043 hours of transcribed speech form the multimodal corpus used for preliminary speech and visual experiments.The experiments use selected subsets of VAANI’s speech and visual modalities.
  • ASR fine-tuning: Whisper-small models were independently fine-tuned for Hindi, Bengali, Kannada, and Telugu using 331, 101, 80.2, and 69.0 hours, respectively.The models were evaluated across multiple benchmark datasets.
  • ASR fine-tuning: Consistent performance improvements were observed across all four languages relative to the pretrained open-source baseline.The result supports VAANI’s use for fine-tuning robust ASR models.
  • Region-specific fine-tuning: State-specific Hindi models performed comparatively better on geographically nearby states than on distant states, even within the same language.The evaluation fine-tuned separate models using state-level Hindi data and tested them across multiple states.
  • Image retrieval: SigLIP2 was fine-tuned on approximately 45,000 Hindi image–transcription pairs and evaluated on around 9,500 test images, improving over the original model.The visual experiment used image retrieval as its evaluation task.

B Image Data Specifications and Collection Guidelines

VAANI’s image pipeline uses district-specific, physically captured images as prompts for spontaneous speech, with specifications covering diversity, authenticity, technical format, participants, metadata, and quality control. The collection process links regional visual grounding to consistent audio capture and organized delivery.

  • Collection approach: Images are curated under detailed vendor specifications to ensure quality, diversity, and regional relevance before serving as visual prompts for spontaneous speech.Speakers describe each image in their own words, so image quality and regional authenticity influence the resulting speech corpus.
  • Image requirements: Images must be physically captured in the target district and depict district-specific sites rather than generic topics.The guidelines prohibit photographing existing images or sourcing them from secondary and tertiary platforms.
  • Image requirements: Each topic should contain equal image counts within a ±20% tolerance, with files delivered as .jpg images at 640 × 400 pixels or a 16:9 aspect ratio and under 500 KB.The specifications also require shooting-date metadata no earlier than July 1, 2023.
  • Audio collection: Speech collection uses randomly selected image prompts to elicit spontaneous 10–20 second utterances with varied vocabulary.Audio is collected in a specified format and without transcoding or post-processing.
  • Recording conditions: Audio collection requires quiet, non-echoey settings and stable microphone placement, including a device distance of no more than two feet from the speaker.The protocol specifies front-facing microphone placement and consistent distance and orientation.
  • Participant requirements: Participant sampling targets adults aged 20–70, gender balance, at least 800–820 speakers per district, and no more than 15 minutes of effective speech per speaker.Participants are required to be local natives and encouraged to speak the language or dialect used at home.
  • Data organization: Deliverables are organized hierarchically by district and speaker, with audio files, per-audio TSV metadata, speaker metadata, transcription files, and reference files.Speaker records use fields such as anonymized IDs, device, gender, age, education, pincode, socioeconomic status, and languages.

F Speaker Demographics

The 158,441-speaker enrolled pool is near-balanced by gender but skews young and toward middle education and socio-economic strata. Older, less-educated, and economically peripheral groups remain represented, though more thinly.

  • Gender: 52.81% of speakers were female, 47.15% male, and 0.04% identified as other.
  • Age: 61.21% of speakers were aged 20–30, while representation declined across older brackets.The 30–40 bracket contributed 17.14%, followed by 11.08% for 40–50 and 6.67% for 50–60.
  • Education: 37.99% of speakers were 12th-pass, 33.79% were graduates, and 15.45% were 10th-pass.The remaining approximately 13% comprised postgraduate, below-10th, and primary-level speakers.
  • Socio-economic status: Lower-Middle was the largest socio-economic class at 25.65%, followed by Middle at 20.52% and Upper-Middle at 18.81%.Together, these three strata accounted for almost two-thirds of the pool; Upper-Lower, Lower, and Upper groups were also represented.
  • Overall composition: The enrolled pool was gender-balanced and skewed young and toward middle education and economic strata, with thinner representation in demographic tails.

G District-wise and Language-wise Data Distribution

The released data are organized for district-wise and language-wise reporting, pairing geographic or linguistic coverage with audio and transcription durations. The dataset totals 31,255.45 hours of audio and 2,043.31 hours of transcription.

  • District-wise distribution: Table 4 reports each district’s distinct language count together with audio and transcription duration.The table is explicitly organized around district-level distribution.
  • Language-wise distribution: Table 5 aggregates audio and transcription duration by language and reports the number of districts where each language was collected.This complements the district-wise view with language-level coverage across districts.
  • Overall totals: 31,255.45 hours of audio and 2,043.31 hours of transcription were released in total.
Loading 2603.28714v3…