Source-linked AI summary
WAXAL: A Large-Scale Multilingual African Language Speech Corpus
Abdoulaye Diack, Perry Nelson, Kwaku Agbesi, Angela Nakalembe, MohamedElfatih MohamedKhair, Vusumuzi Dube, Tavonga Siyavora, Subhashini Venugopalan, Jason Hickey, Uche Okonkwo, Abhishek Bapna, Isaac Wiafe, Raynard Dodzi Helegah, Elikem Doe Atsakpo, Charles Nutrokpor, Fiifi Baffoe Payin Winful, Kafui Kwashie Solaga, Jamal-Deen Abdulai, Akon Obu Ekpezu, Audace Niyonkuru, Samuel Rutunda, Boris Ishimwe, Michael Melese, Engineer Bainomugisha, Joyce Nakatumba-Nabende, Andrew Katumba, Claire Babirye, Jonathan Mukiibi, Vincent Kimani, Samuel Kibacia, James Maina, Fridah Emmah, Ahmed Ibrahim Shekarau, Ibrahim Shehu Adamu, Yusuf Abdullahi, Howard Lakougna, Bob MacDonald, Hadar Shemtov, Aisha Walcott-Bryant, Moustapha Cisse, Avinatan Hassidim, Jeff Dean, Yossi Matias
TL;DR
Speech technology remains constrained by the scarcity of large, high-quality, permissively licensed corpora for Sub-Saharan African languages. WAXAL addresses this gap with partner-collected ASR and TTS datasets spanning 24 languages, while documenting coverage and ethical limitations. Its release provides a foundation for model development, evaluation, and linguistic analysis.
Problem
Sub-Saharan African speech technology lacks large-scale, high-quality, permissively licensed corpora despite substantial linguistic diversity and community needs.
Method
WAXAL constructs ASR and TTS resources through four African partners, using image-prompted natural speech, local transcription, quality control, and phonetically balanced studio recordings.
Results
WAXAL releases ASR data spanning 14 languages and TTS data spanning 13 languages, occupying 1.7 TB and 99 GB respectively.
Takeaways & Limitations
The collection provides a foundational resource for building and evaluating models, conducting linguistic analysis, and supporting technologies for represented communities.
Takeaways & Limitations
Only 10% of collected ASR audio is transcribed, while dialectal variation, unintended content, privacy risks, and voice misuse remain documented limitations or ethical concerns.
Abstract
from arXiv · showhide
The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages representing over 100 million speakers. The collection consists of two main components: an Automated Speech Recognition (ASR) dataset containing approximately 1,250 hours of transcribed, natural speech from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with around 235 hours of high-quality, single-speaker recordings reading phonetically balanced scripts. This paper details our methodology for data collection, annotation, and quality control, which involved partnerships with four African academic and community organizations. We provide a detailed statistical overview of the dataset and discuss its potential limitations and ethical considerations. The WAXAL datasets are released at https://huggingface.co/datasets/google/WaxalNLP under the permissive CC-BY-4.0 license to catalyze research, enable the development of inclusive technologies, and serve as a vital resource for the digital preservation of these languages.
1. INTRODUCTION
WAXAL addresses the scarcity of large-scale, high-quality speech corpora for underserved Sub-Saharan African languages by releasing ASR and TTS resources across 24 languages.
- Speech technologies remain concentrated in high-resource languages, while many Sub-Saharan African speakers lack access in their native tongues.
- WAXAL introduces a large-scale resource covering 24 Sub-Saharan African languages, collected with four African academic and community partners.
- Approximately 1,250 hours of transcribed, image-prompted natural speech support ASR training and evaluation.
- The collection also provides over 180 hours of studio-quality, phonetically balanced recordings for TTS in 10 languages.
- The datasets include speaker metadata, local-expert transcriptions, and a CC-BY-4.0 license for academic and commercial research.
2. RELATED WORKS
Existing multilingual speech datasets broaden linguistic coverage, but African-language resources remain limited in scale, speaker diversity, or language coverage. WAXAL unifies substantially larger ASR and TTS resources for 24 Sub-Saharan African languages.
- Common Voice, FLEURS, BABEL, and CMU Wilderness provide multilingual speech resources spanning many languages and use cases.
- Existing Sub-Saharan African datasets often remain limited in scale, scope, or speaker diversity.
- WAXAL introduces a unified dataset for 24 Sub-Saharan African languages, expanding available ASR and TTS data.
- The release combines approximately 1,250 hours of multi-speaker ASR speech with over 180 hours of single-speaker TTS recordings.
3. DATA COLLECTION METHODOLOGY
WAXAL combines partner-led collection, image-prompted natural speech for ASR, and phonetically balanced studio recordings for TTS, with local transcription and quality control.
- Partner-led collection: Four African partners collected WAXAL data during a multi-year effort from January 2021 to March 2024.
- ASR data collection: ASR participants described diverse images covering at least 50 topics in their native languages to elicit natural speech.
- ASR data collection: ASR recordings lasted at least 15 seconds, targeted gender and age diversity, and underwent local-expert transcription and quality control.
- TTS data collection: TTS used approximately 108,500-word phonetically balanced scripts for each of 10 target languages.
- TTS data collection: Seventy-two community voice actors recorded approximately 16 hours each in a professional studio-like environment.
4. DATASET STATISTICS
WAXAL comprises separate ASR and TTS releases with broad language coverage and substantial storage requirements, summarized in the dataset statistics.
- The ASR dataset spans 14 languages, while the TTS dataset spans 13 languages.
- The released ASR data occupies 1.7 TB, compared with 99 GB for the TTS data.
5. LIMITATIONS AND CONSIDERATIONS
WAXAL’s release involves limitations in coverage and use-case fit, alongside ethical risks concerning privacy, consent, and downstream voice use. The authors describe mitigation measures while acknowledging remaining boundaries.
- Limitations: 10% of the collected ASR audio currently has transcriptions, limiting the coverage of the released transcribed corpus.
- Limitations: The dataset may not capture the full dialectal and socio-linguistic variation within each represented language.
- Limitations: The unscripted ASR data may contain offensive or inappropriate speech, although manual quality control mitigates this risk.
- Limitations: The diverse ASR dataset is not well suited to training high-quality single-speaker TTS models.
- Ethical Considerations: Participants provided informed consent, and personally identifiable information was removed from transcripts and metadata.
- Ethical Considerations: Language and accent may still reveal ethnicity or race, while released TTS voices could be used in unforeseen ways.
6. CONCLUSION
WAXAL addresses speech-resource scarcity for Sub-Saharan African languages by releasing approximately 1,500 hours of annotated ASR and TTS data across 24 languages. The resource is intended to support model development, linguistic analysis, and inclusive technologies under a permissive license.
- Approximately 1,500 hours of annotated ASR and TTS data across 24 languages provide a foundational resource for speech research.
- The dataset supports building and evaluating models, conducting linguistic analysis, and developing technologies for represented language communities.
- WAXAL was developed through collaborative methodology emphasizing local expertise and ethical data handling.
- The datasets are publicly released under the CC-BY-4.0 license.