WAXAL: A Large-Scale Multilingual African Language Speech Corpus
Abstract
The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages (spanning 19 for ASR and 13 for TTS, with overlaps) representing over 100 million speakers. The collection consists of an Automatic Speech Recognition (ASR) dataset containing approximately 2,340 hours of transcribed, image-prompted natural speech (over 14,100 hours of total collected audio) from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with over 235 hours of high-quality, single-speaker studio recordings of phonetically balanced scripts. Crucially, this effort was executed in partnership with four African academic and community organizations, with the explicit goal of building the local ecosystem for speech technology and seeding data collection capabilities within these institutions. To demonstrate the dataset's utility, we benchmark three architecturally distinct ASR models---Gemma 3n-2B, Whisper-Large-v3, and MMS-1b-all---across the 19 ASR languages. Fine-tuning on WAXAL yields substantial improvements, reducing the macro-average Word Error Rate (WER) by up to 61\% (e.g., from 1.23 to 0.48 for Whisper). Models fine-tuned solely on WAXAL also generalize to the independently collected FLEURS benchmark, confirming the dataset's value for domain-transferable ASR adaptation. Furthermore, we employ an LLM-as-judge framework to evaluate semantic meaning preservation, demonstrating that fine-tuned predictions successfully capture core intents of native speakers. The WAXAL datasets are released at https://huggingface.co/datasets/google/WaxalNLP under the CC-BY-4.0 license to catalyze research and the development of inclusive speech technologies for speakers of African languages.