Task-Lens: Cross-Task Utility Based Speech Dataset Profiling for Low-Resource Indian Languages
Abstract
The rapid growth of publicly available Indian speech datasets has accelerated research in multilingual speech processing. However, these datasets are typically developed for individual downstream tasks, making it difficult for researchers to assess their broader applicability across multiple speech applications. We present Task-Lens, a systematic framework for profiling the cross-task utility of Indian speech datasets. Following a PRISMA-guided review, we curate 50 publicly available datasets comprising 91k+ hours of speech across 26 Indian languages and evaluate their suitability for nine downstream speech tasks using a task-feature relevance matrix. Task-Lens investigates three key research questions: (1) identifying the downstream tasks supported by each dataset, (2) quantifying task-wise resource availability, and (3) analyzing language-wise coverage across Task-Ready datasets. Our analysis reveals substantial disparities in metadata completeness, task coverage, and language representation. No dataset was Task-Ready for all nine downstream tasks, and only 9 of the 50 datasets satisfied the requirements of seven tasks. While Automatic Speech Recognition (ASR), Speaker Recognition (SR), Monolingual Text-to-Speech (MO-TTS), and Multilingual Text-to-Speech (ML-TTS) account for the majority of available speech hours, Speech Emotion Recognition (SER) remains the least-resourced task due to the scarcity of emotion annotations. Furthermore, although Task-Ready datasets span 26 Indian languages, most speech hours are concentrated in just eight high-resource languages, exposing significant gaps for low-resource languages. Task-Lens provides researchers with a practical resource for dataset discovery, cross-task reuse, and identifying high-impact directions for future multilingual speech dataset development. Future work will extend Task-Lens through automated metadata quality assessment, continual dataset integration, and benchmarking for multilingual foundation models.