EXPOSING AND CORRECTING SPEAKER LEAKAGE IN YORUBA DIALECT IDENTIFICATION WITH SELF-SUPERVISED SPEECH REPRESENTATION
Ayodele Awokoya
Abstract
Benchmark datasets released without disclosing the training/test partitioning details pose a risk that is easy to miss and costly when missed (Kuparinen, 2026). The evaluation of splits that leak identifying information about individual data-generating entities, in this case, speakers between training and test partitions, inflates reported model performance and misleads downstream comparisons (Kuparinen, 2026). We document and quantify this problem using YORÙLECT (Ahia et al., 2024), a recently released Yorùbá dialect speech corpus spanning four regional varieties, using a three-class dialect identification task (Standard Yorùbá, Ifè, and Ìlàje, which have approximately 9 hours and 4 speakers per dialect) as a case study. Although YORÙLECT's public metadata does not provide an explicit speaker identifier, we reconstructed the speaker identity from the available fields and showed that its official train/validation/test split reuses the same speakers in all three partitions, such that every speaker who appears in training also appears in both validation and test. We built a corrected, verified speaker-independent evaluation protocol, a four-leave-one-speaker-out fold, and benchmarked dialect classification using frozen self-supervised speech representations (AfriHuBERT (Alabi et al., 2025), an African-language-adapted encoder) with a logistic regression classifier, alongside classical MFCC and majority-class baselines. $$\text{Table 1: Classification performance comparison (mean ± standard deviation across 4 folds for S1, or single-split for SD).}$$ | Method | Evaluation Split | Accuracy (Mean) | Macro F1 (Mean) | | :--- | :---: | ---: | ---: | | Majority Baseline | Speaker-Independent (SI) | 33.34% ± 0.01% | 16.67% ± 0.01% | | MFCC + SVM | Speaker-Independent (SI) | 73.40% ± 10.29% | 72.73% ± 10.68% | | AfriHuBERT + LogReg | Speaker-Independent (SI) | 89.89% ± 9.23% | 89.82% ± 9.33% | | AfriHuBERT + LogReg | Speaker-Dependent (SD) | 98.27% | 98.27% | Under our speaker-independent protocol, accuracy is 89.9% (± 9.2 across folds). Upon evaluating our best model on YORÙLECT's original split, the accuracy increased to 98.3%, an 8.4-point inflation attributable entirely to speaker leakage. We further showed, via silhouette analysis of the embedding space, that individual speaker identity is nearly as strong a clustering signal (0.152) as dialect (0.165) in the frozen representation, indicating that self-supervised speech encoders entangle speaker identity with dialect-relevant content. This is consistently in the same direction with concurrent evidence from unrelated language families (Finnish and Norwegian (Kuparinen, 2026)), suggesting that speaker leakage in dialect benchmarks may be a broader, underappreciated evaluation problem rather than a Yorùbá-specific artifact. While our self-supervised model performs best at feature extraction, correcting speaker leakage remains an active research concern. Future work will investigate the use of the Factorized Hierarchical Variational Autoencoder (FHVAE) to separate phonetic content from speaker factors (Shon et al., 2018) or the use of retrieval-based voice conversion to map all speakers to a uniform target voice (Fischbach et al., 2025). The built corrected speaker-independent fold dataset definitions for YORÙLECT will be openly released to support reproducible evaluation, and we strongly opine that entity-level leakage auditing should be a standard step when adopting newly released low-resource benchmarks, regardless of modality.
Chat is not available.
Successful Page Load