Dialect ASR based on Multi-View Pseudo-Parallel Augmentation and Noise-Robust Contrastive Learning
Abstract
Recent advancements in Automatic Speech Recognition (ASR) have significantly improved performance for high-resource languages, yet ASR systems still struggle with low-resource languages and dialects due to the scarcity of annotated training data. To address this challenge, we propose a novel framework that enhances the utilization of limited data through a two-pronged approach: (1) a novel data augmentation technique based on a multi-view pseudo-parallel strategy, and (2) a robust multi-level contrastive learning framework capable of jointly leveraging semantic and dialect-specific attributes to improve model effectiveness under noisy conditions. Our approach effectively generalizes across dialectal variations when dialect data are non-parallel, allowing acquisition of shared linguistic structures and dialectal distinctions. Furthermore, our method could be simply transferred to scenarios where dialect data are parallel. We validate our method by adapting the Whisper-large model for Swiss German, Chinese and Arabic dialect ASR tasks, demonstrating substantial performance gains. Specifically, our framework achieves up to a 48.59% reduction in WER and 23.16% reduction in CER under non-parallel conditions, outperforming state-of-the-art baselines. These results highlight the robustness and adaptability of our approach for low-resource dialectal ASR. The data and codes will be public available upon acceptance.