Towards Robust Context Utilisation in Multilingual ASR for Indian Languages
Abstract
Automatic speech recognition (ASR) is increasingly deployed where auxiliary context, such as domain terminology, named entity lists, or descriptions of spoken content, is available at inference time. Audio Large Language Models (AudioLLMs) provide a natural interface for such information through textual prompts, but it remains unclear whether they reliably use supplied context or rely on parametric knowledge. IndicContextEval [1] reveals substantial variation across eight Indian languages: models may benefit from relevant context, show limited sensitivity to it, or become susceptible to misleading context. One open model degrades by 9.22 WER points with adversarial, wrong-domain entities, while gains from correct context are largest when entities are supplied in the native script. These findings motivate training ASR systems to use relevant context while remaining robust to incorrect context, particularly for low-resource Indian languages. Building on IndicContextEval [1], we extend the problem from evaluation to training by developing a multilingual, context-aware ASR model for Indian languages. We build on an open AudioLLM with a prompt-and-audio interface. Starting from existing speech and transcript data, we construct contextual training examples by associating each utterance with relevant native-script entity lists and natural-language domain descriptions. Context-conditioned training, joint multilingual training, and adversarial context-robustness training are used to encourage relevant context utilisation while reducing reliance on misleading context. We will evaluate using the controlled L0-L6 protocol of IndicContextEval, reporting Word Error Rate (WER), Named Entity Error Rate (NEER), and Orthographically-Informed Word Error Rate (OIWER). We will examine improvements in relevant context utilisation, robustness to incorrect context, and named entity recognition across eight Indian languages. Ablation studies will measure the contribution of the proposed training components while keeping the benchmark evaluation set separate from training. Experiments are in progress.