Dialect Gaps and Lexical Biasing in Multilingual Speech Models: A Paired Evaluation of Erode and Common Spoken Tamil
Abstract
Speech-recognition systems for Tamil are usually evaluated with aggregated scores, so their performance on regional speech is not always clear. We present a paired audit of automatic speech recognition (ASR) performance on Erode (Kongu) Tamil and common spoken Tamil using recordings of the same everyday meanings. Five native Erode Tamil speakers each recorded 45 meaning-matched pairs and five identical-text control pairs, producing 500 clips with speaker and intended meaning matched across forms. We evaluated four multilingual speech models: Whisper large-v3, Gemini 3 Flash Preview, Sarvam saaras:v3, and AI4Bharat IndicConformer-600M using character error rate (CER) and human judgments of meaning preservation and dialect-word retention. The observed gap varied sharply by system: Gemini's CER increased from 9.3\% on common spoken Tamil to 18.7\% on Erode Tamil, while Sarvam and IndicConformer showed smaller increases from 7.9\% to 10.5\% and 9.6\% to 12.7\%, respectively. Whisper remained near 19.0\% in both forms but preserved meaning in only about half of the sentences in either form. We then tested lexical biasing by supplying a list of 10 Erode Tamil words as inference-time context and measured whether it improved dialect-word recognition. For Gemini, listed-word retention increased from 60\% to 100\%, while retention across all evaluated Erode-dialect words increased from 58\% to 71\%. However, lexical biasing also led to unsupported insertions of listed words that were not spoken, at a rate of 0.10 per clip, and on 15 silent room-noise probes, every Gemini transcript contained words from the supplied list. Our results show that dialect gaps, the benefit of lexical biasing, and its insertion cost all depend on the system, and that these three quantities need to be measured separately.