Model Capacity Dominates Language Resourcing in Multilingual ASR Quantization
Abstract
Multilingual Automatic Speech Recognition (ASR) accuracy is correlated to how much pretraining data a language had. A previous study reports that quantization further compounds the gap between high- and low-resource languages. We show that this conclusion is a property of how degradation was measured, an absolute threshold on the error increase and we ask what governs quantization fragility instead. On 18 ASR and audio-language models, we evaluate the effect of post-training quantization (HQQ at 8, 6, 4, and 3 bits) against the BFloat16 baseline across 25 European languages. We find that the amount of pretraining data has little explanatory power for the quantization penalty, while model capacity is significantly more important. Across the Whisper scaling ladder, the relative CER penalty at 4 bits falls from 129.3% for Whisper-T to 2.8% for Whisper-Lv3. This capacity effect becomes stronger as the bit-width decreases. Overall, model capacity has a greater impact on quantization robustness than the choice of HQQ bit-width, particularly for the larger models.