VocalGrad: Evaluating Acoustic Perception in Audio Language Models
Abstract
Large audio language models (LALMs) have recently demonstrated strong capabilities in tasks requiring paralinguistic and non-linguistic understanding of audio. However, they still struggle to reliably utilize such information in tasks such as emotion recognition. Existing benchmarks often entangle acoustic perception with higher-level reasoning, making it difficult to identify the source of these failures. To address this issue, we introduce \textsc{VocalGrad}, a benchmark designed to evaluate the perceptual ability of acoustic features. VocalGrad formulates this problem as a simple binary classification of temporal change direction, enabling evaluation largely independent of higher-level reasoning or semantic understanding. Our benchmarking results show that, despite high accuracy by human annotators, state-of-the-art LALMs perform near chance level, revealing substantial limitations in their acoustic perception ability. We further conduct linear probing to identify structural bottlenecks in these models. Our analysis reveals two distinct failure modes: for some categories, task-relevant information is preserved in internal representations but not utilized by the language model head; for others, the information progressively diminishes across model layers. Overall, our results reveal fundamental limitations of LALMs in modeling acoustic variation and expose representation differences not observable from text outputs.