AcousticBench: Measuring Acoustic Perception in Large Audio Language Models
Abstract
Large Audio Language Models (LALMs) are designed to leverage non-lexical features in audio that Automatic Speech Recognition (ASR)-based cascade systems discard. Yet, existing LALM benchmarks often reward improvements in both language and audio understanding, making it difficult to isolate performance on acoustic perception. We introduce AcousticBench, a benchmark of 3,200 two-choice questions targeting non-lexical audio properties across three task families: Relative Acoustic Discrimination (RAD), Speech Emotion Recognition (SER), and Relative Speech Quality (RSQ). The benchmark draws from 5,800 unique recordings totaling 11.6 hours across three audio collection protocols: locally captured single-speaker voice-acting sessions, WebRTC multi-speaker conversations in 17 languages, and synthetic sound generation. We evaluate 15 LALMs alongside an ASR-based cascade and a held-out human reference. The cascade system performs at near-random levels across task types, confirming AcousticBench cannot be solved from transcripts alone. Gemini-3.1-Pro leads on every aggregate task but no model reaches human performance; open-weight models lag most on relative differences in loudness, dynamic range, and channel degradation. As a diagnostic, we freeze Kimi-Audio's audio encoder and fine-tune its backbone with task-specific LoRA adapters. We find that RAD lifts from 69.2% to 98.7%, RSQ from 54.8% to 98.4%, and pairwise SER from 72.3% to 92.5% on their held-out evaluation sets, with only modest gains on single-clip SER. This suggests that, for Kimi-Audio, targeted post-training can substantially close the gap to human acoustic perception even with the audio encoder held fixed.