What Do Audio Models Really Hear? Layer Selection and Mechanistic Structure in Sound Representations
Abstract
Layer-wise depth dependence is known for speech SSL, but it remains unclear how readout depth changes across audio pretraining paradigms and how that choice can support pruning rather than only post hoc probing. We present the Audio Interpretability Atlas, a readout-centered diagnostic study over a wide variety of frozen encoders and complementary sound, music, and speech benchmarks. The atlas connects layer wise transfer, CKA, geometry priors, classical descriptors, sparse-feature concentration, transcoder routing, robustness, steering, and targeted feature ablation on the same encoder-layer-task cells. Its operational goal is to identify the earliest layer or prefix that preserves task evidence, so later blocks can be treated as pruning candidates and, when labels are available, learned layer weighting can be restricted to the selected prefix. A consistent pretraining-associated pattern emerges: audio-text encoders expose compact upper-stage category features, ASR-supervised encoders route speech information densely into mid-to-late layers, and masked or denoising SSL exposes reusable acoustic structure earlier. Final-layer extraction loses at least 10 score points in about half of evaluated encoder-task settings, with the largest single-layer gaps reaching 24--38 points; speech-affect readouts, by contrast, often justify later depth. In low-resource ASR, \textsc{QuickLayer}'s label-free geometry mode, based on isotropy and participation ratio, improves 11 of 12 language-encoder settings, averaging 21.3\% relative CER reduction, while few-shot probes recover 11--13 points over final-layer extraction. Sparse features, transcoders, perturbations, and interventions then turn selected layers into audit records, identifying when a pruned readout is accurate, stable, localized, and editable.