Structure Is Not Mechanism: Auditing Shortcut Features in Medical Vision Models
Abstract
Interpretability methods increasingly reveal coherent structure in learned representations, but it remains unclear when such structure provides evidence for an underlying mechanism. We present a mechanistic auditing framework that moves from discovered structure to independently tested mechanistic interpretations using causal, semantic, and acquisition-based evidence. On the CovidQU-Ex COVID-19 chest X-ray dataset, we identify a non-pathological but causally dominant (NPC) feature cluster. It comprises only 11\% of the learned feature dictionary yet exhibits 3.6-fold greater mean causal influence than pathology-aligned features. Despite minimal spatial overlap with infection, further analyses link the cluster to acquisition cues such as brightness and GGO-like textures and reveal competition with pathology-related representations. Across four architectures and four datasets, related structural patterns recur but independent evidence supports different functional interpretations across settings. Together, these results show that recurring internal structure can motivate a mechanistic explanation without establishing one. More broadly, they suggest a role for interpretability in turning observations about model internals into testable accounts of how models compute.