Auditing Attention Head Masking for Out-of-Distribution Detection: Cross-Architecture Wins, Failures, and Polarity Inversions
Vishnu Bindu Balachandran
Abstract
Post-hoc scoring functions for out-of-distribution (OOD) detection in transformer models are typically validated on a single architecture and a single modality, leaving cross-architecture behavior largely uncharted. We report a cross-modality empirical study of attention-head-masking-based OOD detection across more than 1,200 AUROC measurements, spanning LayoutLMv3 on document classification, ViT-B/16 on natural images, and RoBERTa-base on text. Each setting is evaluated against seventeen baselines under a uniform protocol with five fixed seeds, including the Attention Head Sensitivity Scoring (AHSS) and AHM-augmented distance methods of Constantinou et al. (2024) and the recent X-Mahalanobis from NeurIPS 2025. Paired statistical tests are run on every head-to-head claim. Within the AHM family, mahAvgAvg_AHM significantly beats X-Mahalanobis on Tobacco-Cross (Wilcoxon $p=0.031$) and on RoBERTa SST-2$\to$20-Newsgroups ($p=0.031$); on standard CIFAR vision benchmarks, X-Mahalanobis remains state-of-the-art. Three architecture-dependent phenomena surface that current evaluation practice tends to miss: Monte Carlo Dropout collapses to chance-level AUROC across all four ViT settings tested; the AHSS displacement score inverts polarity on RoBERTa, yielding a mean AUROC of 0.296 even though it functions on vision and document transformers; and on near-OOD splits where ID and OOD share visual or semantic structure, every method including the strongest baselines degrades by 25 to 42 percentage points relative to standard OOD. We release per-seed measurements, AHM masks, and code to support replication.
Chat is not available.
Successful Page Load