DetectViT: Test-time Backdoor Detection for Vision Transformers via Inter-Head Attention Discrepancy
Siquan Huang ⋅ Yijiang Li ⋅ Xingfu Yan ⋅ Ningzhi Gao ⋅ Leyu Shi ⋅ Ying Gao
Abstract
Vision Transformers (ViTs) have been widely adopted as visual encoders in multimodal models; however, the reliance on third-party pretrained checkpoints exposes these systems to backdoor attacks. Existing test-time detection methods rely on either the predicted label or intermediate embeddings, limiting their applicability across diverse tasks and incurring substantial computational overhead. In this work, we investigate how backdoor triggers affect the multi-head attention mechanism of ViTs and identify two distinct anomalous patterns: $\textit{complete attention hijacking}$, causing the attention to exhibit abnormally high inter-head spatial overlap, and $\textit{partial attention hijacking}$, causing an abnormally large variance in the concentration of the attention distribution. Motivated by these findings, we propose $\textbf{DetectViT}$, a test-time backdoor detection method that requires neither training data, model outputs, nor any prior knowledge of the trigger. DetectViT quantifies the two hijacking phenomena via an inter-head consistency score and an inter-head entropy variance score, with thresholds estimated from a small out-of-distribution (OOD) calibration set drawn from any source. Extensive experiments on four representative attacks, spanning diverse architectures (DeiT, CLIP, LLaVA) and downstream tasks (image classification and captioning), demonstrate that DetectViT substantially outperforms all baselines with as few as 64--128 OOD calibration samples, attaining a true positive rate of $100\%$ and a false positive rate of $0.17\%$ on BadCLIP. Notably, it introduces negligible additional overhead, thanks to its reliance solely on attention weights already computed during inference. We release our code at: https://anonymous.4open.science/r/DetectViT-81AC.
Chat is not available.
Successful Page Load