Backdoor Attacks Rerouted: BatchNorm as a Sink for Adversarial Signals
Abstract
Deep neural networks (DNNs), particularly CNN-based classification systems, are widely deployed due to their strong performance. However, they remain vulnerable to backdoor attacks, where imperceptible triggers can induce targeted misclassification while preserving high accuracy on clean inputs. These triggers may also distort model explanations. Such vulnerabilities raise serious concerns for both currently deployed and real-world applications, highlighting the need for deeper understanding and training-free defense mechanisms. In this study, we extensively investigate the attacking mechanisms in models with Batch Normalization (BN). We provide the first comprehensive theoretical analysis of the relationship between backdoor attacks and BN, showing that trigger-related information is strongly encoded in BN layers even under full model fine-tuning. We further prove that BN’s affine parameters and running statistics jointly influence both predictions and explanations, offering a unified explanation of backdoor behavior. Building on this insight, we introduce a simple training-free defense that re-estimates batch feature statistics and recomputes normalization at inference time, mitigating backdoor effects while preserving clean performance. Extensive experiments on 8 black-box and 3 explanation-aware attacks, compared against 9 defenses, demonstrate that our method reduces attack success rates from 100\% to 1\%, improves true-class recovery by 89\% (+23\% over prior work), and boosts explanation fidelity by up to 91\%, all without retraining or degrading accuracy. Code will be provided upon acceptance.