A Theoretical Analysis of Backdoor Learning as Simplicity-Biased Optimization Dynamics
Abstract
Backdoor attacks implant a trigger-target association into a model, causing malicious behavior at test time while largely preserving clean performance. Despite extensive empirical study, a unified explanation for why standard training dynamics learn backdoors so effectively remains largely missing. We argue that backdoor learning can be understood as a consequence of simplicity-biased optimization dynamics: trigger features are dynamically simpler than semantic features under stochastic gradient descent (SGD) because they induce more coherent gradient alignment, more stable gating behavior, and stronger early-stage amplification. To formalize this perspective, we introduce a directional data model that separates semantic and trigger features and study a one-hidden-layer ReLU network trained by SGD. Our analysis identifies a frozen-gate signal that governs group-wise early-stage loss decrease up to controlled gate-drift and gate-flip errors, yielding a quantitative explanation for faster poisoned-sample fitting and its dependence on the poisoning ratio. Beyond optimization dynamics, we show that this early-stage bias can induce trigger-dominant neurons with selective activation on poisoned inputs, explaining why vanilla clean fine-tuning can attenuate backdoor behavior without necessarily erasing the underlying trigger-related representation. Empirically, we validate the predicted signatures of early-stage optimization bias, representation-level trigger dominance, and post-training residual behavior using ResNet-18 on standard image classification benchmarks.