Natural Jammers: When Non-Robust Features Become Antagonistic
Abstract
Deep neural networks are known to rely on non-robust features that align poorly with human perception. Prior work has largely framed these features as opportunistic shortcuts that help standard models on clean data. We show instead that non-robust cues can be antagonistic: they can act as Natural Jammers that actively mislead standard models on clean, unmodified images. Using disagreement between a standard model and robust models as an analysis lens, we identify a Jammed Set on which robust models outperform standard models. We provide causal evidence for Spectral--spatial Locking: jamming arises only when high-frequency components are precisely aligned with specific spatial structures. Disrupting this lock through spectral filtering or small geometric shifts can restore correct predictions. Crucially, we uncover a Substitution Paradox: while individual jammers are fragile to geometric shifts, the aggregate prevalence of jamming remains stable and is redistributed rather than eliminated under such transformations. Our results reframe natural non-robustness as a structural property of the model-data manifold: one that redistributes rather than disappears under simple interventions.