Learning Recoverable Neural Networks against Weight Corruption via Simple Zero-Sum Projection
Abstract
Deep neural networks are highly vulnerable to bit corruption in stored weights, whether from naturally occurring memory faults or deliberate fault-injection attacks. Existing defenses remain fundamentally limited: fault-tolerance methods degrade under cumulative corruption and remain vulnerable to strong white-box attacks, while ECC-based recovery methods are often tied to specific precisions or architectures and expose concentrated vulnerable surfaces. We propose a simple training-based recovery framework that endows model weights with recoverable structure through block-wise zero-sum constraints enforced by differentiable projection. The resulting method incurs no parameter-space overhead and generalizes naturally across architectures and numerical precisions. We further develop adaptive adversarial bit-flip attacks tailored to recovery-based defenses, covering both our method and prior ECC-based approaches. Across diverse architectures, precisions, datasets, and threat models, our method consistently delivers strong robustness and recoverability while preserving clean performance.