Adversarial Attack and Defense for Machine Learning in Statistical Physics
Abstract
Recently, various machine learning (ML) systems have been proposed for physical science. Their powerful mathematical expressiveness offers an inherent advantage in maintaining the desirable statistical properties of physics. However, they are placed in ideal and virtual environments that seldom consider noise, perturbations, or various real-world experimental limitations. This exposes significant vulnerabilities in several aspects: 1) The input of an ML system can be subject to human or natural, deliberate or unconscious perturbations, which may sabotage the running baseline of the ML system. 2) A small input perturbation may be amplified by the maximum physical sensitivity, which may lead to an extreme output or a divergent loss function value. 3) A perturbed input may violate some strict statistical assumptions that the ML system relies on, which may lead to incorrect scientific results and findings. In this work, we propose a complete attack-defense methodology for ML systems in statistical physics. In the attack side, we launch a black-box attack (such as a small Gaussian noise) at the input, which stays within physically plausible limits and preserves physical properties. In the defense side, we develop a functional local variation regularized learning scheme to capture and suppress the adversarial perturbation. Theoretical analysis and extensive experiments show the effectiveness of the proposed method. This finding may shed new light on the systemic vulnerability and scientific security of such ML systems.