Spatio-Frequency Adversarial Attack Is All You Need: Finding Vulnerable Regions in Pathology AI
Abstract
Whole-slide image (WSI) foundation models and computational pathology systems can achieve strong diagnostic performance, yet their robustness to localized perturbations remains poorly understood. In this work, we introduce a spatial–frequency framework for discovering adversarially sensitive regions in pathology images. Rather than perturbing an entire WSI, our approach first identifies regions that are intrinsically sensitive to model decisions using three complementary strategies: Spatial, which captures localized spatial sensitivity; Frequency, which characterizes sensitivity to high-frequency image structure; and Spatial+Frequency, which jointly models both domains. We then perform localized adversarial attacks restricted to the identified regions and quantify the resulting degradation in diagnostic performance. We evaluate the methods under controlled perturbation budgets and cross-node distribution shifts on CAMELYON17-CLEAN, comparing spatial, frequency, joint, and random-region selection using AUROC, AUPRC, and performance degradation. This experimental design enables us to distinguish regions that are merely visually salient from regions that are genuinely consequential to model predictions. Our central hypothesis is that combining spatial and frequency information can identify substantially more influential vulnerability regions than either domain alone, while requiring perturbations over only a small fraction of the image. Beyond measuring adversarial robustness, the proposed framework provides a mechanism for localized failure discovery, revealing where and what type of image information pathology AI systems rely upon when making predictions. These findings aim to establish adversarial sensitivity as a practical tool for evaluating the reliability, interpretability, and clinical deployment readiness of pathology foundation models.