From Patches to Whole-Slide Images: Robustness Analysis of Pathology Foundation Models
Abstract
Pathology foundation models have achieved strong performance across histopathology tasks, but reliable clinical use requires robustness beyond clean test data. Robustness analysis is also relevant to model development because self-supervised training often relies on perturbation-based invariance, meaning poorly tolerated transformations may reveal weaknesses in learned representations and inform future augmentation choices. A combined evaluation was conducted at both patch and whole-slide image levels to examine how pathology foundation models respond to image degradation, distribution shift and spatially heterogeneous corruption. At patch level, twelve pathology foundation models and ResNet baselines were evaluated across clinically motivated perturbations and deliberately dissimilar data splits, revealing substantial variation in robustness between model families and showing that larger model size alone does not guarantee stronger robustness or generalisation. Whole-slide robustness was then evaluated for oestrogen receptor (ER) and progesterone receptor (PR) prediction using 1,006 diagnostic H&E whole-slide images from TCGA-BRCA. Six frozen encoders, Virchow2, Phikon, Phikon-v2, Kaiko ViT-B/8, RetCCL and ResNet18, were combined with target-specific Attention-MIL and evaluated using five-fold patient-level cross-validation. Classification heads were trained on clean data and kept fixed during robustness testing so that changes in performance reflected sensitivity to image corruption. Virchow2 achieved the strongest clean whole-slide performance, with ER AUROC of 0.908 ± 0.029 and PR AUROC of 0.807 ± 0.034, although clean performance did not consistently correspond to greater robustness. Across perturbation sweeps, mild degradation was generally well tolerated before performance declined more sharply at higher severities, with the onset and magnitude of deterioration varying across encoders and perturbation types. Gaussian blur showed little deterioration at lower tested severities followed by substantially larger losses at stronger settings, while stain perturbations were generally stable at low severities before deterioration became apparent from approximately a one-stop concentration shift. Spatial analysis further showed that whole-slide robustness depended on both the extent and spatial organisation of corruption. At 75% coverage, spatially clustered corruption produced lower AUROC than randomly distributed corruption in 31 of 36 comparisons across six encoders, three perturbation types and two prediction targets. Joint severity-and-extent analysis also showed mean harmful decision-flip rates increasing from approximately 5.7% at 25% coverage to 31.7% at full coverage. Together, these findings show that pathology foundation model robustness depends on model design, perturbation severity, spatial extent and data distribution, while systematic perturbation analysis may inform both model evaluation and future self-supervised training.