Input- and Feature-Space Correction for Cross-Hospital IHC Grading
Abstract
Pathology foundation models (FMs) are trained predominantly on hematoxylin and eosin (H&E) images but are increasingly used as frozen feature extractors for immunohistochemistry (IHC), where cross-hospital differences can impair grading. However, systematic comparisons of correction methods for cross-hospital IHC grading remain limited. To address this gap, we benchmark label-free transformations applied to input images before frozen-FM encoding and to features after frozen-FM encoding using UNI and Virchow2, three IHC markers, and six directed cross-hospital transfers, yielding 12 settings. Among the evaluated input-space methods, our proposed method FIX achieved the largest macro-F1 improvement over no correction, outperforming well-known stain normalization techniques. Among feature-space methods, a simple feature-wise standardization achieved the best improvement over no correction, outperforming CORAL. Under patient-averaged tile-level evaluation, combining FIX with feature-wise standardization improved macro-F1 and quadratic weighted kappa (QWK) by 0.181 and 0.208, respectively, and reduced ordinal error by 0.209 relative to no correction. These results show that input- and feature-space transformations can provide complementary benefits for cross-hospital IHC grading with frozen pathology FMs.