HistoFID: Calibrating Fréchet-Distance Evaluation across Pathology Foundation Models
Abstract
The Fréchet distance (FD), the basis of the Fréchet Inception Distance, is increasingly computed using pathology foundation models instead of Inception to evaluate synthetic tiles, cohort or scanner drift, and image compression, but this substitution makes reported distances non-comparable across studies. We show that the within-cohort FD floor for an identical comparison varies by approximately 15–31× across nine commonly used patch encoders (Inception-v3, CONCH, Phikon-v2, UNI2-h, Virchow2, Prov-GigaPath, Lunit-DINO, Midnight, and Hibou-L), independent of embedding dimension, making raw FD values uninterpretable without specifying the encoder. Benchmarking all nine encoders on an internal cohort of ~500,000 H&E tiles and six external cohorts spanning multiple organs, institutions, scanners, and controlled stain perturbations, we find that expressing FD as a ratio to each encoder's own within-cohort floor reduces the pooled across-encoder coefficient of variation from 0.88 to 0.40, outperforming per-dimension standardization (0.55), trace normalization (0.49), and KID (2.43). The remaining variation is structured rather than noise: under both cohort and stain shifts, encoders separate into a sensitivity group (CONCH, Phikon-v2, Inception-v3, Lunit-DINO, Midnight, and Hibou-L) and an invariant group comprising the large DINOv2 models (UNI2-h, Virchow2, and Prov-GigaPath), which are consistently the three lowest-ratio encoders under every nuisance shift (exact permutation test, p = 0.012), providing a practical basis for selecting encoders that either detect or suppress nuisance variation. The floor ratio also serves as a feature-space fidelity metric for lossy compression, revealing feature drift missed by pixel-based measures: at matched pixel fidelity, a proprietary codec remains below the within-slide floor on every encoder, whereas JPEG, JPEG2000, and JPEG XL exceed it on the most compression-sensitive encoder. Finally, the calibrated metric aligns closely with expert perception, with feature-distance ordering matching similarity judgments from four board-certified pathologists (100 tiles; Fleiss' κ = 0.98; Spearman ρ = 0.96), while preserving this ordering through its strictly monotonic normalization. We therefore recommend reporting pathology FD as the floor-normalized ratio at a fixed sample size to obtain a comparable, human-aligned metric for clinical-scale evaluation, and we release the accompanying code.