When Less Is More: Robust Data Compression for Cross-Site Pathology Anomaly Detection
Tushar Shinde
Abstract
Large-scale computational pathology relies on heterogeneous image corpora whose storage, preprocessing, and transfer costs grow with dataset size. We ask whether such corpora can be compressed by two orders of magnitude while preserving cross-site anomaly-detection performance. We introduce \method, a site-aware representation-space coreset strategy that combines proportional acquisition-site coverage, local-density weighting, and within-site diversity. On CAMELYON17-Clean, we evaluate $1\%$ retention ($100\times$ reduction) using leave-one-node-out evaluation, with each held-out node treated as an unseen acquisition domain. \method\ achieves AUROC $0.5913\pm0.1101$ and AUPRC $0.7694\pm0.0914$, compared with $0.5890\pm0.0969$ and $0.7513\pm0.0887$ for the full-data reference, $0.5864\pm0.0998$ and $0.7476\pm0.0923$ for random sampling, and $0.5062\pm0.0988$ and $0.7332\pm0.1029$ for $k$-center selection. Given substantial fold-to-fold variability, we interpret these results as performance preservation rather than statistically established improvement. Moreover, \method\ has a higher FPR95 than both full-data and random references. Overall, the results suggest that carefully structured data compression can preserve aggregate cross-site ranking performance under extreme retention, while revealing important site-level and operating-point trade-offs.
Chat is not available.
Successful Page Load