Auditing Reporting Practices in Cancer Pathology Foundation Model Papers: A Reproducible Rubric and Checklist
Abstract
Foundation models for computational pathology are introduced and compared at a fast pace, but the literature reporting them has not itself been audited for whether these comparisons rest on adequate evaluation practice. We conduct a structured audit of 64 papers that introduce or benchmark a cancer pathology foundation model, drawn from a stated, reproducible multi-angle search over arXiv, medRxiv/bioRxiv, and major venues (NeurIPS, CVPR, MICCAI, Nature-family journals) spanning 2022-2026. Each paper is scored against a 7-item rubric fixed before any paper was read, covering external validation, unit-of-analysis leakage checks, scanner/stain/site confound disclosure, label provenance and review, variance reporting, prospective framing, and code/weight availability; a blind independent re-scoring of a random 21-paper subsample gives 85.7% exact agreement and quadratic-weighted kappa=0.86. External validation and variance reporting are more common than expected (67% [95% CI 55-77] and 55% [43-66] of papers fully meet them), but two practices are rare field-wide: only 5% [2-13] document a labeling protocol with pathologist review or adjudication, and only 2% [0-8] include any prospective or externally time-split validation, with 89% [79-95] neither performing nor flagging the absence of one. Both gaps survive leave-one-group-out and access-level sensitivity analyses. Two further results bear on what happens next: adherence is statistically flat between 2022-2024 and 2025-2026 (no criterion improves after correction for multiple comparisons), and peer-reviewed papers score no better than preprints - so neither time nor peer review is currently closing these gaps. We release the rubric, the full per-paper coding, and a minimal reporting checklist as concrete artifacts for authors and reviewers.