Can Your Benchmark Tell First from Second? Reporting Resolution in Pathology Model Evaluation
Abstract
Pathology foundation models (PFMs) are increasingly selected for downstream clinical use on the basis of leaderboard position: model A is preferred over model B because it reported a higher area under the receiver operating characteristic curve (AUC) on a benchmark cohort. We ask a question that precedes any comparison of models, and that a benchmark can answer about itself: can this evaluation cohort tell first place from second? We call that property the resolution of a benchmark, the smallest performance difference its cohort can distinguish from sampling noise. We treat the leaderboard as an estimator, measure its sampling distribution by resampling evaluation cohorts, and study resolution as a protocol-level property governed by cohort size, performance gap, prevalence, and the correlation between model predictions on shared cases. We deliberately separate the evaluation protocol from PFM architecture: real-data experiments use fixed prediction pipelines on two pathology-derived cytopathology cohorts (569 and 683 fine needle aspirate cases) with different feature and annotation protocols, and calibrated simulations isolate the same quantities in the regime where PFMs are compared. A 200-case cohort recovers the reference top-ranked pipeline only 41 to 57 percent of the time, orderings separated by less than 0.01 AUC reverse on 21 percent of resampled cohorts, and a paired DeLong test between the top two entries reaches significance in at most 1 percent of draws at any size we study, while coarse discrimination survives: pairs separated by more than 0.05 AUC reverse at most 0.7 percent of the time. Resolution is also being wasted. Model scores on shared cases are strongly correlated (median 0.784), yet comparing per-model intervals rather than a paired interval on the difference inflates interval width by a median factor of 2.01, which in simulation turns 0.78 power into 0.13 at a 0.02 AUC gap and 1000 cases, so benchmarks can extract substantially more resolution from data they already hold. Our deliverable is a four-item Resolution Report, adoptable without collecting a single new case, so that a leaderboard states not only who ranks first, but whether its cohort could tell first from second.