Change the Metric, Change the Winner: Auditing How We Decide Which Clinical Foundation Model to Deploy in Cancer Pathology
Abstract
Change the metric and the winner changes. We audit an evaluation protocol rather than a model: the recipe the efficient-adaptation literature uses to recommend a fine-tuning strategy for a clinical foundation model, which fixes a benchmark, adapts a backbone under several strategies, reports accuracy over a few seeds, and takes the best. Run faithfully on a metastasis-detection grid in histopathology, five adaptation strategies by three annotation budgets by three seeds, that recipe admits twelve defensible specifications, and every one of the five strategies is selected as best under at least one of them. A published comparison reports a single draw from this table, and nothing in the protocol tells the reader which. Ranking by accuracy and ranking by calibration name different winners in all three regimes, and where labels are scarce the accuracy ranking does not reproduce even itself: redraw the seeds and its winner survives 53% of the time. The obvious rejoinder is that this is an artifact of a poor calibration estimator, and the Brier score answers it uncomfortably. A proper scoring rule ranks with accuracy at every budget and agrees with it on the winner at two of three, so adding one would have concealed the disagreement in exactly the regimes where labels are plentiful, and surfaced it only where they are not. We then turn the same checks on ourselves, and two of the six conclusions our own protocol licensed do not survive: an equivalence claim that was under-powered by construction, and a calibration ranking inside its estimator's noise. Neither needed a new experiment. Both were found in minutes from numbers already reported. And none of this is a standard medicine lacks, since TRIPOD+AI already requires calibration reporting for the regression models these systems are displacing. We propose a claims ledger, recording for each conclusion what its evidence resolves and which population it covers, as what evaluations of medical foundation models should report in place of a leaderboard row.