Evidence-Gated Claim Ledgers: A Reporting Standard for Pathology Foundation Models
Abstract
Pathology foundation models underwrite clinical-adjacent decisions across grading, subtype prediction, and biomarker inference, yet the field has no shared surface for recording what a model asserted about an individual slide and what verdict a common standard would assign to that assertion. Current practice reports aggregate metrics and cites confidence scores that carry no shared operational meaning across backbones. We introduce EGCL — Evidence-Gated Claim Ledgers — a lightweight, model-agnostic reporting standard for per-slide pathology-FM outputs. EGCL treats every prediction as a typed Claim, routes it through an Evidence Gate that returns a Verdict from a four-value release-policy vocabulary (SUPPORTED, UNSUPPORTED, NOT_APPLICABLE, CONTESTED), and persists the outcome to an append-oriented Claim Ledger. An Anti-Fabrication Guard refuses to release any downstream artefact whose numeric claims cannot be resolved to a SUPPORTED ledger entry or a declared ledger aggregation. On the public PANDA prostate cohort with the Phikon-v2 backbone we demonstrate an existence proof: on an N=500 cohort evaluated out-of-fold, the verdict vocabulary separates operationally. Under attention pooling (ABMIL) the gate releases 90 slides at SUPPORTED accuracy 0.811, whose 95% Wilson interval [0.718, 0.879] does not overlap CONTESTED [0.412, 0.514]; a simple mean-pool head releases only 9 slides (coverage 1.8%, accuracy 0.778) and is reported as the low-coverage floor. Both orderings of per-verdict subset accuracy run SUPPORTED > CONTESTED > the head's own full-cohort balanced-accuracy base rate > UNSUPPORTED — the ordering a calibrated model should produce, which is the sanity check the gate must pass rather than a gain it creates. The gate is stable across three MIL heads (mean-pool, ABMIL, CLAM-SB). The standard also surfaces a clinically-critical blind spot aggregate metrics hide: the SUPPORTED channel is class-monopolised by benign predictions under both heads — all 9 mean-pool releases and 83 of 90 under ABMIL — so it releases 2 and 15 missed cancers respectively.