Certified Rank-Stability: Risk-Controlling Leaderboard Claims Under Distribution Shift
Abstract
Benchmarks turn complex model behavior into compact leaderboards. A common but rarely tested inference then follows: if model (A) beats model (B) on the clean benchmark, the same comparison is treated as reliable under plausible deployment shifts. This paper studies that inference directly. We define a rank-reversal loss over a declared population of shift units and audit pairwise leaderboard claims under two statistical regimes. When the shift population is a fully enumerated finite grid, reversal risk is computed exactly. When shift units are sampled exchangeably from a larger or implicit population, conformal risk control calibrates an adaptive margin with marginal expected-risk control. We evaluate four object detectors on COCO and four semantic segmentation models on Cityscapes across 15 corruption families and five severities, with ACDC used as an external real-weather diagnostic. Over the exact 75-cell synthetic grid, 12 of 30 clean pairwise claims have reversal risk at most (\alpha=0.10), all on COCO detection; no Cityscapes segmentation claim satisfies this exact-grid criterion. A sampled-shift CRC diagnostic gives nonnegative margins primarily for the same detector comparisons. ACDC further shows that the near-tied DeepLabV3+ versus DeepLabV3 clean ranking reverses on 3 of 4 real adverse-condition units. The contribution is a statistical audit of what a leaderboard claim supports under an explicitly declared shift population.