Hidden Tails: Certifying Tail-Risk Claims under Selective Labels
Abstract
High-stakes prediction systems are often evaluated only on labels revealed by past decisions. This selective observation can hide the losses that determine full-population tail risk: a predictor may appear safe on revealed labels while its worst compatible failures remain unseen. We study certification of conditional value-at-risk (CVaR) under selective labels. First, we prove non-identifiability: even with overlap, two predictors with identical observed selected-loss distributions can have reversed full-population CVaR rankings under compatible hidden-label laws. Under an outcome-dependent odds-ratio sensitivity model, we derive sharp upper and lower CVaR envelopes. A minimax interchange shows that worst-compatible hidden-label completion commutes with the Rockafellar--Uryasev threshold optimization, giving an exact certificate rather than a loose robust surrogate. We then define the tail-risk certification frontier: the minimum passive revelation cost needed to reduce hidden-tail ambiguity below a target tolerance. For a fixed predictor and threshold, this frontier is a fractional-knapsack problem whose value density, tail-identification value, pinpoints labels that can move the CVaR tail. Finally, we give finite-sample guarantees for cross-fitted conservative certificate learning and a lower bound governed by the effective number of revealed tail labels.