The Evaluation Gap: Joint Bounds under Distribution Shift, Selective Labels, and Outcome Measurement Error
Abstract
We study three ways the data used to evaluate a predictive model can differ from the target setting: the regression function the model approximates differs between the evaluation sample and the target (distribution shift), outcomes are recorded only for units given some prior decision (selective labels), and the recorded label is a proxy for the construct of interest (outcome measurement error). We model all three as unobserved confounding between the true outcome and the binary indicator of each obstacle (sampled, labeled, correctly measured), so that the gap between observed and target performance is a sum of three terms with one sensitivity assumption per obstacle, and we compare five methods that apply the assumptions separately or jointly and bound the terms separately or in total. Composing single-obstacle bounds at face value does not yield a bound; composing them validly yields a bound that on some settings adds nothing to the no-assumption bound; joint analysis is never wider than the valid composition and is sharp; assuming the obstacles' latent drivers are independent is found, numerically, to tighten it further on a conjectured region.