Recovering Clean Evaluation Metrics from Contaminated Benchmarks
Abstract
When benchmark examples overlap with a model’s training data, benchmark contamination can substantially inflate reported evaluation metrics, yet training provenance is often unavailable. We study post-hoc recovery of uncontaminated evaluation metrics in supervised tabular classification and regression using only prediction–label pairs from a potentially contaminated benchmark. We introduce MILAN (Metric Integrity via Leakage-Aware Normalization), a meta-trained framework that models a contaminated benchmark as a mixture of training-origin and clean evaluation examples and estimates soft posterior weights for the clean component to correct evaluation metrics. MILAN supports a broad class of metrics through weighted estimators and metric-specific correction procedures, and can optionally incorporate contamination-rate priors and sample-wise difficulty scores. To quantify uncertainty, MILAN additionally provides split-conformal prediction intervals with marginal coverage across exchangeable recovery problems. Across a large suite of simulated recovery problems with held-out datasets and model families, MILAN consistently improves clean-metric recovery over thresholding, clustering, mixture-model, anomaly-detection, and membership-inference baselines. In a polygenic risk score case study using UK Biobank evaluation data, MILAN improves agreement with independent FinnGen reference metrics. We provide a scikit-learn-style implementation at \url{hidden-for-submission}.