Selection-Induced Optimism Corrupts Calibration, Not Just Discrimination: A Code-Verified Audit of Hidden-State Hallucination-Detection Pipelines
Abstract
When a checkpoint, layer, threshold, or hyperparameter is selected on data later reused to report performance, the reported number is optimistic. This is well documented for discrimination metrics such as AUROC. We show the same selection step also corrupts calibration – how reliable a system's stated confidence is, not just how well it ranks – and that this breaks a guarantee practitioners actually rely on. Selecting one layer by argmax over 32 candidates – with no fitting on the evaluation data at all – inflates the error budget a selective-prediction threshold is believed to buy by 10–16% relative. At 70% coverage over 100,000 queries a day, a believed risk of 0.128 against an achieved 0.142 is 980 additional unflagged errors daily. We establish the underlying calibration gap two independent ways: on two structurally different selection mechanisms in our own harnesses (Brier gaps +0.008 to +0.026, ECE gaps up to +0.019, bootstrap-confirmed, with nine of the ten cells established per-cell – six of them cleanly, three resting on a non-negativity our own method section disqualifies – and all nine surviving joint family-wise correction), and as a real bug in an externally published, pinned-commit-verified pipeline's own reported ECE – evidence from a source we did not build. A regularization sweep stress-tests this: the Brier gap survives but shrinks, the ECE gap does not survive, and the risk violation is unchanged or larger – we report all three. Measured severity is a property of a (mechanism, metric, control) triple, not of a mechanism alone: one case-study cell spans 1.44× across two admissible controls. We release a five-mechanism leakage taxonomy, three mechanisms verified against pinned third-party source code, and a checklist extending Kapoor & Narayanan and REFORMS – plus a negative result: a regex-based scanner built to automate four of its five items finds zero true positives across seven repositories and misses both leaks it was built to catch. Our object of study is the evaluation protocol, not the model.