When the Label Measures the Estimator: Threshold Leakage in Estimate-Scaled Evaluation Labels
Abstract
Evaluation labels are often built from an estimate. In financial machine learning a common construction is the tail-event label: a stock's forward return counts as an event when it exceeds a multiple k of that stock's estimated volatility. Conditional on the estimation error, the stock's own volatility drops out of the crossing probability and the error is what remains, so any score correlated with that error earns apparent ranking skill, measured here by AUC (the area under the receiver operating characteristic curve). We call this threshold leakage. We derive the leaked AUC in closed form from estimator noise and bias, event rarity and window overlap; validate it in a simulated market with nothing to predict; and stress-test the label on a panel of U.S. equities under a protocol registered before a single held-out pass. The score under test is the ratio of a stock's three-year to its nested sixty-day volatility estimate; the latter also sets the threshold. Replacing the volatility estimator in both the label and the score removes between a quarter and more than half of the score's AUC excess above chance on the thousand most liquid stocks, a lower bound on the leakage because the swap cancels whatever artifact the estimators share. The same estimator bias makes a volatility-targeted loss limit, applied without recalibration, breach at nearly twice its designed rate under three of the five estimators. On real prices the closed form reproduces the direction and shape of each effect but not its level. A diagnostic control, an oracle that re-scores the same stock-days against a threshold built from the label window's own realized volatility, removes the artifact in sample where the score's two windows overlap heavily and leaves about 0.04 of AUC where they barely overlap, against a residual near 0.06 that the closed form leaves unexplained; we report the pair as a floor and a ceiling on real skill.