Ranking Is Not a Deployment Policy: Coverage Drift in Learned Reliability Gates for Prediction-Market Inputs
David Lewis ⋅ Enrique Zueco ⋅ Ling Yi
Abstract
Ranking is not a deployment policy: prediction-market prices may aggregate information, but they are also emissions from a time-varying social channel. We test whether a learned pre-outcome loss gate adds practically useful selective accuracy beyond simple and linear selectors. A frozen synthetic stress test uses 48 training streams, 24 threshold-calibration streams, and 384 confirmation streams (267,264 eligible events) spanning six temporal channel mechanisms. At nominal 80\% coverage (556 of 696 events per stream), the nonlinear gate attains Brier loss 0.18746 versus 0.18849 for the strongest pooled operational baseline, a paired difference of $-0.00103$ (95\% stratified stream-bootstrap $[-0.00127,-0.00078]$). The gain is resolved but fails the predeclared $-0.002$ practical-effect gate. Moreover, a threshold calibrated for 80\% coverage realizes 77.8--91.4\% mean coverage across mechanisms, and retained-set calibration slopes range from 0.773 to 0.930. Thus, small risk-ranking gains do not establish a robust deployment policy. The experiment is synthetic: it validates an evaluation contract, not real-market performance or a new foundation model.
Chat is not available.
Successful Page Load