The Reliability Horizon: A Temporal Stress Test for Forecast Recalibration
Adarsh Agrawal
Abstract
Forecast recalibration can look useful retrospectively and fail when time moves forward. We introduce the reliability horizon, a balanced-panel audit of calibration across forecast lead times coupled to a two-stage deployment rule. On 1,481 resolved Polymarket markets, including 633 observed at every horizon, neither omnibus nor ordered-trend tests detect degradation from one day to six months. Direct horizon equivalence cannot be certified at a strict $\pm 0.02$ ECE margin, while recalibration itself is equivalent to a no-op within that margin at every horizon. Metaculus shows the same no-trend pattern; Kalshi does not. We then evaluate nine language-model forecasters with strict as-of-date gating, per-horizon cross-fitting, and utility guardrails. One retrospective candidate emerges, but its group estimate depends on a single model and reverses on later events. Frozen predictions on 312 resolved outcomes from a 390-market prospective panel produce no reliable gain for any model tier. The main result is therefore procedural: recalibration should be deployed only when its apparent gain survives utility checks, model heterogeneity, and time.
Chat is not available.
Successful Page Load