Clustered Misses Are Not Miscalibration: A Controlled Audit of TSFM Prediction Intervals
Abstract
Reliable uncertainty estimates for time-series foundation models require distinguishing clustered forecast misses from miscalibration. We examine whether clustering diagnoses miscalibration and whether interval corrections improve quality under overlapping horizons and delayed outcomes. We audit 80\% prediction intervals across 332 monthly series using a calibrated Gaussian oracle, a factorial experiment separating feedback timing from fitting-data availability, and a static calibration baseline. Increasing overlap raises the oracle's median worst-window coverage deficit relative to overall coverage from 0.150 to 0.339 despite exact 80\% conditional coverage. At a 12-month horizon, delayed feedback increases this deficit by median paired differences of 0.111--0.170 for Chronos-Bolt and Chronos-2 on tourism and M4, with fitted parameters fixed. Under feasible timing, the GARCH volatility correction reduces Chronos-2's deficit on M4 by a paired median of 0.060 at the default window but worsens normalized interval score by 0.040 versus the static baseline. These findings show that neither clustered misses nor reduced clustering alone establishes interval quality. Our audit provides a protocol for evaluating temporal foundation models through joint coverage, width, and score comparisons under explicit feedback constraints.