Calibration and Discrimination in Time-Series Foundation Models
Lucas Fugate
Abstract
Time-series foundation models are predominantly ranked by a single aggregate score: mean CRPS, or the closely related WQL. These scores indicate which forecaster scores better, not why. A model can win because its distributions track the outcome more closely, or merely because it states its uncertainty more honestly, and the two call for different fixes. We separate them by decomposing CRPS into miscalibration, discrimination, and uncertainty across five open-weight TSFMs and a seasonal-naive baseline on four GIFT-Eval configurations. Models that are close on CRPS can fail in very different places. On M4-weekly, TimesFM-2.5 and Chronos-Bolt show no detectable difference in discrimination yet differ by $0.12$ CRPS, a gap our decomposition assigns almost entirely to calibration, while a CRPS-only ranking separates them by three places. Whether such a gap is worth acting on varies by dataset: recalibrating out of fold closes $74\%$ of the seasonal-naive baseline's deficit on solar/H, so most of its apparent disadvantage is correctable rather than missing information, while on the other three configurations the same procedure leaves every foundation model slightly worse. Aggregate CRPS cannot tell these cases apart. We release a tool that computes the decomposition from any quantile-forecast file, and report the approximations our benchmark-scale implementation requires.
Chat is not available.
Successful Page Load