What a Forecasting Leaderboard Can Establish: Resolution Reporting for Zero-Shot Time-Series Models
Arya Sikder ⋅ Sriram Pillutla
Abstract
A foundation model for time series is pretrained once on a corpus of unrelated series and then asked to forecast series it has never seen. Evidence for a claim of that shape has to be held-out data from many domains. Independent public forecasting datasets are scarce, so the field has concentrated its comparisons on one board. GIFT-Eval, the public leaderboard for zero-shot forecasting, scores every entry on 97 dataset configurations drawn from 28 underlying datasets. It displays a strict order over 130 entries and publishes every per-configuration score. It does not publish how small a difference 28 datasets can detect. We name that quantity the ${resolution}$ of the board and introduce $\textbf{Vernier}$, a report computed from files a board already publishes. Vernier states the resolution, the entries inseparable from the leader, and whether the displayed order changes when unrelated parties submit. GIFT-Eval resolves 6.2 rank places against a median gap of 0.91 between displayed neighbours, so no neighbouring pair in its top ten separates under either rank rule it publishes. Replaying its history, third-party submissions reversed 20 pairs of entries without a re-run. The geometric mean of scaled error, which depends only on an entry's own scores, reversed none. On fev-bench, the main alternative board, removing any single entry reversed no pair in its top ten. Boards also do not record how much history a model is given. Shortening it to 512 steps worsened TimesFM-2.5 and Toto by 22.2\% and 20.6\%, more than the 9.3\% spread of the displayed top ten on the same configurations. We release Vernier, the pinned snapshot, and the disclosure list.
Chat is not available.
Successful Page Load