Rethinking Long-Horizon Forecasting Leaderboards: Most Published Margins Fall Within the Noise of Their Own Test Set
Barry Meng
Abstract
We show that most published long-horizon forecasting margins fall inside the noise of the test set that produced them. The literature keeps a league table, a single mean squared error for each model, dataset and forecast horizon, and reviewers accept papers and practitioners adopt architectures on the third decimal of one entry, the way a race is decided on the margin at the line. Call one entry a cell. How fine a margin can a cell resolve? The benchmark scores each model on windows slid one step at a time, so consecutive forecasts share every target point but one and a test period that looks like thousands of observations holds only as many independent ones as the horizon divides into that period. We call that count the cell's resolving power. One division gives the count, the tables omit the count, and cells the table prints alike split from 100% of their margins certified to 0%. Of 293 published-scale margins the convention certifies 99% and a validated estimator 41%, and held to the cell's own resolution the three flagship tables leave the 2021 Autoformer tied with the model they propose on ETTh1 at $H=720$.
Chat is not available.
Successful Page Load