When Forecast Stability Changes Decisions: Evaluating Time-Series Foundation Models under Intermittent Demand
Jiekai Wang ⋅ Jiamei Meng ⋅ Ziwei Jiang ⋅ Adam Gilbert
Abstract
Forecast stability is often treated as a single reliability property of a time-series foundation model. On mostly-zero demand, however, rankings depend on what counts as a material revision, which forecast quantile is tracked, and whether sampled forecasts are reproducible. We evaluate six open-weight probabilistic TSFMs against intermittent-demand baselines in a paired sparsity experiment, with external checks on three public demand panels. Exact-change prevalence mainly records numerical identity. Percentage normalisation amplifies small movements near zero, and quantile rankings shift around the moving zero mass. On Online Retail, Toto 2.0 ranks first by common-support predictive accuracy but 16th by mean Symmetric Forecast Percentage Change (sFPC). Identical-context controls show that sampled plan disagreement can be as large as observed cross-origin change, so cost rankings use reproducible deterministic plans. At occurrence probabilities $p=.2/.3$, all five deterministic TSFM plans reduce free-replanning demand cost relative to the always-zero plan. Six of ten TSFM–Croston-SBA comparisons significantly favour the TSFM when replanning is free. When each plan change has fixed cost $\kappa=.25$, all six change direction and all ten significantly favour Croston-SBA. Public panels confirm that predictive accuracy and revision behaviour need not align, and stockout controls show that target provenance changes accuracy evaluation. We therefore recommend evaluating predictive quality and revision behaviour jointly, with native support, inference reproducibility, and target provenance explicit.
Chat is not available.
Successful Page Load