Opportunity Without Recovery: Routing Time-Series Foundation Models on a Freshly Assembled Cohort
Abstract
Time-series foundation models (TSFMs) often make different forecasts for the same series, creating an apparent opportunity for routing. We evaluate how much of that opportunity a rule can recover on 135 series groups assembled in 2026 from nine US public-agency sources, using five hash-pinned checkpoints and 2,700 forecasts. The cohort was constructed from agency feeds rather than an existing benchmark, but pretraining exposure cannot be ruled out. We separate an outcome oracle's per-group opportunity from the held-out gain of a rule that must choose before observing the target. The oracle improves by 12.81% over a pre-specified TimesFM anchor on 125 of 135 groups. The best single expert selected post-hoc gains 2.33%. Within known domains, a leave-one-group-out domain chooser gains 6.79% (95% interval [4.27, 9.45]), while fixed-weight stacking gains 3.34% under a random group split. On the 12-action pool, a reduced-feature, TimeRouter-inspired comparator gains 2.61% ([0.94, 4.42]) in that split, but its estimate becomes -1.11% ([-3.84, 1.45]) when an entire domain is held out. This exploratory study shows that substantial oracle opportunity need not yield supported cross-domain routing gain.