When to Call a Time-Series Foundation Model: Metric-Dependent Routing and Long-Horizon Limits
Teeratorn Kadeethum ⋅ Vignesh Srinivasan ⋅ Amin Shojaeighadikolaei
Abstract
For hourly electricity load in 19 European countries, two zero-shot foundation-model (FM) services cost much more to serve than a structural forecaster whose fitted coefficients fit in 13\,kB, and neither FM is more accurate everywhere. We therefore propose a router that decides, request by request, whether an FM call is worth its cost. The usual measure of what such a router could gain is the \emph{oracle gap}: the loss a fixed model choice incurs whenever the other model would have won. We show that the gap equals $(\mathbb{E}|D-\lambda| - |\mathbb{E}[D]-\lambda|)/2$, where $D$ is the FM's per-request advantage and $\lambda$ is the call cost. The gap is positive whenever the winner changes from request to request, whether or not that change is predictable: two datasets can have the same gap when the winner is predictable from the request in one and a coin flip in the other. The gap therefore measures how often a clairvoyant would switch, and it only upper-bounds what a router can actually recover. On a frozen 2025 window of 27,735 requests, with pretraining overlap unknown, three regimes emerge. (1) At one-hour and one-day horizons, the FM is the default, routing reduces the number of calls, and a post-release quarter from 2026 reproduces the gains. (2) At the one-week and one-month horizons, the structural model is the default, routing adds calls, and the 2026 quarter does not reproduce the gain. (3) At one- and two-year horizons, cap-tuned Chronos-2 trails seasonal naive. Within the two routed regimes, a frozen router beats the best static policy for each (country, horizon) in all 8 cells, by a median of 0.19 percentage points (pp). Scoring the same decisions by panel-median MdAPE instead of pooled mean APE still yields a 0.10\,pp gain. Letting each metric choose its own threshold shifts the deployed gain by 0.10\,pp; the interval that accounts for threshold selection is $[0.011, 0.176]$, and only one cell is significant on its own.
Chat is not available.
Successful Page Load