The Arbitration Gap Is Mostly Noise: On Complementary Expertise in TSFMs
Abstract
Work on arbitration, routing, and ensembling of time-series foundation models rests on one observation: an oracle picking the best model per window beats every individual model by a wide margin. We show that gap is mostly an artifact of measurement. A null in which the models differ only by noise reproduces almost all of it, and a selector that must commit in advance recovers little. The pool is also less diverse than its provenance suggests, with independently developed models barely less error-correlated than checkpoints of one lineage spanning three orders of magnitude in size. Combination works where selection does not: a faithful reimplementation of a published dynamic arbitrator merely matches a median ensemble that estimates nothing. Textual context, proposed to close the remainder, adds little, and what it adds is covariate information a covariate-capable forecaster could read directly. Arbitration work should report a realizable baseline alongside its oracle.