The Unused Input: The Initial State as a Conditioning Interface for Recurrent TSFMs
Abstract
Time series foundation models (TSFMs) forecast unseen series without task-specific training. Increasingly they are built on recurrent backbones, whose constant per-step cost makes long-context inference cheap, and which, unlike transformers, carry a state. A recurrent backbone computes its prediction from two inputs: the observed context window and the initial recurrent state. In practice only the first is utilised and the initial state is unused by convention. We ask whether the initial state can be utilised in recurrent TSFMs. Using the xLSTM-based zero-shot forecaster TiRex as a case study, we freeze the backbone and train only a small head that writes a context vector into the initial state, so that any change in accuracy is attributable to the injected content rather than to adapted weights. A spectral summary of the input window, learned once and applied unchanged to every series, yields a small but consistent improvement on GiftEval that neither a parameter-matched LoRA adapter nor a learnable static state reproduces. An embedding of text attached to each series, evaluated on six MoTime datasets, reveals two roles for the state: as a per-dataset adaptation surface, where a static state recovers part of what full fine-tuning achieves, and as a channel for exogenous information, where the text-conditioned state improves on the static state on four of six datasets and, combined with a LoRA adapter, can exceed full fine-tuning, which has no access to the text. The effect concentrates in the mLSTM blocks and is strongest at short context. These results are preliminary and the improvements modest, but they are consistent, and they establish the initial state as a conditioning interface that recurrent forecasters have so far left unused.