Into the ORBIT for Time Series: Training Regimes for Foundation Models
Abstract
Time series foundation models are usually compared by architecture and parameter count, although the data pipeline determines the forecasting problems actually seen during pre-training. We present a compact study of two coupled design choices. First, Falcon-2.0 is a 585M-parameter encoder-only forecaster with missingness-aware triple-channel patches, parallel future-query prediction, and direct quantile outputs. Second, ORBIT (Omni-Range Bootstrap Incremental Training) makes the effective pre-training distribution explicit: domain-aware source quotas are materialized as a low-discrepancy stream, and each example is indexed by a sampled record, variable, context window, and feasible horizon. Variable context-horizon pairs are then interleaved throughout one step-based training run. We conduct extensive experiments on GIFT-Eval and fev-bench to demonstrate the effectiveness of the proposed model. Controlled ablations show that replacing stochastic construction with sliding windows consistently degrades both point and probabilistic forecasting performance. We also identify an operational boundary: the gains reverse when future covariates are available but cannot be consumed by the univariate interface. These results highlight the importance of jointly considering model and training-distribution design.