Time Flies When You Parallelise: Pareto-efficient Time Series Foundation Models with Linear Recurrences
Abstract
Time-series foundation models have demonstrated strong zero-shot forecasting performance, but computational efficiency remains a central architectural challenge. Temporal self-attention incurs quadratic cost in the number of temporal tokens, while nonlinear recurrent state transitions restrict parallel computation along the temporal dimension. We introduce CI-Proto, a family of time-series foundation models based on Mamba-3, whose linear state recurrences support chunk-wise parallel-in-time computation during training and inference costs that scale linearly with sequence length. The architecture alternates recurrent temporal processing with cross-variate attention, supporting multivariate probabilistic forecasting with past-only and future-known covariates. Across the 100 forecasting tasks in fev-bench, all four CI-Proto variants lie on the Pareto frontier of mean scaled quantile loss versus both mean active parameter count and median runtime, outperforming several substantially larger Transformer and nonlinear recurrent baselines.