SWAFT: Sink-Augmented Sliding Windows for Time-Series Foundation Models
Abstract
Long-context time-series foundation models commonly retain quadratic temporal attention or replace it with recurrent and linear mechanisms. We study causal sliding-window attention augmented with persistent sinks. SWAFT is a 1.842M-parameter TimesFM-3-style transformer in which only temporal attention is replaced; scaling, the quantile patch head, and variate attention are retained. It uses a 2,048-step context, 48-step evaluation horizon, length-32 patches, eight learned sinks, and eight recent patches. We extend SarSim0 and CauKer-style generation with multivariate series, covariate roles, and diverse distributional modes. Five 1.83--1.97M-parameter models first receive the same 103k updates and 1.648 million synthetic tasks under final-origin supervision. A SWAFT training ablation instead mixes final-origin and dense per-patch supervision equally. It obtains the best GIFT-short CRPS (0.1406) and TIME CRPS/MASE (0.2683/0.9596); KDA retains the best native FEV SQL/MASE (0.7574/0.8798). Its primary-metric ranks are 1, 2, and 1 across GIFT, FEV, and TIME. Because the baselines were not trained densely, these gains belong to the combined architecture and training recipe.