Native Sparse Time Attention for High-Dimensional Multivariate Forecasting
Balthazar Courvoisier ⋅ tristan cazenave
Abstract
Scaling attention-based forecasting models to high-dimensional multivariate time series often comes at the cost of cross-channel interactions. Existing efficient architectures avoid the quadratic cost of dense attention by restricting communication to compressed representations or removing multivariate interactions altogether. We introduce Native Sparse Time Attention, which enables fine-grained interactions while scaling linearly in the number of channels for fixed sparsity budgets. This efficiency exposes a tunable accuracy–compute frontier and enables longer contexts, which can improve forecasting performance under memory constraints.
Chat is not available.
Successful Page Load