Why Do Time Series Models Need Long Context Windows?
Luca Butera ⋅ Giovanni De Felice ⋅ Andrea Cini ⋅ Cesare Alippi
Abstract
The effectiveness of modern deep learning models in forecasting *groups of time series* is often attributed to their ability to capture (long-range) dependencies across long observation windows. In this paper, we show that this forecasting task involves two objectives: (i) *generative process identification* (GPI), i.e., inferring the specific process generating the input sequence, and (ii) *conditional forecasting* (CF), i.e., predicting future values given input observations. From this perspective, optimal predictions can be interpreted as an average over plausible data-generating processes, weighted by their likelihood given the input window. This suggests a different explanation for the benefits of long context windows: they reduce the uncertainty about which specific process is generating the input time series during operation. We prove that even for processes with memory length $P$, an input window size strictly larger than $P$ is *necessary* to achieve the minimum attainable error. Finally, we show how decoupling GPI and CF can improve computational scalability without compromising accuracy. Experiments on synthetic and real-world data validate our insights and their relevance for designing forecasting architectures.
Chat is not available.
Successful Page Load