WILD: Widely Linear Conditioning for Time Series Forecasting
Abstract
Transformer-based time series forecasting models often suffer from block-wise attention collapse, where temporal smoothness and cross-variable correlations concentrate attention locally and restrict global modeling capacity. We argue that this collapse is partly a pre-attention representation-conditioning problem: attention logits inherit the spectral concentration and low effective rank of the input representation. We propose Widely Linear Frequency-Domain (WILD) conditioning, a lightweight, model-agnostic module inserted before self-attention. WILD maps inputs to the frequency domain, applies data-dependent spectral preconditioning to redistribute over-dominant frequency energy, and then uses asymmetric real/imaginary projections to realize a widely linear transform beyond fixed time-domain channel mixing. The preconditioning step is essential: it constructs the usable normalized frequency basis on which the widely linear branch can operate. Thus, WILD uses frequency-domain modeling to condition an existing attention layer, not to replace it with a new forecasting architecture. Extensive experiments on eight benchmark datasets show positive gains on 105 of 112 model-dataset-metric configurations across seven backbones. Attention diagnostics show that WILD substantially increases effective rank over the baseline, and a Wr=Wi ablation identifies spectral preconditioning as the primary mechanism, with widely linear mixing providing a secondary, dataset-dependent gain.