Text-Informed Gated Estimation with Residuals for Multimodal Time-Series Forecasting
Abstract
Multimodal time-series forecasting can benefit from paired text, but incorporating text at every forecasting step may hurt performance when the numerical signal is already sufficient. We propose TIGER (Text-Informed Gated Estimation with Residuals), which keeps a strong numerical forecast from Time series foundational model as the default and admits paired text only as a content-gated residual correction. The numerical default is a frozen time-series foundation model (TimesFM, Moirai, Sundial, Chronos, or TimesMoE), then fit only the trainable TIGER adapter (multimodal) without updating foundation-model weights. On Time-MMD benchmark, this setup improves average MSE and MAE for all five backbones over their uni-modal baselines and also shows competitive performance against other multimodal timeseries model like TaTs and MM-TSFLib.