Scale Invariant Training for Time Series Foundation Models
Ignacy Stepka ⋅ Willa Potosnak ⋅ Kin Gutierrez ⋅ Artur Dubrawski
Abstract
TSFMs trained across diverse datasets face significant scale variation across series. Affine scaling methods like ReVIN normalize inputs but invert the transform before computing the loss, which we show makes each series' gradient proportional to $b^p$ (scaling denominator $b$, loss degree $p$), causing high-scale series to dominate training. We call this setup *scale-contaminated training* (**ScaleCon**) and show that computing loss directly on scaled targets yields *scale-invariant training* (**ScaleIn**). We prove that for any scale-equivariant scaler and degree-$p$ homogeneous loss (MSE, MAE, QL), arbitrary rescaling of training series leaves gradients and the optimization trajectory unchanged. With a set of synthetic experiments we show that under **ScaleCon**, artificially inflating scale of datasets causes disparate scale-induced per-dataset convergence rates, while having no such effect under **ScaleIn**. Furthermore, across TSFM pretraining and supervised forecasting we show that **ScaleIn** improves accuracy, with MASE reductions of up to 57.8\% and WQL reductions of up to 39.3\%. In most standard forecasting pipelines, **ScaleIn** is an easy-to-add one-line code change.
Chat is not available.
Successful Page Load