Depth Compression in Time-Series Foundation Models: Saturation, Recovery, and Measured Speedups in Chronos-2
Egor D. Serebriakov ⋅ Zibo Shang ⋅ Amartya Mitra ⋅ Alexis Roger ⋅ Kashif Rasul ⋅ Irina Rish
Abstract
Time-series foundation models (TSFMs) support zero-shot forecasting but can be too expensive for latency- and memory-constrained deployments. Most compression methods reduce parameter count or numerical precision. In this work, we ask whether model depth itself can be reduced by skipping later layers. Using Chronos-2, we find that forecasting performance using a direct linear readout saturates before the final representation under a 5\% criterion on all tested datasets, although the saturation depth varies across datasets. Effective rank and CKA reveal clear layerwise structure but do not identify a universal transition associated with saturation. We then attach lightweight linear adapters to intermediate representations and reuse the frozen Chronos-2 quantile head. Dataset-specific adapters often recover much of the native forecasting performance, whereas cross-dataset transfer is less consistent. We show that truncating at layer 3 drops $71\%$ of the active parameters, reduces peak GPU memory by $34\%$ (on batches of 256), and provides a $\times 3.23$ measured speedup, with a MASE increase of $7\%$ averaged across three datasets. Depth compression can therefore provide substantial efficiency gains, with the optimal truncation depth and associated gains varying by dataset.
Chat is not available.
Successful Page Load