An Extractor Without a Tunnel: Layerwise Probing and Spectral Analysis of a Time-Series Foundation Model
Abstract
Time-series foundation models pretrained for forecasting can also provide representations for forecasting and classification through newly fitted prediction heads. Which layer provides the most useful features remains an empirical question. In supervised image classifiers, the tunnel effect describes late-layer compression that preserves task accuracy while reducing transferability. We ask whether this pattern appears in frozen Chronos-2. Our primary comparisons fit a separate linear probe at each encoder layer and use the final block before output normalization as their reference. We adapt the tunnel test by locating the first layer reaching 95% of final-layer probe accuracy and measuring representation dimensionality with effective rank. On two datasets documented in pretraining, probe accuracy peaks at intermediate layers. However, dataset-pooled effective rank, which measures how many dimensions contribute to the representations, rises roughly threefold from the threshold layer to a later peak. The measured intermediate-layer advantage does not consistently increase when transferring to other forecasting datasets or classification tasks, or perturbing inputs. Final normalization removes most of the last-block drop in pooled effective rank, while the original forecasting head performs best at the final layer. These findings depart from the compression pattern predicted by the adapted test and show why layer comparisons must specify the prediction head and normalization used.