What Do Time Series Foundation Models Learn About Physical Structure?
Abstract
Time-series foundation models for physical domains are typically trained with self-supervised learning objectives that do not explicitly encode the physical processes generating the data. Yet their representations are increasingly used for downstream physical tasks such as activity recognition, road-surface classification, and fault detection. In this work, we compare time-series foundation models on physical tasks and investigate whether their downstream performance is related to how linearly recoverable physical structure is from their representations. This question is motivated by recent work showing that representations learned by self-supervised objectives can recover underlying factors up to a linear transformation. We evaluate three models using linear probes: Newton, trained with a JEPA-style objective; MOMENT, trained with a masked-reconstruction objective; and Chronos-2, trained for forecasting. We find that Newton achieves the strongest linear probe performance, and that this performance corresponds to a measurable representational property: the linear recoverability of underlying physical factors. Newton achieves the highest linear-probe accuracy on 12 of 14 open source datasets. To test whether this performance reflects how well each model encodes the physical structure of the underlying processes, we use a controlled setting in which the ground-truth physical factors are known. Specifically, we generate synthetic signals from a small set of physical parameters, including frequency, decay, amplitude and noise, and for each model, fit a linear map from its frozen embeddings to these ground-truth parameters, measuring the coefficient of determination (R2). Overall, Newton recovers the physical factors most completely, ahead of MOMENT and Chronos-2, and all three recover them far better than the raw-signal and random-projection baselines. Thus, Newton leads both in downstream classification and in recovery of known physical factors. These findings are empirical and correlational: although linear recoverability is associated with downstream performance across our models, establishing whether this relationship is causal requires a more principled evaluation, which we leave for future work.