Saturn: Self-Supervised Time Series Models Enable Classification and Event Forecasting
Abstract
Time-series foundation models are pretrained one task at a time: classification encoders by contrasting augmented views or reconstructing masked spans, forecasters by predicting future values. A shared task interface can already serve both, but no single self-supervised objective learns one representation for both. We present Saturn, a joint-embedding predictive architecture in which the mask split that sets up the prediction task also supplies the contrastive negatives. One horizon is drawn per batch, so an episode's own future is its positive and every other episode's future at that same horizon forms a negative pair: one construction both carries the predictive signal and rules out the constant solution, with no prior on the latent distribution and no separate teacher. One representation, learned by that objective alone and joined to no task head, then serves both readouts. On event forecasting it beats the same encoder left untrained on ranking, timing and early warning at every seed, and the early-warning gain persists when the evaluation's dynamics families are withheld from pretraining. Probed frozen for classification it beats that untrained encoder on the univariate archive while trailing the published field. A simpler future-value pretext forecasts better on timing but worse on ranking and early warning, so the unification is a trade rather than a free win. One objective that both recognises the present and forecasts what is coming is a step toward a physical time-series world model, and we report where it falls short.