Breaking feature collapse in self-supervised time series encoder
Özgün Turgut ⋅ Daniel Rueckert
Abstract
General-purpose time series encoders such as \texttt{OTIS} and \texttt{MOMENT} promise a single backbone that transfers across domains, but they do not scale and rely on manually tagged domain labels at tokenisation. We trace both symptoms to a single pathology: under masked-data-modelling supervision the encoder's output patch tokens \textit{partially collapse}, with mean off-diagonal pair-wise cosine similarity stabilising at $0.68$ in \texttt{OTIS} and the encoder's effective rank capped accordingly. This is the natural consequence of supervising the encoder only in data space, which leaves no constraint against redundant token directions in the representation space. We introduce \texttt{OTISv2}, which suppresses the collapse by complementing the masked-reconstruction objective with two representation-space terms --- a self-distillation loss against an exponential-moving-average teacher and a KoLeo regulariser --- and replaces \texttt{OTIS}'s per-domain variate embeddings with register tokens to remove the label dependency. The resulting encoder pulls mean patch-to-patch cosine similarity from $0.63$ to $0.58$, scales monotonically from $7.1\,$M to $40\,$M parameters where its predecessor plateaus, and dominates general-purpose baselines on $143$ uni- and multi-variate datasets from the UEA and UCR archives from a single label-free checkpoint. We release code and pre-trained weights upon acceptance.
Chat is not available.
Successful Page Load