W-JEPA: Wavelet Joint-Embedding Predictive Architecture for Wearable Activity Recognition
Abstract
Wearable activity recognition is challenged by heterogeneity across datasets in sensor placement, channel count, and sampling rate, which limits both cross-dataset transfer and generalization to unseen participants. To address this challenge, we introduce W-JEPA, a self-supervised framework that exploits the multiscale structure of inertial motion and leverages a joint-embedding predictive architecture. It has a fixed wavelet tokenizer that summarizes each sensor window as tokens indexed by scale, sensor axis, and time, so that each token carries a physically meaningful address. It then encodes the visible tokens with a Transformer and predicts continuous latent targets at masked token positions, supplied by an exponential-moving-average (EMA) target encoder. Recovering a hidden token from the visible ones then encourages relating motion at one scale on one axis to motion at other scales on other axes. Experiments on six wearable activity benchmarks demonstrate that W-JEPA not only outperforms state-of-the-art methods under the random-window training splits but also retains the strongest performance when participants are held out.