Masking Is Not Free: Masked Reconstruction Failure in EMG Foundation Models under Corpus Scaling
Abstract
Surface electromyography (EMG) foundation models learn transferable representations through masked reconstruction, but existing models are typically pretrained on only a small number of datasets. Expanding the pretraining corpus is a natural direction toward improving generalization, yet we find that existing pretraining architectures fail to recover meaningful EMG signals when pretrained on the expanded corpus. Our analysis identifies two failure modes. First, shared mask tokens lack an explicit position-specific component at the token input, leading to insufficiently differentiated masked representations. Second, zero-padded invalid channels are included in encoder computation and reconstruction loss, diluting the contribution of valid EMG channels. To address these limitations, we propose Position- and Validity-Aware Pretraining (PVP), which introduces a learnable positional embedding on signal and masked tokens and excludes invalid channels from encoder computation and reconstruction loss, without altering the backbone or overall pretraining framework. Across EMG foundation models, PVP improves masked reconstruction under corpus scaling, restores masked representation diversity, and improves reconstruction on valid EMG channels. These improvements also lead to better downstream performance. These results highlight that explicitly modeling masked token position and channel validity is important for robust masked pretraining as EMG pretraining corpora scale.