Temporal-Scale Sensitivity in Time-Series Tokenization and Scale-Robust Token Estimation by Gated Sum
Abstract
Transformers for temporal sequences require mapping sampled signals into sequences of vectors, called tokens. This is done by partitioning the sequence into local temporal windows (patches). We show that the tokenization step induces an approximation--estimation tradeoff. Smaller windows yield higher-variance statistical estimates because each token is supported by fewer samples, while larger windows stabilize estimation but increase the amount of information that must be summarized. Motivated by this tradeoff, we propose a lightweight tokenizer that constructs multiple estimates for each token and sums them, without changing the Transformer backbone. Across 12 datasets spanning four modalities, our method improves performance and reduces sensitivity to temporal scale across tasks.