Do Time Series Need Words? Encoding Numeric Context for Time Series Foundation Models
Abstract
Auxiliary information can improve time-series prediction, but its usefulness depends on both its content and representation. We compare identical numerical auxiliary values encoded as discrete tokens, continuous vectors, or template sentences processed by frozen BERT with trainable input embeddings. Using two frozen time-series foundation models across vibration, ECG, weather, and hydraulic classification, we examine signal-derived features, additional measurements, and metadata. Across the evaluated domains and backbones, sentence encoding offers no consistent advantage over numerical representations, and no encoding is uniformly best. Substantial gains are achievable without natural-language encoding: discrete tokens improve weather classification accuracy by up to 14.7 percentage points over signal-only. Additional measurements and persistent metadata yield different benefits, suggesting that useful inputs should be physically relevant and reflect conditions that change across time-series windows. Our findings emphasize selecting informative auxiliary inputs and evaluating their contribution separately from how they are encoded.