Put the World Back: Timestamps Matter in Time Series Reasoning
Abstract
Large language models (LLMs) are increasingly applied to time series reasoning. However, current evaluations often provide temporally ordered and manually aligned context, making it unclear whether LLMs reason by linking time series to the event sequence or simply exploit latent shortcuts. To expose this ambiguity, we reformulate time series reasoning from a world-to-series projection perspective, viewing timestamps as explicit indices that connect numerical observations and external events to shared world states. We identify three structures that timestamps can make explicit: temporal causal structure, time semantics, and event--series alignment. Building on this perspective, we develop a diagnostic protocol that distinguishes timestamp-grounded reasoning from correlational shortcuts, a distinction that accuracy-based benchmarks cannot make. Controlled interventions on 150 questions reveal that apparent temporal competence can persist even when clock values are removed, while breaking chronological context induces systematic failures. Our analysis shows that models can substitute sequence order for temporal grounding, construct false local narratives from temporally displaced events, and privilege those narratives over authentic numerical evidence. These findings expose a limitation of current evaluation: high accuracy on chronologically ordered inputs does not establish that a model uses timestamps to synchronize observations with the world. We release our code and reasoning trajectories to support future research: \url{https://anonymous.4open.science/r/timestamps-matter}.