ChronosAlign: A Large-Scale, Updatable Benchmark for Decomposing Temporal Alignment in LLMs
Abstract
Large Language Models (LLMs) struggle with temporal alignment, often producing outdated time-sensitive facts as real-world knowledge evolves. Yet the field's ability to measure this misalignment is itself fragile. Existing benchmarks are static snapshots, rely on rigid templates, are small, lack provenance, conflate willingness-to-answer with recall, and use phrasings that newer models have likely memorised. We argue that progress on temporal alignment is bottlenecked less by data scarcity than by evaluation design, and we contribute on both fronts. We present ChronosAlign, a natural language, programmatically updatable dataset and evaluation framework for temporal alignment in LLMs; generated from a Wikidata and Wikipedia corpus of ~500,000 question-answer pairs covering ~27,000 unique questions across sport, politics, and culture from 2000 to 2025. Each question carries a Wikipedia URL and a documented update procedure, so the same evaluative claim can be re-tested against future models without re-authoring the benchmark. We couple ChronosAlign with a framework focussed on three prompting modalities, explicit (year-anchored), relative ("current" anchored), and invariant (untimed control), whose disagreement isolates which component of temporal alignment a model fails. We benchmark sixteen LLMs surfacing three findings: (i) a sharp post-2022 decay in explicit recall even for models with later training cutoffs, (ii) systematic misalignment between explicit and relative recall, with some models better aligned to explicit year prompts and others to relative anchors, and (iii) a confounding effect of refusal behaviour. Our generation and update code can be found at: https://github.com/sg-sy/ChronosAlign. Our dataset can be found at: https://huggingface.co/datasets/sg-sy/ChronosAlign.