In Search of Lost Time: Temporal Competence in Large Language Models
Abstract
Today's large language models (LLMs) increasingly operate as agents that schedule work, coordinate with people and other AI systems, and trade-off accuracy against latency. Yet LLMs only encounter time secondhand---in descriptions rather than experience---and whether they can nonetheless reason reliably about time remains unclear. We find that their competence bears the mark of this origin: models know the rules of time, which can be read, but struggle to predict durations, which must be lived; they misestimate their own completion times, and can disregard time costs even when they are stated outright. We establish these through a three-level evaluation paradigm. At the level of structural knowledge, we synthesize findings across the literatures of AI, psychology, operations research, and economics into a shared ontology and test three frontier models on 224 axiom-grounded vignettes. Accuracy is high (91.9--95.8\%), with small yet consistent errors involving fatigue, skill decay, and interruptions. At the level of calibration, against observed completion times, models estimate human durations moderately well but underestimate their own execution times by 37--99.5\%, and fail to identify which of their tasks require longer runs. At the level of decision-making, in tool use experiments with explicit time costs, models use tools more when they expect larger accuracy gains, but GPT-5.6 Sol ignores a 1,200-fold increase in wait time entirely. Together, these results locate where temporal competence breaks down, revealing a new set of limitations for agent scheduling, coordination, and responsiveness.