Answers Expire: Delivery-Time Evaluation of Spatial Reasoning in Egocentric Assistants
Abstract
Embodied spatial reasoning is evaluated as if it lives in a frozen world: a model is scored against the environment exactly as it was when the question was asked. However, in real-world deployment, time passes during inference, and the world changes. We introduce delivery-time evaluation, a framework that scores spatial accuracy at the moment a model is delivering an answer, making the model's own latency a correctness property. Testing this across two scenarios whose answers expire for different physical reasons (an object being relocated by someone else (Reloc), or the wearer walking (Dist)), we show that delivered-accuracy loss grows with how often a model's latency outlives the queried fact's time-to-change. We then test methods often assumed to improve answers, such as Chain-of-Thought reasoning: they leave query-time accuracy unchanged while lowering delivered accuracy, because longer thinking means later delivery. Changes caused by others remain largely unrecoverable, while loss from the wearer's own motion can be recovered in principle with geometric odometry, a recovery we bound with an oracle correction. Because a second of thinking is a second in which an answer can stop being true, leaderboards that underweight latency will keep favouring computationally expensive systems over assistants that arrive while their answers are still true. Real-world embodied AI cannot simply scale test-time compute. It must explicitly co-optimize for reasoning latency.