Driving on Memory
Abstract
End-to-end autonomous driving models learn to plan the vehicle's future trajectory from raw sensor observations. While earlier driving benchmarks often measured deviation from the human trajectory, current benchmarks such as NAVSIM and Bench2Drive evaluate models with richer simulation-based metrics intended to capture safe and compliant driving. A high benchmark score should reflect that a model can understand the scene in front of it and act accordingly. But how much of that score specifically comes from reacting to the dynamic part of that scene? To probe this, we remove a model's camera input and replace it with memories from prior drives of the same location. The retrieved memories can provide persistent scene information, including road layout and location-conditioned regularities, but because they are historical, they do not reveal the current actors, signal states, or transient conditions. Surprisingly, memory is nearly sufficient on NAVSIM, reaching or even exceeding the performance of leading end-to-end methods without actually observing the evaluated scene. Our results suggest that a high NAVSIM score does not require a planner to react to the current traffic scene and should therefore be treated with caution. However, we do not observe the same behavior on Bench2Drive: driving from memory remains far from camera-based performance, suggesting that Bench2Drive better preserves interaction demands through longer closed-loop routes. To ensure reproducibility, we provide our code at anonymous.4open.science/r/MemoryDrivoR-E64C.