Beyond Raw Observations: Distilling and Storing Invariant Driving Memories for Generalizable Autonomous Driving
Abstract
Modern end-to-end autonomous driving systems often suffer from fragile out-of-distribution (OOD) generalization and limited decision transparency, largely due to their reliance on raw sensory observations and susceptibility to spurious visual correlations. Although memory-augmented approaches have been explored to improve generalization, existing methods typically store raw observations, symbolic rules, or entangled latent features, making retrieval sensitive to surface-level similarity rather than decision-relevant driving structure. To address this limitation, we propose DriveMem, a plug-and-play framework that moves beyond raw observations by distilling and storing invariant driving memories for generalizable driving systems. DriveMem integrates two core components: Invariant Feature Abstraction that first distills perturbation-stable, decision-relevant representations from visual tokens, and Prototype Memory Bank that then stores and retrieves prototypical driving experiences grounded in invariant representations. Extensive open-loop and closed-loop experiments demonstrate that DriveMem improves OOD generalization and safety-related planning metrics across diverse unseen scenarios. Notably, under severe sensor disruptions, DriveMem achieves up to a 39.5\% relative reduction in Collision Rate compared with a strong domain-generalization baseline. These results suggest that building prototype memory in an invariant representation space is a promising approach for improving robustness, while offering prototype-grounded traceability for model behavior.