Retrieve-then-Rerank Inference for Vocabulary-Based End-to-End Driving
Hyunjun Kim ⋅ Donggue Kim ⋅ Kichun Jo
Abstract
Vocabulary-based end-to-end driving selects a future trajectory from a fixed candidate set, enabling explicit comparison among multiple plausible plans. However, existing planners typically process the scene-invariant trajectory vocabulary and the scene-dependent driving context together in a single online forward pass. This repeatedly recomputes reusable candidate-side representations and applies expensive scene-conditioned evaluation to the full vocabulary. We revisit vocabulary-based planning as a retrieve-then-rerank problem, where the trajectory vocabulary is a static candidate collection, lightweight scene cues form a retrieval query, and the scene-conditioned evaluator acts as a reranker. Based on this view, we propose $\textbf{Decoupled Retrieval-Style Inference (DRSI)}$, centered on two modules: $\textbf{Offline Candidate Indexing (OCI)}$ and \textbf{Scene-aware Candidate Retrieval (SCR)}. OCI caches scene-invariant trajectory embeddings offline, while SCR retrieves a compact set of route-consistent and dynamically reachable candidates online. The scene-conditioned evaluator is then applied only to this retrieved subset for fine-grained reranking. Experiments on NAVSIM show that DRSI substantially improves the latency--quality trade-off without adding trajectory proposals or increasing evaluator complexity. Compared with the Hydra-MDP++ baseline, our large DRSI model achieves 86.1 EPDMS on NAVSIM v2 with 23.9 ms latency, corresponding to about a 6.8$\times$ speedup. On the challenging navhard split, DRSI further improves EPDMS from 37.6 to 40.2. These results demonstrate that retrieval-style decoupling is an effective inference principle for efficient vocabulary-based end-to-end driving.
Chat is not available.
Successful Page Load