DriveMind: Mind-Evolving Belief Tracking for Closed-Loop Autonomous Driving
Abstract
Recent multimodal large language models (MLLMs) can perceive a traffic scene, reason about it, and propose a driving action within a single model. However, existing MLLM-based drivers reason from scratch at each frame and commit to a single explanation, discarding earlier hypotheses about the scene. When the scene is ambiguous, the agent locks in on the wrong interpretation and acts on it before evidence accumulates. We propose DriveMind, an MLLM-based driving framework that keeps several hypotheses about the scene alive at the same time and updates their weights as new frames arrive. Inspired by particle filtering, DriveMind propagates each hypothesis through a VLM, scores it against the current observation, and resamples to maintain diversity. On Bench2Drive, DriveMind raises the Driving Score (DS) from 71.36 to 81.37 and the Success Rate (SR) from 45.45% to 54.09%, with the largest gains on ambiguous scenarios such as unsignalized intersections (+20.10 DS) and oncoming-traffic interactions (+20.50 DS). Beyond the overall metrics, DriveMind also improves fine-grained interactive driving abilities such as merging, emergency braking, and give-way. The code will be released soon.