Evaluating the Readiness of VLMs for Real-World Autonomous Driving Tasks
Abstract
VLMs are increasingly being explored in autonomous driving for scene understanding, driving reasoning, decision-making, and end-to-end driving. While recent models have shown promising results on standard benchmarks, their reliability under degraded visual conditions remains less understood. In real-world driving, camera inputs are rarely ideal and are affected by environmental and sensor-related challenges such as illumination changes and adverse weather. These perturbations can influence not only prediction accuracy but also model uncertainty and calibration, which are particularly important in safety-critical settings. In this work, we study the robustness of three general-purpose VLMs, Gemma4-4B, LLaVA OV-8B, and Qwen3.5-9B, along with DriveFusionQA-3.75B and NVIDIA Alpamayo1.5-10B. For our analysis, we consider four common camera perturbations. Specifically, we select one representative example from each corruption category, including glare for illumination changes, fog for weather conditions, motion blur for sensor-related effects, and lens occlusion for occlusion. This setup allows us to examine how different forms of visual degradation affect model behaviour. To evaluate performance across different driving scenarios and reasoning tasks, we use four autonomous driving VQA benchmarks, including DrivingVQA, STRIDE-QA, NuScenes-QA, and Open Spatial Reasoning. These datasets cover different tasks such as traffic-rule and safety reasoning, egocentric spatial understanding, multi-view scene reasoning, and monocular 3D spatial reasoning. Across these benchmarks, the effect of perturbations on accuracy is mixed and dataset dependent. However, confidence calibration consistently worsens under corrupted inputs, showing that models become less reliable in expressing uncertainty when the visual input is degraded.