Automotive-ENV: Benchmarking Multimodal Models in Automotive Cockpit Environments
Abstract
Multimodal models have demonstrated strong capabilities in web, desktop, and mobile GUI environments, but their performance in automotive cockpit systems remains largely unexplored. In-vehicle GUIs introduce unique challenges beyond traditional GUI interaction, including vehicle-state reasoning, implicit intent understanding, and safety-critical decision making. To address this gap, we introduce Automotive-ENV, an interactive benchmark for evaluating multimodal models in automotive cockpit environments. Built on Android Automotive OS, Automotive-ENV combines cockpit-screen interaction with vehicle-state signals and deterministic state-based evaluation. The benchmark contains 420 carefully constructed tasks covering routine cockpit operation and safety-critical scenarios. Unlike prior GUI benchmarks that rely on action matching or LLM-based judging, Automotive-ENV evaluates model behavior through reproducible validators over UI states, vehicle signals, and safety-related conditions. We benchmark representative general-purpose multimodal models and GUI-specialized vision-language models on Automotive-ENV. Experimental results show that current models remain far from reliable in automotive environments, with substantial performance degradation in implicit and safety-critical tasks. Further analysis reveals consistent failure modes in safety-aware reasoning, vehicle-control understanding, and implicit hazard recognition. We will release Automotive-ENV, together with tasks, validators, and evaluation tools, to support future research on automotive multimodal systems.