ViZDoom-IDM: A Benchmark for Inferring Actions from Video Games
Abstract
Action-conditioned world models learn how an environment responds to a player's controls, but training them requires trajectories with action labels. Gameplay recordings are abundant, yet they often omit those controls. Can vision-language models (VLMs) recover the missing labels from video-game footage? We introduce \textbf{ViZDoom-IDM}, a benchmark for inferring six navigation actions from short first-person clips. Across closed and open-weight VLMs, parameter-efficient VLM adaptations, and a supervised CNN evaluated on the same 600 examples, the CNN reaches 73.86\% macro-F1 after one completed epoch. Across the three prompted closed-source models, the mean macro-F1 is 13.95\%, while prompted Qwen3.5-4B and Qwen3.5-9B score 4.62\% and 6.04\%; QLoRA-adapted Qwen2.5-VL reaches 17.91\%. The CNN results show that the clips contain a strong learnable action signal, but the tested VLMs do not reliably recover explicit controls. Improved action grounding, potentially through targeted pretraining, could enable automated labeling pipelines for scaling action-conditioned world-model datasets.