Can VLMs Reason When to Stop for Human Safety?
Abstract
Since Vision-Language Models (VLMs) are the core of the perceptual and reasoning capability in Vision-Language-Action models, which are deployed in environments shared with humans, evaluating VLMs' safety reasoning from video is critical for preventing human harm. Current video-based benchmarks, however, evaluate only whether a model can recognize dangers that are already visible, and overlook two capabilities that VLMs require: reasoning under hierarchical principles in which human safety takes precedence over obedience to user instructions, and anticipating harm rather than identifying it after it has occurred. We introduce VESPER, a video benchmark that targets both capabilities by scenarios in which a robot is executing a user-assigned task when a person unexpectedly moves into harm's way, requiring the model to predict the potential harm and decide whether to continue or stop. The benchmark comprises 458 videos covering matched safe and dangerous variants, built through a human-in-the-loop pipeline from diverse taxonomies. We evaluate 21 VLMs, and the results indicate that the current VLMs fail to properly decide the actions under such a scenario.