Beyond Risky Activities: Bridging the Supervision Gap for Situational Risk Reasoning
Abstract
Situational risks refer to safety risks that arise from the interaction between a user activity and its surrounding visual scenario, where the same activity may be safe in one situation but unsafe in another. Identifying such risks is particularly important for deploying VLMs as real-time assistants, where models are expected to assist users based on a shared visual context. However, existing safety alignment data is dominated by examples of intrinsically harmful activities, leaving activity--scenario risks underrepresented. We refer to this missing activity--scenario supervision as a supervision gap in situational risk reasoning for VLMs. To address this gap, we introduce Situational Risk Reasoning (SRR), the first post-training dataset for improving VLMs' situational risk awareness. SRR is constructed through a hybrid pipeline that combines vision-language reasoning, language-based generation, and text-to-image synthesis. Each example includes a target response in a safety chain-of-thought format, guiding models to reason about activity--scenario interactions rather than simply learning refusal patterns. Experiments show that supervised fine-tuning on SRR substantially improves VLMs' ability to reason about situational risks, while maintaining strong performance on conventional safety benchmarks and general VQA tasks. We also demonstrate the effectiveness of SRR across different VLM backbones and validate the importance of both the synthesized SRR data and the safety chain-of-thought format through ablations. The SRR dataset is available at https://huggingface.co/datasets/Anonymous-SRR/SRR.