A$^3$Bench: A Benchmark for Compositional Reasoning over Aggressive Interactions in Videos
Austin C Baker ⋅ Raiyaan Abdullah ⋅ Daniel Z Yoffe ⋅ Shruti Vyas ⋅ Yogesh Rawat
Abstract
Recent Multimodal Large Language Models (MLLMs) have achieved strong performance on general video understanding benchmarks, yet their ability to reason about participant roles in aggressive interactions remains largely unexplored. Existing violence and crime datasets primarily evaluate whether aggression occurs, without testing whether models can identify who performed the action, who received it, and how participants are related within the scene. To address this gap, we introduce **A$^3$Bench**, a benchmark to study *aggressive action and association* in videos. The benchmark evaluates whether models can jointly recognize aggressive actions, distinguish aggressors from victims and bystanders, and correctly bind actions to participants under challenging distractor settings. We benchmark 15 recent MLLMs and find that current models struggle substantially on fine-grained aggressive-scene understanding. Performance degrades consistently as reasoning complexity increases, with models frequently confusing aggressors with victims or bystanders, revealing a systematic failure in role-action association rather than coarse aggression recognition alone. We further observe that increasing model scale or applying generic reasoning prompts provides limited improvement. Finally, we explore a simple role-graph prompting strategy that encourages structured reasoning over participant interactions, leading to consistent gains on compositional reasoning. Our findings highlight that reliable aggressive-scene understanding remains a major challenge for current MLLMs and motivate future research on role-aware video reasoning beyond coarse violence detection.
Chat is not available.
Successful Page Load