MMA-SafetyBench: A Benchmark for Multimodal Agent Safety Evaluation
Abstract
As multimodal agents increasingly rely on visual perception to navigate complex digital workflows, the vulnerability of their reasoning-action cycles to cognitive manipulation has emerged as a critical security priority. This systemic reliance on visual-semantic grounding introduces a fundamental susceptibility to executive hijacking. Unlike traditional pixel-level adversarial noise, this threat operates at the cognitive-semantic level by synthesizing high-fidelity, deceptive UI elements—such as forged system alerts or mandatory compliance overlays—that seamlessly subvert an agent's safety alignment. We propose a targeted evaluation methodology that injects context-aware payloads into specific vulnerability windows within an agent's reasoning trace. By specifically changing the attack content within the observation space while strictly maintaining a static attack strategy, we demonstrate the consistent subversion of executive intent across diverse task trajectories. Building upon this, we introduce MMA-SafetyBench, the first comprehensive safety evaluation suite for multimodal agents. Comprising 1,083 adversarial trajectories, the benchmark evaluates these vulnerabilities across five ubiquitous environments: web automation, desktop/mobile OS control, visual document understanding, and long-form video reasoning. Our evaluation of nine state-of-the-art foundational models reveals a pervasive risk of executive hijacking, where agents disproportionately prioritize deceptive observations over benign textual constraints. With the Semantic Compromise Rate (SCR) reaching an alarming 98.09%, our findings expose a critical blind spot in current safety alignments and underscore the urgent need for cross-modal cognitive defenses in autonomous multimodal architectures.