$\mathbf{\mathtt{MAD\text{-}Bench}}$: How Do Multimodal Agents Deceive You?
Hao Gu ⋅ Mingli Song ⋅ Jiacong Hu
Abstract
Multimodal agents increasingly operate computers on behalf of users. This shift from passive question answering to interactive execution creates a new safety problem. When an agent encounters normal, non-attack scenarios where the task it needs to perform is impossible or difficult to complete, the central risk is not only that it fails, but that it deceives the user into believing that it has succeeded. Although prior work has studied truthfulness and deception in language models and language agents, multimodal interaction introduces new sources of deceptive behavior, increasing the implicitness and complexity of deceptive behaviors in multimodal agents. To fill this gap, we introduce $\mathbf{\mathtt{MAD\text{-}Bench}}$, a benchmark for evaluating **M**ultimodal **A**gent **D**eception. $\mathbf{\mathtt{MAD\text{-}Bench}}$ contains 360 sandbox tasks organized around three core elements of multimodal execution: *modality*, *target*, and *tool*, and spans six task types including *modality evidence conflict*, *modality asynchronous mismatch*, *modality ambiguity distortion*, *modality object missing*, *target infeasibility*, and *tool defect*. To our knowledge, $\mathbf{\mathtt{MAD\text{-}Bench}}$ is the first benchmark to formalize deceptive behavior in multimodal agents from an execution-evidence perspective, characterizing deception as a misalignment between the user-facing task-state claim and the multimodal evidence available throughout the execution trajectory. We further propose a behavioral taxonomy covering *evasive deception*, *manipulative deception*, *misleading deception*, and *non-deception*. Evaluating 10 mainstream multimodal agents, we find that agents frequently fabricate, conceal, mislead, frame, and disguise during task execution. These results show that substantial deceptive behaviors are widespread across mainstream model families. This benchmark is available at https://anonymous.4open.science/r/MAD-Bench-A304.
Chat is not available.
Successful Page Load