AIM: Adversarial Interface Manipulation
Abstract
Web agents complete tasks on websites that often contain misdirections and cues to distract the user from their original intent. Agents often follow these cues and end up doing something the user did not ask for. Existing benchmarks study this behavior using researcher-defined families of manipulative patterns, so the set of manipulations an agent is tested on never grows. We introduce Adversarial Interface Manipulation (AIM), a meta-agent method that automatically generates manipulated web environments to test whether web agents stay on the user's task when a webpage steers them elsewhere. Given a user task and a working website, a teacher agent proposes manipulations and an implementer agent builds them into a copy of the site, while automatic checks reject changes that are broken, unfair, or make the original task impossible. A student web agent then attempts the same request on each manipulated site. When the manipulation steers the student into an action the user did not request, we define this as a capture. Across ten websites and 100 tasks, AIM with GPT-5.6 Luna produced 50 validated captures. We then tested these manipulated environments on four other agents, and all were captured, including GPT-5.6 Sol at 59% and Qwen3.5-9B at 90%. AIM is a meta-agent loop that builds web environments hard enough to capture state-of-the-art agents, and those environments transfer beyond the agents used to generate them.