Relation-Integrity Attacks in Browser Agents: Separating Unsafe Proposals from Execution
Abstract
Browser-agent security is mostly studied as instruction integrity: untrusted page text carrying a command that competes with the user's request. But a page is also an executable interface. Its labels, images, elements, and destinations can disagree with each other while no instruction appears anywhere. A browser agent must keep three things in agreement: what the page shows, which element it picks, and what that element actually does. Break one of those links and a harmless-looking click performs a prohibited action. A control labelled Share can still trigger Publish Ad. We call this an instruction-free relation-integrity attack. We contribute a threat model for these attacks, a benchmark that separates model failure from system failure, and a trusted check placed between the model's choice and the browser. We evaluate on a local classified-ads site under three ways an agent can see a page: a DOM tree, a plain screenshot, and Set-of-Mark, a screenshot with its clickable elements numbered. Every run records five stages on its own: what the model proposed, what the parser read, what the check decided, whether the browser acted, and whether the task succeeded. Under a hidden Share/Publish label swap, agents followed the label and proposed the prohibited Publish action in 84 of 95 units, including on DOM, where the element's true destination was in view. Before the click reached the browser, a trusted monitor read the destination from the chosen element itself, rather than from what the page displayed. It knows which elements to trust because it registered them before the attack altered the page. It blocked all 84, and none reached a browser attempt or execution. A screenshot cannot reveal this attack, and the models did not check the DOM evidence that could; the defence therefore belongs at the action boundary, where it held. A second experiment shows the same corruption costing utility rather than safety. Binding a displayed image to the wrong listing cut task success from 105/162 (64.8%) to 79/162 (48.8%), against a matched harmless change that left image and listing correctly paired. That cost fell almost entirely on one grounding: 27.8 points for a plain screenshot against 1.9 for DOM. The corruption rebinds an image, which is a screenshot's only identity evidence, while DOM also carries listing text the attack leaves intact. These interface figures are post-hoc diagnostics over a fixed design. The 84/84 block rate has a narrow meaning. The monitor applies one fixed policy naming two destinations, one permitted and one prohibited. It never decides whether a page is under attack.