PopupBench: Diagnosing Popup Blindness in Vision–Language GUI Agents
Abstract
Real interfaces contain foreground dialogs that may either advance or interrupt the current task, yet GUI agents are commonly evaluated on clean screens. We present PopupBench, a controlled first-click benchmark for this decision. It contains 9,494 Ubuntu and Windows cases built from 197 real dialog templates, plus 368 cases from 150 naturally rendered web screenshots. Holding the screen fixed while varying the task creates paired conditions: an agent must engage a related popup but dismiss an unrelated one. Across nine open and three proprietary models, related popups are usually easy, whereas general VLMs score below 5\% on unrelated OS cases and click the background in 70--86\% of them. Fixed-screen and click-geometry analyses show that most models do not apply a context-sensitive popup policy. PopupBench deliberately isolates this first action rather than end-to-end recovery.