Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Abstract
Computer-use agents (CUAs) increasingly carry out enterprise workflows through desktop GUIs. Existing benchmarks measure end-to-end task success or single-screen grounding, but not whether models understand what changed after an action, as needed to reject stale observations, verify progress, and recover from failure. We introduce Desktop-Delta Bench (DDB), an offline benchmark with 2,013 human-verified Linux samples across 50 task domains and ~15 applications. DDB comprises 2 tasks: (1) temporal ordering requires models to order 3 shuffled screenshots and reject a cross-recording decoy when present; (2) single-action reconstruction requires models to recover the action family and exact payload from immediate before-after frames. It contains 463 ordering samples, 105 with decoys, and 1,550 action pairs. Across 8 model families, best exact ordering is 65.1% without decoys and 65.7% with them. Task context raises decoy identification but also lowers exact ordering. Some models make the critical error of copying presentation order: A → B → C accounts for 37.7–45.6% of GLM-4.6V and MiniMax M3 failures. Models recognize clicks more reliably than drags (F₁: 0.96 versus 0.76), although recognized drags are localized well. DDB complements, rather than replaces, end-to-end benchmarks, filling the gap between GUI grounding and task-level success and targets CUA verification, reliability, and recovery.