From Click Imitation to Transition Equivalence: Rethinking Supervision for GUI Agents
Abstract
GUI agents are increasingly expected to complete real tasks through language-based interaction, yet most supervised training still treats a single demonstrated click or action as the unique ground truth. This creates a mismatch between how agents are trained and how they are evaluated: task success depends on reaching the intended interface state, not on reproducing the demonstrator's exact surface action. We address this gap by introducing Transition-Equivalent Supervision (TES), a framework that trains GUI agents around goal-relevant state changes rather than raw action identity. TES converts demonstrated interactions into transition descriptors, supervises agents to predict the intended transition, and learns an executor that selects any verified action capable of realizing that transition. At inference time, the agent further verifies whether the executed action produces the intended state change, reducing repeated no-op or erroneous behavior. Across offline web prediction, executable web navigation, mobile control, and desktop-use benchmarks, TES consistently improves equivalent-step success, alternative-action recall, and online task completion while reducing NoOp/Error transitions. These results suggest that GUI-agent learning should move beyond click imitation toward transition-level reasoning, offering a more robust supervision principle for agents operating in diverse and dynamic interfaces.