PrecisionCUA: Iterative Visual Refinement for Pixel-Precise Cursor Grounding in Code Editors
Abstract
Computer Use Agents (CUAs) rely on GUI grounding to translate language instructions into screen actions, yet fine-grained grounding in dense coding interfaces such as VS Code and Cursor remains underexplored. We study pixel-precise cursor localization and introduce an iterative grounding approach that uses visual feedback from previous predictions to correct localization errors. We evaluate this approach across Claude, Qwen, and GPT models on a benchmark spanning VS Code and Cursor, demonstrating that multi-turn refinement significantly outperforms state-of-the-art single-shot models in both click precision and task success rate. Our results suggest that iterative visual reasoning is a critical component for the next generation of reliable software engineering agents.