PointerBench: A Diagnostic Benchmark for Fine-Grained GUI Grounding
Valentino Sacco ⋅ Saurabh Pandey ⋅ Alexander F Jercher ⋅ Maximilian Schilling
Abstract
Multimodal Large Language Models have shown promising performance in computer-use tasks, from planning user workflows to grounding natural-language actions to precise UI targets. Existing end-to-end benchmarks primarily emphasize long-horizon task completion, while GUI grounding benchmarks remain focused on conventional controls and a limited set of professional interfaces. We introduce PointerBench, a diagnostic suite of 1,500 screenshot--instruction pairs probing three complementary capabilities: thin sub-element identification and geometric grounding in spreadsheets (PointerBench}-Sheets), character-level grounding in long-text documents (PointerBench-Text), and mixed GUI grounding across 100 professional applications (PointerBench-Pro). We carefully curate each task with pixel-exact target geometry and evaluate models of different scales and capabilities, ranging from lightweight systems such as Gemini 3.7 Flash to frontier models such as Claude Fable 5. We find text grounding to be the most challenging setting, with Claude Fable 5 and GPT-5.6 Sol achieving $69.6\\%$ and $53.4\\%$ success despite respective $83.2\\%$ and $77.7\\%$ average success across PointerBench. These results reveal substantial headroom in fine-grained GUI grounding, even for state-of-the-art models.
Chat is not available.
Successful Page Load