Controlled Perturbation Reveals Grounding and Faithfulness Failures in Computer-Use Agents
Abstract
GUI grounding models report over 85% accuracy on standard benchmarks, yet1 drop 27–56 percentage points when instructions require spatial reasoning rather2 than direct element naming. Current benchmarks miss this because they evaluate3 each screenshot once with a single fixed instruction. We introduce GUI-CP1,4 a controlled perturbation framework that independently varies visual scenes and5 instructions to measure grounding robustness. Evaluating three 7B models from6 the same SoTA architecture lineage, we find that relational instructions cause7 systematic accuracy collapse across all models, mild adjustment to browser zoom8 produces statistically significant degradation, and rank-8 LoRA fine-tuning with9 augmented data degrades performance rather than improving it, suggesting system-10 atic overfitting. By perturbing along independent axes, GUI-CP isolates which11 specific capability axes are affected—spatial reasoning, visual robustness, reason-12 ing calibration—providing diagnostic signal that aggregate benchmarks cannot. We13 release the augmentation pipeline and evaluation harness, from which the dataset re-14 generates deterministically; the rendered dataset and fine-tuned checkpoints follow15 on publication.