Argus: A Cross-Regime Benchmark for the Transferability of Uncertainty Quantification in Computer-Use Agents
DIVAKE KUMAR ⋅ Sina Tayebati ⋅ Devashri Naik ⋅ Amanda Rios ⋅ Nilesh Ahuja ⋅ Omesh Tickoo ⋅ Ranganath Krishnan ⋅ Amit Trivedi
Abstract
Computer-use agents turn visual predictions into executable GUI clicks, so uncertainty estimates must support rejection, calibration, miss-severity ranking, and spatial safety regions. Existing post-hoc uncertainty quantification (UQ) evidence is fragmented across isolated model and dataset pairs, leaving unclear whether UQ rankings remain stable when the agent, benchmark, or observable interface changes. We present a cross-regime benchmark for post-hoc UQ in single-step executable GUI grounding, covering a 27-method, seven-family open-weight matrix over 4 GUI-grounding VLM agents and 4 datasets, plus an 8-method API-compatible closed-source matrix across 3 frontier vendors where logits, hidden states, and attention maps are unavailable. The main finding is selective transfer: UQ rankings are stable across datasets for a fixed model, but degrade across model classes and observable interfaces. Hidden-state and density methods form the most stable open-weight family, while CoCoA-1MCA, Focus, sampling-based scores, and verbalised self-assessment win in specific regimes. Ranking transfer is strongest within a fixed model across datasets, reaching Spearman $\rho=0.969$ and averaging $\rho=0.705$ over 120 open-weight pairs. In contrast, open-weight to closed-source transfer is the strongest break: cross-tier Spearman $\rho$ averages only $+0.08$ over 12 vendor x dataset pairs on the shared 8-method intersection, indicating that UQ recommendations do not survive removal of internal model signals. Model-class transitions further reshape UQ preferences: attention, verbalised, and VLM-native families lose AUROC on every dataset, while density methods remain stable; a scale-only baseline shows that logit-family degradation on ScreenSpot-Pro is partly confounded by scale rather than grounding fine-tuning alone. Finally, conformal click regions show that score-level discrimination is not enough for deployment: locally weighted disks can shrink radii by 40-60% when the plug-in UQ is calibrated, but coverage can degrade under calibration-test or interface mismatch. We release per-item records, calibration/test splits, UQ scores, and analysis scripts as a reproducible basis for regime-aware UQ selection in GUI agents.
Chat is not available.
Successful Page Load