GUI-Libra: Data-Efficient Post-Training for Reliable Reasoning-and-Acting in Native GUI Agents
Abstract
Open-source native GUI agents still lag behind closed-source systems on long-horizon tasks. One reason is the direct reuse of generic post-training pipelines that ignore GUI-specific failure modes: standard supervised fine-tuning (SFT) with long chain-of-thought (CoT) reasoning often degrades grounding, and stronger offline optimization in RL does not necessarily translate to better online performance. This offline-to-online mismatch arises in part from \emph{partial verifiability}: multiple actions may validly advance a task, but supervision typically marks only one demonstrated action as correct, creating reward ambiguity. We present \textbf{GUI-Libra}, a data-efficient post-training recipe for reliable reasoning-and-acting in native GUI agents. GUI-Libra combines a construction and filtering pipeline for a curated 81K GUI reasoning dataset, \emph{action-aware supervised fine-tuning} that mixes reasoning-then-action and direct-action supervision with action-aware token reweighting, and KL-constrained RL with success-adaptive scaling to improve offline-to-online predictability under ambiguous rewards. Across web and mobile benchmarks, GUI-Libra consistently improves both step-wise accuracy and end-to-end task completion. GUI-Libra-4B and GUI-Libra-8B improve their base models by +15.6\% and +12.2\% on AndroidWorld, +4.0\% and +8.7\% on Online-Mind2Web, and +12.5\% and +11.3\% on WebArena-Lite-v2. These results show that careful reasoning data curation and tailored post-training can substantially improve long-horizon task solving without costly online data collection. We release our dataset, code, and models to support future research.