ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
XuHao Hu ⋅ Xi Zhang ⋅ Haiyang Xu ⋅ Kyle Qiao ⋅ Jingyi Yang ⋅ Xuanjing Huang ⋅ Jing Shao ⋅ Ming Yan ⋅ Jieping Ye
Abstract
Computer Use Agents (CUAs) can act through both atomic GUI actions ($\textit{e.g., click, type}$) and high-level tool calls ($\textit{e.g., API-based file operations}$), but they are often confused by this hybrid action space: they do not know when to continue with GUI actions and when to switch to tools, and finally fail to select the optimal execution path. We view this orchestration problem as GUI-Tool path selection: deciding when the agent should continue with GUI actions and when it should switch to tool calls to form an effective execution trajectory. This difficulty stems from two issues. First, high-quality interleaved GUI-Tool trajectories are scarce, and collecting real tool trajectories is expensive and brittle. Second, existing supervision provides limited guidance for GUI-Tool path selection, as most methods focus on step-level action imitation or final task completion and offer little trajectory-level feedback on whether GUI-Tool switching leads to a more effective execution path. In this paper, we propose $\textbf{ToolCUA}$, an end-to-end agent designed to learn optimal GUI-Tool path selection through a staged training paradigm. We first introduce an $\textbf{Interleaved GUI-Tool Trajectory Scaling Pipeline}$ that repurposes abundant static GUI trajectories and synthesizes a grounded library of tools, making it possible to scale diverse GUI-Tool trajectories without manual engineering or real tool-trajectory collection.Based on this data, we perform Tool-Bootstrapped GUI RFT, which combines warmup SFT with single-turn RL to improve decisions at critical GUI-Tool switching points. Finally, we further optimize ToolCUA with $\textbf{Online Agentic RL}$ in a high-fidelity GUI-Tool environment, using a Tool-Efficient Path Reward that encourages both appropriate tool use and shorter execution paths. Experiments on OSWorld-MCP show that ToolCUA achieves 46.85% accuracy and outperforms the baseline by over 60% relatively, establishing a new state of the art among models of comparable scale. It also improves by 3.9% over GUI-only settings, demonstrating effective GUI-Tool orchestration. The results further suggest that training in a hybrid action space is a promising paradigm for real-world digital agents.
Chat is not available.
Successful Page Load