SETA: Scaling Environments for Terminal Agents
Qijia Shen ⋅ Zhiqi Huang ⋅ Vamsidhar Kamanuru ⋅ Aznaur Aliev ⋅ Jay Rainton ⋅ Ahmed Awelkair ⋅ Boyuan Ma ⋅ Qizheng Zhang ⋅ Jiwei Fu ⋅ Yuzhen Mao ⋅ Wendong Fan ⋅ Ping Nie ⋅ Philip Torr ⋅ Bernard Ghanem ⋅ Changran Hu ⋅ Jonathan Li ⋅ Urmish Thakker ⋅ Guohao Li
Abstract
Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the terminal command line provides a text-based, general-purpose interface, covering tasks from system operations to data science and machine learning. However, scaling terminal-agent training remains challenging, as it requires diverse and coherent task instructions, executable environments, and reliable verification, while lacking naturally grounded supervision data. In this work, we propose SETA, a scalable framework for generating verifiable terminal environments for reinforcement learning (RL). The framework consists of two pipelines sharing a unified verification mechanism: SETA-Synth converts diverse sources into standardized RL environments, and SETA-Evol further expands from existing environments with adaptive control of difficulty and diversity. Together, we construct and release SETA-Env, the largest open-source verifiable terminal RL dataset to date, containing over $4{,}500$ environments. We evaluate our dataset by training a Qwen3-8B model with GRPO on SETA-Env, achieving $12$\% pass rate on Terminal-Bench 2.0, the best reported result for an RL-trained model at the 8B scale. The results demonstrate that SETA-Env provides high-quality training environments for terminal agents and serves as a valuable resource for advancing research on terminal-based agent learning.
Chat is not available.
Successful Page Load