Explorer-Zero: Self-play from Environment Exploration for Long-Horizon Tasks
Abstract
Training LLM agents on long-horizon tasks with reinforcement learning requires human-annotated tasks with reliable verification, which are expensive to construct for stateful environments. Self-play, in which a proposer generates tasks for a solver, offers a way around this bottleneck, but existing proposers generate tasks in a single shot from static artifacts, yielding tasks that are often unachievable, leave most of the environment unexplored, and provide only a sparse binary reward. We propose Explorer-Zero, a self-play framework in which the proposer explores the environment before committing to a task, declaring subtasks through tool calls that snapshot the environment as reference states. Completion is verified by matching the solver's states against these snapshots, so every task is solvable by construction, and the subtasks form a checklist that supplies dense rewards. On ALFWorld and AppWorld, without any human-annotated tasks, Explorer-Zero raises the success rate of a Qwen3-4B agent from 0.17 to 0.84 and from 0.15 to 0.37, recovering most of the performance of reinforcement learning on the full human-annotated training sets (0.96 and 0.55).