Canopy: Tree-Aware Rollout Scheduling for Agent Reinforcement Learning
Abstract
Rollout generation bottlenecks large-scale RL training for LLM agents. Tree rollout has emerged as an important agentic-RL strategy: by branching from shared intermediate states across reasoning turns, tool calls, and environment observations, it avoids regenerating entire trajectories from scratch and exposes long prefixes for context reuse. However, existing RL and LLM-serving infrastructure remains largely tree-unaware: it treats sibling branches as independent generation requests, making the rollout tree invisible to routing, preemption, and KV-cache management. Consequently, shared prefixes are scattered across servers, discarded with request-private suffixes, or recomputed after cross-server spill. We present Canopy, a tree-aware rollout scheduling system that makes the active rollout tree a first-class serving abstraction without changing sampling, rewards, advantage estimation, or optimizer updates. Canopy consists of three mechanisms: tree-aware routing and retention, which keep siblings near reusable prefixes and protect active prefixes; prefix-preserving partial preemption, which separates shared-prefix KV from private-suffix KV under pressure; and spill-aware transfer-or-recompute admission, which transfers shared-prefix KV to a spilled sibling only when its estimated transfer cost is lower than destination-side recomputation. On long-horizon multi-turn agent workloads, Canopy achieves up to 2.11× rollout-generation speedup; across four representative tree-rollout algorithms, mechanism studies show higher prefix locality and a 51.1% average, 94.1% maximum reduction in full-release shared-prefix invalidations. These results show that exposing rollout-tree structure to the rollout-serving stack can reduce wasted prefill and KV-cache churn, enabling more efficient RL training for LLM agents.