Content-Addressed Workspace Branching for Agentic RL Rollouts: What the Blob Store Alone Does Not Buy You
Dipankar Sarkar
Abstract
Agentic reinforcement learning for coding agents repeatedly forks mutable workspaces, and recent systems have already shown that filesystem-level branching itself can be made cheap. We study a narrower question: what happens when a content-addressed, copy-on-write workspace store is used as the environment backend of an agentic-RL trainer? We integrate one environment interface with SkyRL end to end and wire the same core into verl's rollout loop, without establishing reward-path execution in the latter. Three results are negative and actionable. First, a per-branch full-text index duplicates file content and destroys metadata-only fork scaling; deferring it holds per-vault metadata at 209 KB across a $64\times$ range of content size and cuts median fork latency at 64 MiB from 106.8 to 4.2 ms. Second, cross-branch deduplication reduces storage as sibling edits converge, but at $G=32$ linear interpolation places break-even against a shared-base file-granular baseline at 83.6% convergence, because per-branch metadata offsets the saving below that point. Third, on a sequential, provisioning-heavy synthetic workload, forked workspaces have roughly an order of magnitude lower per-environment lifecycle latency than container baselines, including a warm pool, but do not beat git worktree. Integration also surfaced three failure modes silent in the trainer's own output: a tool-call parser mismatch, an evaluation harness that manufactures plausible learning curves from random task sampling, and a filesystem truncation defect. Cheap workspace branching depends as much on metadata, indexing and integration semantics as on the blob store underneath.
Chat is not available.
Successful Page Load