Code Agents Automate the Two Manual Bottlenecks of Sim-Based Robot Learning: Digital Twin Scene Construction and Demonstration Collection
Abstract
Sim-based robot learning is throttled by two manual processes: constructing realistic digital twins of scenes, and collecting demonstrations by teleoperation. Code agents are gaining real momentum as embodied agents for zero-demonstration trajectory generation. Here, I first show that an off-the-shelf code agent can now automate visually realistic 3D scene generation with surprisingly high efficacy—to my knowledge, the first photo-referenced digital twins built by an unmodified code agent in a robot simulator—using outdoor agriculture as the pilot domain, where deployment must be proven against field-specific variations before any robot enters the field. My second contribution tests whether code-agent trajectories can train a control-rate policy under the demonstrated task conditions. A code agent spends 5–20 s of inference per conversation turn, far too slow for a deployed controller. I therefore evaluate using a code agent to seed demonstrations in a canonical behavior-cloning pipeline. The resulting perception-to-action policy reaches 56% in-distribution success on the 14 demonstration fruits. Together, these results suggest that a generalist code agent can relieve both manual bottlenecks of sim-based robot learning.