Distributionally Robust Domain Randomization with Learned Risk-Sensitive Dynamics Samplers
Abstract
Domain randomization (DR) improves sim-to-real transfer by training policies over randomized simulator dynamics, but standard DR optimizes average performance under a fixed sampler and can miss rare failure-prone domains. We propose Risk-Sensitive Domain Randomization (RSDR), a distributionally robust formulation of episodic DR that evaluates each policy against the worst-case distribution over dynamics parameters within a KL neighborhood of a reference sampler. The resulting soft adversary has an exponential-tilting form, yielding a single temperature-controlled sampler family: negative temperatures emphasize low-return domains for robustness, zero recovers uniform DR, and positive temperatures give an optimistic sampler closely related to curriculum-style DR. Because this target changes with the policy and returns are observed only through rollouts, RSDR learns an amortized dynamics sampler by reverse-KL variational inference while reusing on-policy PPO trajectories. Across six domain-randomized MuJoCo Playground tasks, robust RSDR improves CVaR and minimum-return metrics over uniform and adaptive DR baselines, and outperforms best-tuned EPOpt.