Warp RL: Reshaping Base Policy Distributions for Dynamics Adaptation
Abstract
Robot policies that perform well under nominal training conditions can degrade after deployment when dynamics, contact, calibration, or hardware differ from the source domain. We introduce Warp RL, a lightweight output-space adaptation method that reshapes sampled actions from a frozen stochastic base policy rather than retraining the policy itself. In controlled simulation, we compare a translation-only residual, a positive affine transform, and Warp, which applies a state-conditioned monotone spline directly to sampled base actions. Across four manipulation tasks with held-out dynamics shifts, Warp achieves the strongest aggregate recovery under evolution strategies (ES); matched spline and MLP controls indicate that its structured transformation matters beyond sample access or generic nonlinear capacity. A complementary policy-gradient study compares Warp and residual adaptation against full-policy and last-layer fine-tuning, showing that frozen output-space adaptation remains competitive while Warp's relative benefit depends on the optimization regime. Standardized action marginals further confirm that trained Warp policies use non-affine reshaping. Finally, on physical SO101 peg insertion under an off-policy residual adaptation protocol, Warp achieves comparable success to a strong stochastic residual while reducing median successful cycle time by approximately 30\%. Together, these results support structured output-space transformations as a practical approach to adapting frozen robot policies under deployment-time dynamics shift.