High-Dimensional Robotic Reinforcement Learning with Developing Synergies
Abstract
High-dimensional robotic reinforcement learning (RL) is often bottlenecked by inefficient exploration in large, redundant actuator spaces. We introduce Devyn (Developing synergies), a lightweight framework that learns a structured low-dimensional action parameterization directly during policy optimization. Unlike methods that require precomputed synergies or structured exploration in the full action space, Devyn starts from a randomly initialized latent-to-action map and shapes it online through task interaction. The policy acts strictly in a reduced latent space, while a learnable synergy matrix decodes latent actions into actuator commands. To make this evolving decoder stable and useful for policy learning, Devyn combines slow decoder evolution with structural regularization, reducing latent-action semantic drift while encouraging compact coordination patterns. Across diverse continuous-control benchmarks, including high-dimensional humanoid and overactuated musculoskeletal systems, Devyn improves upon standard SAC and achieves competitive or better performance than recent methods designed for high-dimensional robotic control. Ablations show that the gains do not come from a generic latent-action bottleneck alone, but rather from the slow, structured evolution of the synergy matrix. Our results suggest that learning the action parameterization itself can be a simple and effective route to scaling RL for high-dimensional robots.