Control First, Robustness Next: Decoupled Representation Learning for Visual RL Generalization
Abstract
Visual reinforcement learning (RL) is promising for real-world problems such as autonomous driving, robotic locomotion, and manipulation, but remains highly vulnerable to visual distribution shifts between training and deployment. Existing methods improve robustness through Q-consistency regularization, masking, and auxiliary objectives, yet most of them follow a coupled representation learning paradigm in which control-oriented and robustness-oriented representation learning are jointly optimized within the same training framework. We argue that this coupled structure can introduce harmful interference between control learning and robustness learning, degrading original-environment control performance and limiting generalization under visual shifts. To address this issue, we propose Separate Then Align Representations (STAR), which reformulates visual RL generalization as a problem of decoupled representation learning. Specifically, we first learn a control-relevant reference representation through online RL under weak augmentation, and then perform an offline post-representation alignment stage that maps strongly augmented inputs back to this reference representation. Experiments on RL-ViGen benchmarks spanning DMControl and Robosuite show that STAR achieves strong generalization under visual distribution shifts while better preserving performance in the original environment than prior coupled approaches. Our source codes are available in the supplementary material.