When Alignment Fails: Stabilizing Cross-Dynamics RL with Prototype Trust Regions
Dong Uk Kim ⋅ Ji Su Yoon ⋅ Eui-Nam Huh ⋅ Choong Hong
Abstract
Alignment-based representation learning for cross-dynamics RL admits a degenerate global minimizer: per-domain feature covariances collapse to a common point, driving the alignment loss to zero while transfer regret remains large and rendering the standard Bures-Wasserstein (BW) transfer regret bound vacuous. We document this collapse across multiple environments, alignment kernels (BW, Frobenius, MMD), and methods (BW-CORAL, VGDF), where per-domain effective rank drops from p to $\approx 1$ within a few epochs—structurally the same failure as representation collapse in self-supervised learning. We propose a minimal remedy: a per-domain Prototype Trust Region (TR) that anchors each empirical covariance to a slow-moving EMA prototype, introduces no learnable parameters, and acts as an explicit second-moment stabilizer, mirroring momentum targets and covariance regularizers in SSL. Because TR depends only on per-domain covariances, it is combined with any BW-based method; adding it to VGDF inherits the same resistance to collapse. Theoretically, we prove a BW non-collapse lower bound that restores the transfer-regret guaranty to a non-vacuous form, and present a Representation Stability Transfer Bound that decomposes regret into coverage and stability terms TR directly controls. Empirically, TR delivers statistically significant gains under pre-registered tests in collapse-dominated regimes (finite-sample, lower-dimensional) and is correctly neutral when other bottlenecks dominate. We frame TR not as a universal improvement but as a targeted fix for an identifiable and broadly shared failure mode.
Chat is not available.
Successful Page Load