Prevailing Bisimulation Metric Learning Is Biased: Implicit Regularization and Its Remedy
Abstract
Deep bisimulation metric learning has emerged as a principled framework for learning robust state representations in reinforcement learning. However, prevailing methods rely on sample-based regression objectives that exhibit a critical stability flaw in stochastic environments. We identify this flaw as a manifestation of the double sampling problem: minimizing mean squared error against single-sample stochastic targets introduces an irreducible variance bias. Through the lens of stochastic differential equations (SDEs), we rigorously prove that this bias leads to a Variance Trap, wherein target variance scales with the encoder's sensitivity. This variance effectively acts as an implicit Jacobian regularizer, driving representation collapse in the presence of transition noise. To address this issue, we propose Bisimulation Saddle-Point Optimization (BSPO), a debiased framework that reformulates the bisimulation metric learning objective as a primal-dual saddle-point problem. By introducing an Auxiliary Network (AuxNet) to estimate the expected Bellman error, BSPO decouples the structural error from transition noise, eliminating the bias without requiring physically infeasible double sampling. To the best of our knowledge, this work is the first to identify that prevailing sample-based bisimulation metric learning objectives are systematically biased and to provide a practical debiasing framework. Our empirical evaluations on stochastic continuous/discrete control tasks and visually distracting environments demonstrate that BSPO significantly outperforms representative baselines, particularly in high-noise regimes where conventional methods fail to converge.