Exploration as Constrained Policy-Space Optimization
Abstract
Reinforcement learning (RL) agents face a continuous trade-off between exploiting known strategies and exploring novel states to obtain higher rewards. While exploration is often undirected, it can also be guided by intrinsic signals such as novelty or uncertainty. However, combining these signals with task rewards is challenging, as they can sometimes drive the agent away from optimal trajectories and reduce final performance. We find that this degradation arises from a structural occupancy conflict: as the policy improves, exploration signals naturally begin to anti-correlate with task advantages. In this work, we introduce Kantorovich Dual Shaping (KDS), a framework that resolves this conflict through constrained policy-space optimization. KDS formulates exploration as an Optimal Transport problem and applies the Sinkhorn-Knopp algorithm on the empirical batch manifold to reweight intrinsic rewards in a geometry-aware, point-by-point manner. When combined with standard RL objectives, it acts as a spatial filter that suppresses exploration in conflicting regions while enhancing it elsewhere. We evaluate this general approach across multiple off-policy and on-policy algorithms, achieving stable exploration that consistently avoids task degradation in complex continuous control environments.