Bootstrapped Bipartite Actor-Critic for Diffusion RL
Abstract
Diffusion policies have achieved strong performance in reinforcement learning (RL) due to the exceptional expressivity in capturing complex action distributions. However, existing approaches focus primarily on positive samples and lack explicit modeling of negative samples, thereby failing to fully exploit this representation capacity and ultimately limiting performance. To mitigate this problem, we propose BOBAC (BOotstrapped Bipartite Actor-Critic), a novel method for diffusion policy optimization via pairwise preference learning. Bridging the gap between distribution modeling and preference learning, we reformulate diffusion policy optimization within a preference-aware framework via the Bradley-Terry (BT) modeling, which enables the effective exploitation of negative samples through relative ranking. Since the original BT loss assigns the same weight to all pairs regardless of their varying importances, we further design a soft-gate mechanism to adaptively weight each pair according to its Q-gap. Experimental results on 6 MuJoCo tasks and 4 H1Bench tasks demonstrate that, compared with representative model-free and generative-policy RL baselines, BOBAC achieves state-of-the-art performance with consistent improvements in all continuous-control environments.