Trajectory-Consistent Diffusion Policies for Offline Reinforcement Learning
Abstract
Diffusion policies have emerged as a powerful policy class for offline reinforcement learning (RL) due to their ability to model complex, multimodal action distributions through iterative denoising. However, in offline RL, expressivity alone is not enough. Generated actions are supposed to remain within value-reliable regions of the action space to avoid extrapolation error. We find that standard diffusion policies often produce unstable denoising trajectories whose intermediate and final action candidates drift toward weakly supported regions, making critic-guided action selection unreliable and degrading teacher quality for one-step distillation. This issue is particularly pronounced on mixed-quality offline datasets, where suboptimal behaviors amplify trajectory instability. To address this problem, we propose Trajectory Consistency Rectification (TCR), a framework designed to promote denoising stability during both training and inference. At inference time, TCR aggregates locally consistent, high-value action candidates across denoising steps to suppress critic-unreliable outliers. During training, it adaptively reweights denoising steps to improve trajectory-level coherence. We show that this stabilization provides a more reliable teacher for one-step flow-matching distillation. Extensive experiments on D4RL and OGBench show that TCR consistently improves over diffusion-policy baselines, with significant gains in sparse-reward domains like AntMaze and Adroit. Moreover, one-step policies distilled from TCR achieve competitive performance among efficient single-step methods. These results suggest that enhancing the temporal consistency of denoising trajectories is a key ingredient for robust offline diffusion reinforcement learning.