Distribution Corrected Decision Transformer for Offline Reinforcement Learning with Imbalanced Datasets
Abstract
Decision Transformer (DT) has emerged as a powerful paradigm for offline reinforcement learning. For complex tasks with imbalanced datasets, where high-quality samples constitute only a small fraction of the data while the majority is suboptimal, the performance of DT methods usually degrade significantly due to the the offline data distribution deviates substantially from the stationary distribution of the optimal policy. In this paper, we construct a dual-based correction mechanism to mitigate the distribution shift problem and propose a Distribution Corrected Decision Transformer method. Specifically, we investigate the Lagrangian duality of the reward maximizing objective in the standard Markov decision problem and recover the optimal correction weights between the static offline data and optimal stationary distribution. Note that a double-KL regularizer is integrated in the objective to prioritize high-return state-action pairs while remaining anchored to valid data support. Subsequently, we use the learned weights to correct distributional bias in two critical components of DT: stabilizing policy evaluation through weighted temporal-difference learning, and improving policy extraction by prioritizing DT training on weighted dataset aligned with the corrected target distribution. We theoretically analyze the performance guarantee, proving our method improves upon standard baselines, and empirically demonstrate strong performance on D4RL benchmarks, particularly on highly imbalanced datasets where prior methods fail.