CS-DICE: Offline Reinforcement Learning with Coherent Occupancy Regularization
Jeremías Figueiredo Paschmann
Abstract
Offline RL methods often regularize toward the empirical occupancy of the dataset, but this reference can mismatch the deployed policy. In partially observed or decentralized settings, data may be pooled from hidden modes, histories, or conventions that are unavailable at execution time. The pooled occupancy can then require incompatible action laws for the same deployed input, or joint correlations that a decentralized policy class cannot represent, so regularization toward it can favor a low return projection. We formalize this failure through *reference coherence*: a reference is coherent when it does not rely on information hidden from the deployed policy, and when its action law is representable by that policy. We propose *Closest Slice DICE (CS-DICE)*, a DICE variant that regularizes toward a coherent slice of the training data rather than the pooled occupancy. A hard or soft selector chooses the slice used as the reference term, while the deployed observations, actor class, and factorization constraints remain unchanged. Most of the results focus on the soft reverse-$KL$ case, where the selector only changes the reference used by an otherwise standard DICE update. Controlled examples show that incoherent pooled references collapse to suboptimal behavior, while CS-DICE recovers high return policies when the selected or inferred slice is coherent with the deployed interface. We also identify a boundary case where pooling is benign because the relevant signal is revealed before the conflict. Finally, a CS-CoMA-DICE study on MaMuJoCo shows that selected references can be integrated into a neural cooperative MARL pipeline while preserving decentralized execution.
Chat is not available.
Successful Page Load