Kernel Value Regression in Offline Reinforcement Learning
Jongyeon Lee ⋅ Jaehyoung Jeon ⋅ Kim Haneol ⋅ Myungjoo Kang
Abstract
Value overestimation is a central challenge in offline reinforcement learning (RL), arising when Bellman backups propagate out-of-distribution (OOD) errors that lead to catastrophic inflation of $Q$-values. Prior approaches address this issue by encouraging the learned policy to stay close to the behavior policy, but they often result in suboptimal policies or require significant training complexity. More recent work attempts to tackle this issue through minimal modifications to the critic network, yet it remains limited in either theoretical grounding or practical implementation. In this work, we present Kernel Value Regression (KVR), a simple critic regularization method based on kernel regression. KVR builds on the vanishing extrapolation property of kernel ridge regression, which encourages value estimates to decay toward zero away from the data support. This inductive bias naturally suppresses $Q$-values for OOD actions without auxiliary regularizers, providing a structural mechanism for addressing value overestimation. Using random Fourier features, KVR can be incorporated into standard actor-critic algorithms with only minimal modifications. We empirically demonstrate that these properties persist in complex offline RL environments, leading to strong performance across diverse tasks in the D4RL and OGBench benchmarks.
Chat is not available.
Successful Page Load