DiReCL: Learning Differentiable Reward Code with Inverse Reinforcement Learning
Aoran Wang ⋅ Jingtao Zhang ⋅ Maosen Li ⋅ Zongzhang Zhang
Abstract
In reinforcement learning (RL), reward specification plays the fundamental role in shaping policy optimization. Inverse RL (IRL) offers a data-driven approach to inferring rewards from demonstrations, but prevalent neural-network-based methods often produce rewards lacking interpretability and flexibility for manual inspection or revision. In contrast, human-written or large language model (LLM)-generated reward code provides explicit semantic structure, yet its precise numerical parameters are often difficult to calibrate and prone to specification errors. To bridge this gap, we propose DiReCL (\textbf{Di}fferentiable \textbf{Re}ward \textbf{C}ode \textbf{L}earning), a novel framework that unifies these two paradigms by formulating reward learning in a differentiable reward code space. DiReCL first prompts an LLM to synthesize the reward code templates, then converts the numeric elements in these templates into differentiable parameters and align them with expert demonstrations through IRL. We further introduce a reflection mechanism that uses IRL diagnostics to provide evidence of inadequacies in the reward structure and guide template revision. Extensive experiments on MuJoCo locomotion and autonomous driving benchmarks demonstrate that DiReCL achieves over $1.25 \times$ the performance of both standard IRL and LLM-assisted reward design baselines, while preserving the crucial benefits of explicit, interpretable reward semantics.
Chat is not available.
Successful Page Load