Inverse Reinforcement Learning with Just Classification and a Few Regressions
Lars van der Laan ⋅ Nathan Kallus ⋅ Aurelien Bibaut
Abstract
Inverse reinforcement learning (IRL) seeks to recover a reward function from observed behavior, but rewards are typically only partially identified: many reward--value pairs can induce the same behavior policy. We study this problem in the maximum-entropy, or Gumbel-shock, model under a broad class of statewise affine normalization constraints, with anchor-action constraints as a special case. This leads to Generalized Policy-to-$Q$-to-Reward (GenPQR), a modular approach to normalized reward recovery based on policy estimation and $Q$-evaluation via the Bellman equation; both components can be instantiated with off-the-shelf classification and regression methods. We establish modular finite-sample guarantees under general function approximation, with separate terms for policy-estimation and $Q$-function estimation error. As a concrete instantiation, we study GenPQR with fitted $Q$-evaluation, reducing IRL to policy estimation followed by regression. In experiments, GenPQR improves on DeepPQR in reward recovery while remaining simpler and more modular. Relative to DeepPQR, our theory is broader: it goes beyond anchor actions, accommodates large action spaces, and is not tied to a specific neural-network architecture or training procedure.
Chat is not available.
Successful Page Load