Statistical Analysis of Inverse Entropy-Regularized Reinforcement Learning
Abstract
Inverse reinforcement learning (IRL) seeks to recover a reward function from expert state--action trajectories. The problem is non-identifiable because multiple rewards may induce the same optimal policy. We study \emph{Inverse Entropy-regularized RL} and resolve this ambiguity by selecting a canonical least-squares reward through the soft Bellman residual. Expert demonstrations are modeled as a Markov chain with invariant distribution induced by an unknown policy (\pi^\star), which is estimated by penalized maximum likelihood over a class of conditional action distributions. We derive high-probability bounds for the excess expected conditional KL loss in terms of covering numbers of the policy class and transfer them to reward error, obtaining non-asymptotic minimax-optimal rates that display the dependence on entropy smoothing, model complexity, and sample size. For an unknown transition kernel, a Galerkin approximation combined with two-timescale linear stochastic approximation yields a trajectory-based estimator with separate policy-estimation, operator-approximation, and stochastic-approximation error terms.