Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
Abstract
Teaching large language models (LLMs) to reason during post-training typically relies on reinforcement learning with explicit outcome- or process-based reward functions. However, in many real-world settings, obtaining or defining such reward functions is difficult, especially for complex tasks, making learning from expert demonstrations an attractive alternative. The dominant approach, supervised fine-tuning (SFT), trains models to imitate expert reasoning traces directly, but suffers from the general limitations of off-policy learning: performance can be fragile to inference-time deviations from states explicitly covered by the demonstrations. To address this, we propose \textbf{Reasoning GAIL (ReGAIL)}, a GAIL-style adversarial reward-learning procedure for reasoning traces. Rather than directly imitating the expert's reasoning, ReGAIL learns a reusable, discriminator-derived process reward from expert chain-of-thought traces. Through experiments on GSM8K, MMLU-Pro and MedReason we show that the reasoning reward function learned with ReGAIL can be effectively used throughout the training and inference pipeline: (1) to provide a training signal for \textbf{post-training}, outperforming SFT in most of the considered settings and matching or exceeding a source-based adversarial distillation baseline in a targeted comparison, (2) for \textbf{inference-time reranking}, improving pass@1 by up to 8.7 points, and (3) for \textbf{process-level evaluation}, improving natural-error localisation by up to (+10.8) MAP points over token-probability baselines. Overall, ReGAIL bridges imitation learning and reward-based optimisation, enabling the extraction of meaningful reasoning signals from expert thinking traces.