Counterfactual Distillation: Internalizing Reflective Experience into LLM Agent
Rui Ge ⋅ Yichao Fu ⋅ Yu-Yang Qian ⋅ Junda Su ⋅ Yiming Zhao ⋅ Hao Zhang
Abstract
Large language models are increasingly deployed as autonomous agents that plan and act over long-horizon interactions, learning from both final outcomes and rich environmental feedback. However, optimizing from such interaction experience remains challenging, as feedback is heterogeneous and rarely carries explicit reward information. Existing post-training approaches rely on outcome-driven methods, such as reinforcement learning with verifiable rewards (RLVR), which primarily exploit final success signals while leaving interaction experience and feedback underutilized. In sparse-reward, long-horizon tasks, this often results in **distribution sharpening**: the policy reinforces a narrow set of already-successful behaviors, without substantially improving the feedback-grounded agency needed for broader problem-solving capacity (e.g., Pass@$k$). We propose ***Counterfactual Distillation***, a framework that converts exploration-derived experience into trainable policy supervision. Specifically, we organize exploration as a tree-structured search process, where the agent reflects on past decision points and makes new experience-guided decisions to form alternative branches. These corrections are then distilled as counterfactual targets: we train the model to predict the revised action from the original history without explicitly providing the experience, thereby introducing out-of-distribution behavior patterns and expanding exploration capacity. Across diverse interactive coding and agentic tasks, our method outperforms outcome-driven baselines such as GRPO and experience-based methods such as Early Experience, with gains of up to **14%** on Pass@128. When interleaved with reinforcement learning updates, it further raises the performance ceiling, yielding over **10%** improvement in Pass@1. We provide an [anonymized implementation](https://anonymous.4open.science/r/Counterfactual-Distillation-0E71).
Chat is not available.
Successful Page Load