Interactive Combinatorial Reinforcement Learning for Knowledge Graph Reasoning
Abstract
Knowledge Graph Question Answering (KGQA) increasingly requires multi-turn interaction with knowledge graphs (KGs) to derive final answers. Such interactive capabilities often rely on proprietary large-scale LLMs, which hinders cost-efficient and local deployment. Reinforcement learning offers a scalable alternative for endowing smaller open-source LLMs with multi-turn interaction ability, as it enables models to improve from self-explored interaction trajectories rather than relying on expensive expert demonstrations. However, KGQA presents a distinctive challenge for RL: unlike tasks where exploration can proceed in a relatively unconstrained textual space, KG reasoning is governed by a sparse graph topology. Free-form policy rollouts frequently generate relation sequences that cannot be executed on the graph, making reward signals sparse and unstable. We propose ICOR, an Interactive Combinatorial Reinforcement Learning framework that grounds policy exploration in the combinatorial structure of KGs. ICOR first retrieves a query-specific set of candidate relations and then restricts policy rollouts to relation-path compositions within this graph-derived space. This design increases the likelihood that sampled trajectories are executable and thus informative for policy optimization. Building on this constrained rollout space, ICOR applies GRPO with a staged feedback mechanism that guides learning from structural validity to answer-consistent reasoning. Experiments on multiple KGQA datasets demonstrate the effectiveness of ICOR.