CoRE-RL: Co-evolving Reasoning Trajectories and Evidence Subgraphs with Learned Evidence Projection
Abstract
Large language models have made knowledge-graph question answering more flexible, yet complex multi-hop questions often require evidence that becomes identifiable only during reasoning. Starting from the question alone, narrow retrieval can miss bridge facts, while broad expansion injects distractors and worsens long-context interference. The key bottleneck is dynamic evidence identification: the system must revise a compact evidence view as reasoning exposes unsupported facts. We formulate this process as partially observable control over evidence and propose CoRE, a residual-coupled framework in which reasoning trajectories and evidence subgraphs evolve together. CoRE decomposes each trace into answer-critical claims, treats unsupported claims as residual signals, and uses them to guide targeted retrieval and budgeted subgraph projection. The same formulation exposes a learnable projection decision: after residuals retrieve candidate facts, CoRE-RL learns from rollout feedback which facts should remain visible for the next reasoning pass. Across WebQuestionsSP and Complex WebQuestions, the CoRE family improves answer accuracy, evidence sufficiency, and budget robustness, with CoRE-RL providing further gains through more effective retained-subgraph selection.