When Does Graph Retrieval Become Answer-Supporting Evidence? A Diagnostic Audit of GraphRAG
Abstract
Graph retrieval-augmented generation (GraphRAG) is motivated by the idea that relational structure can help a language model find evidence that flat retrieval may miss. Yet retrieval scores and final-answer scores often compress a multi-stage pipeline into a single method row. We study the missing middle of this pipeline through an evidence-to-answer audit that follows each system from candidate availability, through evidence selection and context construction, to generated answers and cost. Across GraphRAG-Benchmark Medical and Novel tasks, we find a consistent conversion gap. Strong reranking, learned evidence scoring, and graph-based retrieval can improve evidence-side signals, but these gains do not yield significant answer-correctness improvements under the same generator and evaluator. Matched attribution controls show that some apparent graph-side gains are driven by reranking rather than graph expansion alone. Gold-evidence and bridge-aware interventions further show that evidence completeness is insufficient on its own; answerability also depends on salience, organization, and support in the generator-facing context. We therefore argue that GraphRAG evaluation should report the full evidence-to-answer conversion chain, including candidate source, selected context, generated answer, metrics, and cost, rather than relying on retrieval or answer leaderboards alone.