Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering
Abstract
A central bottleneck in multi-hop Question Answering (QA) is that the granularity at which a question is expressed often differs from the granularity at which corpus evidence is retrievable. Existing methods address this mismatch either by imposing fixed graph structures over the corpus or by iteratively reformulating the query, but these strategies do not explicitly decide when a query unit is already supported by evidence and when it should be refined. We formulate this bottleneck as retrievable granularity discovery and introduce Hi-Q, an evidence-conditioned framework for hierarchical query refinement. At each query node, a resolution operator tests whether retrieved evidence supports the current query unit; resolved nodes terminate, while unresolved nodes are expanded by a dependency-preserving binary operator and checked by a semantic coverage verifier. Hi-Q therefore grows a query tree whose topology is determined by corpus support signals rather than by a fixed decomposition template or a pre-built graph. Across three multi-hop QA benchmarks, Hi-Q achieves 57.9 EM and 69.3 F1 on average, outperforming PropRAG, a graph-based RAG baseline, by 5.6 EM / 3.9 F1 and IRCoT, an iterative retrieval baseline, by 13.7 EM / 15.8 F1. Under full-corpus retrieval, Hi-Q maintains its gains with 53.4 EM and 65.2 F1 on average, improving over IRCoT by 16.3 EM / 19.4 F1 without corpus-wide graph construction.