HiddenPathQA: A Benchmark for Knowledge Graph Question Answering When Questions Hide Their Paths
Abstract
Knowledge graph question answering (KGQA) systems are typically evaluated on questions whose surface tokens lexically expose the gold KG path. Real users, however, ask questions without knowing the underlying schema: "Which condition could Andrew have due to family history?" compresses a 4-hop chain (Andrew - father → father → medical condition → subclass of - hereditary disorder) into a single phrase (family history) whose tokens name none of those relations. We show that this collapse of path-question alignment is a blind spot of existing KGQA evaluation. The three dominant Large Language Model (LLM)-KG paradigms — hop-by-hop KG exploration, plan-then-retrieve, and similarity-based retrieval — all implicitly rely on the question's surface leaking the gold path, an assumption rarely tested in practice. We make three contributions. (i) HiddenPathQA, a human-validated KGQA benchmark of 1,055 multi-hop questions (2-6 hops, multiple domains) constructed so that the question surface does not leak the gold KG path's relations. (ii) Holistic Path Selection, a reasoning paradigm that exposes all candidate relation chains around the topic entity to the LLM at once, replacing per-hop relation scoring with question-against-full-chain alignment. Two implementations, TieredHolistic and FlatHolistic, outperform six strong baselines (ToG, PoG, FiDeLiS, R2-KG, KAPING, KARPA) on HiddenPathQA by +8.8 to +17.9 points over the strongest baseline across all four (backbone × split) settings. (iii) A pseudoword stress test that replaces entity surfaces with placeholder tokens. It both confirms HiddenPathQA items are solvable from KG structure alone — not from entity-level priors — and reveals that several baselines rely on entity memorization rather than KG-grounded reasoning.