Anchoring the Bridge Across Hops: Unlocking Latent Multi-Hop Reasoning under Knowledge Injection
Zichen TANG ⋅ Zhenheng Tang ⋅ Yifan Hou ⋅ Hanwen Xing ⋅ Xinda Qi ⋅ Xiaowen Chu ⋅ Bo Li
Abstract
Test-time training (TTT) keeps parametric knowledge current by writing updated facts into the parameters, but those facts are queried in multi-hop form at inference and their combinations cannot be enumerated during training. Whether a model trained only on the atomic form composes arbitrary chains within a single forward pass remains a critical question. We evaluate this on a two-hop benchmark of updated facts in which no training sample contains both hops of a chain. We show that latent composition remains hard for existing rewriting methods for knowledge injection (paraphrase, keyword diversification, and added context), which help the model acquire the updated facts but leave composed accuracy low (below $36$ at $1.7$B). We attribute this to a mismatched representation of the bridge entity between the two hops. We propose \textbf{Bridge Anchoring by Common Enrichment (BRACE)}, which reuses the same descriptor pool on the bridge in both hops to align its representation across them. Across four pretrained LLMs spanning two families and $1.7$B--$14$B and five domains, BRACE improves composed accuracy over the strongest baseline by $3$ to $31\%$ in relative terms and raises the composition rate, composed direct-QA accuracy restricted to the chains that answer at least half the paraphrases of both hops, up to $78$. A count-matched control attributes the gain in composed accuracy to the reuse of the descriptor across hops, and a logit lens places BRACE's resolution of the bridge $2.0$ to $5.5$ layers earlier in the stack. We further assess two other routes to elicit composition, chain-of-thought prompting and composition demonstrations on held-out entities, and find that neither reliably improves latent composition: prompting lowers composed accuracy under three of the four rewrites, and demonstrations transfer only on BRACE's training text. This paper presents a basis for compositional generalization under continual knowledge injection, where the chains queried at inference are by construction absent from the training data.
Chat is not available.
Successful Page Load