When Hallucinated Claims Become Harder to Retract: Tracking False Claims in Multi-Turn LLM Conversations
Abstract
Large language models (LLMs) can generate hallucinated factual claims that persist across subsequent turns and may become harder to correct as the conversation develops. While prior work has demonstrated hallucination snowballing and declining reliability in multi-turn conversations, less is known about how an individual hallucinated claim becomes embedded as a premise for downstream reasoning and how its conversational history affects its future trajectory. We introduce a controlled framework for studying hallucination cascades at the claim level. Starting from externally verified-false claims, we recursively generate conversational continuations under three matched interaction conditions: dependency-seeking, neutral, and verification-seeking follow-ups. Each resulting response is categorized according to the fate of the original hallucination as DROP, RETRACT, REPEAT, or DEPEND, distinguishing simple persistence from genuine propagation into downstream reasoning. We use the resulting branching conversation trees to investigate how interaction type is associated with hallucination propagation and correction, whether using a hallucinated claim as a premise is associated with its continued use in later reasoning, and how prior interaction relates to later repair. Across HalluHard and FactBench, verification following a dependency-seeking interaction produced lower retraction than verification following a neutral interaction in five of six model-dataset evaluations. Further analysis showed that verification after a false claim had remained active through repetition or further reasoning produced less retraction than verification after the claim had been dropped. This pattern held in every model-dataset evaluation. Together, our results show that later repair depends not only on whether verification occurs, but also on how the false claim was treated earlier in the conversation.