Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap
Abstract
Large Language Model (LLM) interactions are typically underspecified, with users clarifying all necessary details across multiple conversational turns. Yet recent work shows that LLMs perform far worse in this multi-turn setting than in a single turn with all information being available at once, a phenomenon termed ``Lost in Conversation.'' However, bridging this gap effectively and generally remains an open problem. Here we introduce Found in Conversation (FiC), a training framework where a model teaches itself to find back its single-turn competence given underspecified multi-turn prompts. We develop View-Asymmetric Self-Distillation, which runs the same self-distillation backbone on two views of the same information: the teacher sees a single-turn view that concatenates all information revealed across the conversation turns, and the student sees the multi-turn view itself; distillation aligns the student's weaker multi-turn behavior with the teacher's stronger single-turn behavior. Across diverse model sizes (3B–14B) and architectures (Llama, Qwen, Phi, and OLMo), ours achieves a 100% recovery of the single-turn performance on 2 Llama models and recovers over 90% on every model, delivering more helpful and efficient multi-turn conversations without compromising single-turn performance.