Mind2Dialogue: Training Human-Aware Language Models through Shared-State User Simulation
Abstract
Personalized assistants must infer users' unobserved beliefs, goals, affect, and social relationships from interaction histories, yet training data rarely provide direct supervision for these latent states or their evolution. We introduce Mind2Dialogue, a scalable simulation framework that generates user–assistant conversations conditioned on evolving structured user states. These states shape both sides of the dialogue: they condition each user message and are observed by an Oracle assistant when generating responses. We then train a student model to reproduce Oracle responses from static persona information and dialogue history, enabling it to approximate state-conditioned behavior without access to the evolving state at inference time. Without requiring per-dialogue human annotation, the pipeline produces M2D-Corpus, which contains synthetic multi-turn dialogues and QA pairs in which assistant responses and answers are grounded in the corresponding latent user states. Fine-tuning Qwen2.5-7B-Instruct on M2D-Corpus improves PrefEval-Gen by 33.4 absolute points and PersonaMem-v2 by 10.0 points over the corresponding base model; the resulting model outperforms memory-augmented and personalization baselines, and the personalization gains hold across three model families. Although trained without explicit Theory-of-Mind supervision, the Qwen and Llama models also improve on ToM benchmarks. Together, these results suggest that personalization and Theory of Mind may rely on a shared capability: inferring latent user states from observed interactions.