Implicit Value Probing: Inferring Human Value Preferences via Strategic Multi-Turn Conversations
Abstract
Recognizing human value preferences is crucial for building reliable AI agents in user-centric fields like psychological counseling and strategic negotiation. Existing benchmarks rely on explicit questionnaires, assuming (1) users are always cooperative and willing to self-report; (2) value preferences remain static. However, humans may conceal true preferences if they feel probed, and may adopt distinct value preferences depending on the context. Consequently, explicit queries are insufficient for predicting value preferences reliably. To address this, we introduce Implicit Value Probing (IVP), a benchmark simulating realistic scenarios where agents implicitly infer value preferences through strategic, multi-turn conversations. IVP grounds these interactions using a diverse set of Persona Agents, covering a wide spectrum of personalities and values. We evaluate mainstream LLMs on IVP, including Qwen3, DeepSeek, Doubao, and GPT-5, observing substantial performance gaps. We further develop Value-Prober-8B, a specialized model internalizing structured cognitive reasoning via distillation. It employs strategic questioning to elicit concealed value preferences. Experimental results demonstrate that Value-Prober-8B can effectively balance probing accuracy with social acceptability, achieving capabilities comparable to frontier models.