Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment
Hamidreza H Balyani ⋅ Seyed P Davoudi ⋅ Alireza Amiri Margavi ⋅ Amin Gholami Davodi ⋅ Arshia Gharagozlou
Abstract
Large language models are increasingly deployed as advisors whose stated objective is not perfectly aligned with the user's. Whether such advisors stay truthful when honesty conflicts with their own payoff is a core alignment-evaluation question, and it is hard to answer because the correct answer is usually unknown. We turn the canonical Crawford--Sobel cheap-talk model into a pre-specified benchmark with an \emph{exact oracle}: game theory supplies not only the direction of the effect but the numerical target. A sender observes a state $\omega\in[0,1]$, wants the receiver's action close to $\omega+b$, and sends one costless message to a receiver whose ideal action is $\omega$. For the positive-bias grid $b\in\{0.01,0.04,0.08,0.12\}$ the most-informative equilibrium partitions have $7$, $4$, $3$ and $2$ cells, with exact $20$-bin normalized mutual information $0.5294$, $0.3268$, $0.2205$ and $0.1829$. Running the pre-registered design on four instruction-tuned models ($12{,}000$ logged sender calls), all four \emph{over-reveal} by $1.8$ to $4.2\times$: informativeness stays at $0.78$--$0.94$ where the equilibrium prescribes $0.18$--$0.53$. Informativeness declines with bias as predicted, but never approaches the strategic optimum; rather than coarsening into partitions, models exhibit near-full revelation with a constant upward offset that tracks their bias. The result for this workshop is the evaluator itself: the conclusion is recoverable only when the decoder reads the sender's stated number. A sentence-embedding decoder -- a defensible default, and the one a text-similarity judge would use -- mis-reads the same $12{,}000$ messages as near-babbling ($0.30$ versus $0.86$) and fails the zero-bias sanity check ($R^2=0.53$ versus $0.99$). Because the oracle is exact, we can show this is an evaluator failure rather than a modelling disagreement, and we argue that a cheap zero-conflict identity check should be a standard validity gate for automated evaluators.
Chat is not available.
Successful Page Load