Towards Dissociating Acceptance and Belief in LLMs
Abstract
In conversation, a listener may accept or agree with a speaker's statement for the purposes of the conversation without actually believing it. While this belief-acceptance distinction has been well explored in other fields, including neuroscience and social psychology, we make the case for interpreting LLM behavior through this lens. We operationalize this distinction to analyze LLM sycophancy in a cooperative task, where an LLM erroneously agrees with false statements despite having sufficient evidence to the contrary. Using truth-directional probes to track model representations in goal-oriented dialog, our results indicate a distinction between instances of model acceptance with and without a corresponding update to its truth representation. While incorrect acceptance is well studied in the form of sycophancy, incorrect belief updates potentially signal a more critical failure mode, where models, in the process of agreeing with the user, end up "believing" a false statement. We find that in cases of erroneous acceptance without a belief update, models recover the truth of a proposition later in conversation significantly more often than in cases with a belief update (81% vs. 16%). We make the case that delineating belief from acceptance offers a valuable lens for distinguishing identical model outputs caused by differing internal states, and that researchers should be able to make this distinction in service of making LLMs safer.