How, Not How Much: Decomposing Sycophancy in Large Language Models
Abstract
Sycophancy, a language model's tendency to agree with a user's opinion even when it is incorrect, is an alignment failure that can be exploited to manipulate model behavior. While the naive sycophancy rate (the fraction of responses that agree with a user-endorsed answer under an opinion cue) is a convenient summary statistic, it conflates behaviors with very different implications, e.g., abandoning a correct answer to agree with the user is fundamentally different failure from merely restating an answer that was already incorrect. In this work, we decompose the naive sycophancy rate into four mutually exclusive components, partitioned by the model's plain answer. We show that these components distinguish harmful knowledge-override from agreement that overrides originally incorrect or invalid responses, giving a behavioural picture of not just how much a model agrees with a user, but more importantly, how it agrees. Evaluating 19 language models on our variant of the MMLU multiple-choice benchmark, we show that models with similar naive sycophancy rates can differ sharply in their underlying components, and that the aggregate rate can even rank a more knowledge-overriding model as less sycophantic. We further demonstrate the value of the decomposition by applying it to prompt-framing factors, including opinion position, grammatical person, and claimed user expertise, and find that the components respond to these factors differently from the naive sycophancy rate. Our results show that a single naive sycophancy rate is insufficient to understand or compare models' sycophantic behavior, and that our decomposition provides a more informative and interpretable metric for this purpose.