Sycophancy as Miscalibrated In-Context Updating
Abstract
Sycophancy in large language models refers to an excessive tendency to produce responses that cater to a user's false beliefs at the expense of sound judgment. This failure mode is often described as a form of social reward hacking, in which an LLM expresses a user-aligned response in anticipation of the user's approval. In this work, we present a series of observations that challenge this view and place sycophancy within the model's broader capacity for in-context updating, or revising its response based on information supplied during an interaction. We argue that sycophancy is best understood as poorly calibrated contextual updating, in which an unsupported assertion by the user receives more weight than it warrants and tips the balance toward the answer the user endorses. Behaviorally, we show that this accommodation of false claims is not specific to the user: models also move toward false claims asserted by other sources in the prompt, with larger shifts when the claim is attributed to a more reliable source, even when that source conflicts with the user. Mechanistically, we identify and remove a tiny set of parameters that causally support movement toward unsupported user claims, and find that this intervention also reduces appropriate movement toward credible new information, providing evidence that warranted and unwarranted updating partly rely on a shared parameter-level mechanism. Together, these findings suggest that sycophancy is closely connected to the same process by which models revise their responses to accommodate new information.