Behaving Better, Thinking Worse: Sycophancy Across Post-Training Stages
Sonnet Xu ⋅ Kritika Singh ⋅ Sheharbano Jafry ⋅ Roxana Daneshjou ⋅ Sanmi Koyejo
Abstract
We trace factual sycophancy across three post-training pipelines (OLMo 3 7B Think and Instruct, Llama 3.1 8B Instruct, Tulu 3) on AMPS and MedQuAD with a dual-track evaluation: a GPT-4o-judged generative track and a log-probability track reporting $\Delta$LogOdds separately under \emph{in-context} rebuttals and \emph{preemptive} wrong-answer assertions. Decomposing by challenge type $\times$ context reveals three patterns. (1) \textbf{Challenge type:} the sycophantic shift is alternative-conditioned, so simple pushback with no asserted answer produces no shift, while ethos/justification/citation challenges produce substantial positive shifts. (2) \textbf{Context:} every base model's sycophantic shift concentrates preemptively, with weak or defensive in-context response. On computational, IC diverges into four pipeline-specific endpoints. On medical, the same four pipelines converge with no defensive IC developing on any recipe. (3) \textbf{Behavioral vs log-probability:} post-training produces a \emph{surface-predictive dissociation}: matched-subset behavioral flip rate drops significantly on identical items while preemptive $\Delta$LogOdds on those same items grows.
Chat is not available.
Successful Page Load