Stubborn or Sycophantic? Auditing What Sycophancy Evaluations Measure
Abstract
Anti-sycophancy evaluations often score whether a language model retains its answer after misleading feedback. That outcome does not identify selective correction, because a model that refuses all feedback avoids harmful reversals while also rejecting valid corrections. We audit this construct-validity gap by evaluating frozen system-prompt interventions under both correct and incorrect suggestions on the complete first-party SycoBench-600 protocol. Its exact-scored measures distinguish Update, when a baseline-wrong answer is corrected after a correct suggestion, from WrongFlip, when a baseline-correct answer is lost after a wrong one. Across 14 cells spanning pinned Phi, Llama, and Mistral revisions, a compact prompt on Llama lowers WrongFlip by 53.7 points but also lowers Update by 46.0, so its apparent gain is mostly answer inertia. The same prompt instead improves both components on Phi, while a selected-long prompt makes both worse on Mistral and lowers their difference by 22.7 points with a paired descriptive interval wholly below zero. No prompt improves correction selectivity on all three models, and six off-diagonal transfers show that prompt verdicts do not transfer uniformly. A reversal-only score can therefore reward resistance without establishing selective correction; evaluation should report valid correction and harmful reversal separately on each target model. Any safety benchmark that scores a single resistance behavior has the same gap, so behavioral evaluations should separate the harm they suppress from the capability they must preserve. Code: https://anonymous.4open.science/r/ncpe-5CDE/