From Preference Data to Personalization: Tracing Sycophancy in Large Language Models
Abstract
Sycophancy in large language models is often measured as agreement with the user, yet agreement can be appropriate when the user's claim is valid, their preference is genuine, or their experience warrants acknowledgment. We study sycophancy as inappropriate agreement: validating a claim, preference, or self-assessment when the context calls for correction, qualification, or resistance. We introduce a validity-aware sycophancy grader and SycoDrift, a persona-targeted evaluation set of 3,298 validity-annotated probes with controlled annotations for user openness and stakes. We use this framework to analyze three sources of agreement pressure: human preference data, automated judge specification, and in-context personalization. First, preference for validating responses in human datasets (Chatbot Arena, HH-RLHF, Community Alignment) is weak and concentrated where agreement is often appropriate. Second, automated LLM judges can introduce sycophancy by conflating prosocial values with agreement during reward specification; GRPO training amplifies this rubric-level error into model behavior. Third, inference-time personalization increases inappropriate agreement across 14 models, with the largest effect from matched user profiles placed directly before the query.