Reproducing ELEPHANT: Judge-Dependence and Model Deprecation in an LLM Sycophancy Benchmark
Kristina Sakayeva ⋅ Ana Trisovic
Abstract
Behavioral benchmark scores for large language models (LLMs) are produced by a fragile measurement chain---LLM judges applied to responses from commercial models that drift and disappear---yet no such benchmark has been reproduced end to end. We conduct a full reproduction of ELEPHANT, a benchmark of social sycophancy, one year after publication, recomputing every evaluation from raw responses across its four sycophancy dimensions, four datasets, and eleven evaluated models, and extend it with a sensitivity audit over choices the original leaves undocumented. The reported results reproduce with characterized caveats: the OEQ sycophancy gap, AITA-YTA face-preservation effect, ALP failure-to-challenge rate, and moral both-sides finding hold at the aggregate level, while some model-specific rankings change. Measured rates are nevertheless contingent on undocumented choices---switching the judge produces changes of $-0.16$ to $+0.23$ in the mean sycophancy rate, depending on the dataset--dimension combination---and the benchmark's apparatus is perishable: within a year, five of the eleven evaluated models and one of the paper's own robustness-check judges were retired or no longer accessible through the original provider. We introduce this as the LLM behavior reproduction window, distill reporting recommendations, and release our code, per-example scores, and cached model outputs at https://github.com/Kristina-Sakayeva/elephant-reproduction
Chat is not available.
Successful Page Load