When Evaluation Protocols Outlive Their Assumptions: Stress-Testing Social Desirability Bias Measurement Across LLM Generations
Kensuke Okada ⋅ Yui Furukawa
Abstract
Evaluation protocols are often reused across model generations, although changes in the systems being evaluated may alter what those protocols measure. We examine this problem through the $Q_n$ protocol of Salecha et al. (2024), where $n$, the question batch size, is the number of questionnaire items presented together in a single request and $Q_n$ denotes the corresponding experimental condition. The protocol quantifies social desirability bias (SDB) from score shifts across conditions with different values of $n$. We conducted a replication of the original questionnaire-response and source-identification experiments using two LLMs also examined in the original study and stress-tested the same protocol on seven newer LLMs. For the two previously studied LLMs, the original qualitative pattern was largely replicated: large socially desirable score shifts and source-identification rates first reaching .90 in the $Q_5$ condition. Most newer LLMs, however, identified the evaluation context from fewer items, with the .90 threshold often occurring between the $Q_2$ and $Q_4$ conditions, while their $Q_n$-induced score shifts were generally smaller and less consistently aligned with social desirability. Together, these findings show that $Q_1$ no longer serves as a consistently low-identification baseline across newer LLMs, limiting the $Q_n$ contrast's interpretation as a stand-alone measure of SDB. More broadly, our results provide an empirical case for periodically revalidating the assumptions underlying AI evaluation protocols as models evolve, because a stable procedure can continue to yield comparable numbers without supporting the same diagnostic interpretation.
Chat is not available.
Successful Page Load