LLM judges are easier to talk down than up: a Bayesian item response analysis
Guoyang Zheng ⋅ Susan Wei ⋅ Sevvandi Kandanaarachchi
Abstract
LLM judges are increasingly used to grade student work. For quantitative problems like mathematics, a grade should reflect solution quality, not the language used in the submission. If LLM judges can be manipulated by that language, the process is unfair. We grade 450 faulty mathematics solutions under ten versions each, with six LLM judges, for $27{,}000$ gradings. Four persuasion techniques are each mirrored into a dissuasion technique that argues the grade down, and a length-matched control that asserts nothing is the reference for every effect. Using Item Response Theory we separate whether a cue changes a grade from which direction it moves. We establish the critical importance of having a control in such studies as without one the cue effect is mistakenly overestimated. We show that persuasion increased scores for only two of the six judges, whereas dissuasion reduced scores for all six; moreover, even the smallest dissuasion effect exceeded twice the largest persuasion effect.
Chat is not available.
Successful Page Load