Evaluating the Evaluators: Investigating LLM Judges for Personalized Writing Style Assessment
Natalie Mackraz ⋅ Andrew Silva ⋅ Aini Putkonen ⋅ Brihi Joshi ⋅ Joelle Alcaidinho ⋅ Arnav Arora ⋅ Barry-John Theobald ⋅ Katherine Metcalf
Abstract
We evaluate how good large language model (LLM) judges are for personalized writing tasks. This is important when a user’s preferred writing style is known (e.g., ``upbeat’’) and an LLM judge is used to evaluate whether generated text adheres to this preference. We examine how the writing task, the combination of generator and judge LLM, the evaluation objective, and general commonsense and reasoning ability impact LLM judge performance. To this end, we generate text with four LLMs for three long-form writing tasks, collect human labels evaluating which of two responses is better, collect judge-LLM, pairwise-ranking annotations, and then compare the human and judge LLM annotations. We find that judge quality correlates weakly with general LLM ability measured using MMLU (Pearson's $\rho\leq0.2$) and varies by both writing task and evaluation criteria (i.e., general helpfulness versus style personalization). We also find that LLM evaluators are more consistent and performant using paired comparisons rather than scoring responses independently. Finally, we find that the ``strongest'' LLMs do not most reliably reproduce the human evaluations, which we hypothesize relates to exaggerated sensitivity to subtle details that human annotators and weaker LLMs ignore.
Chat is not available.
Successful Page Load