Ask, Then Grade: When Graded Scales Beat Binary Decomposition for LLM-as-Judge Evaluation
Abstract
Large language models are increasingly used to evaluate generated text, yet there is little agreement on how they should express their judgments. Recent approaches replace holistic scores with decomposed Boolean checks: CheckEval uses engineered checklists, while BinEval generates lightweight task-specific questions. We test whether such binary decomposition improves agreement with human ratings, or whether it discards information needed to evaluate graded, multifaceted qualities. Across SummEval, QAGS, and USR Topical-Chat, we compare single binary judgments, rubric-enhanced binary judgments, six-question BinEval-style decomposition, ordinary five-point ratings, and superlative-anchored scales using six judge models. Ordinary five-point ratings achieve the highest mean Spearman correlation across benchmarks for every judge. Their advantage over BinEval is statistically robust in 31 of 36 planned contrasts under conservative global correction. The advantage is not fully explained by nominal response resolution. The same aggregate ordering is observed after common binary reduction, within-source pairwise comparison, and cross-fitted monotonic alignment. A nominally matched seven-level comparison also favors direct graded judgment for all six judges, although these comparisons are descriptive. A matched five-level ablation shows that superlative endpoints reduce use of both scale boundaries without improving rank agreement or held-out monotonic alignment. Performance varies substantially across judge models, but the present design cannot separate variation in question generation, question answering, and aggregation. Lightweight binary decomposition remains useful for atomic verification and question-level diagnosis, but it is not a generally superior default for matching human ratings of holistic constructs.