Simulating Human Judgment through Multiple Aspects: A Systematic Study
Abstract
LLM-based evaluators increasingly simulate human judgment in natural language generation (NLG) evaluation. When evaluating a target aspect, human evaluators often assess related auxiliary aspects and use the resulting judgments as supporting evidence, an intuition adopted by some recent LLM-based evaluators. However, the mechanisms behind these gains remain unclear: how to construct, score, and integrate this knowledge, how these benefits generalize across tasks and aspects, and how their scale affects evaluation. Existing methods require costly retraining or highly specialized architectures, hindering unified analysis of robustness and component contributions. To systematically study these questions for LLM-simulated human evaluation, we develop MAKScore, a simple, training-free, modular framework that constructs and evaluates auxiliary aspects for each target and integrates their evidence into a final score. We experiment with it across five NLG tasks and nine evaluation aspects, spanning dialogue, summarization, story generation, data-to-text, and machine translation. Results show that multi-aspect knowledge generally improves alignment with human judgments over direct LLM evaluation and conventional unsupervised metrics for overall and most aspect-specific settings. Ablations show that LLM-generated auxiliary aspects outperform predefined alternatives, auxiliary-aspect scores contribute substantially, and LLM-based integration outperforms direct averaging. Varying the number of auxiliary aspects reveals how knowledge scale affects performance. Overall, our study systematically examines how multi-aspect knowledge influences the validity of LLM-based simulation of human judgments in NLG evaluation.