Beyond Proxy Metrics: MLLM-Based Human Surrogate Evaluation for Explainable AI
Abstract
Explainable AI (XAI) offers far-reaching consequences for human-AI collaboration, but the fundamental question of how best to evaluate these systems has proven frustratingly elusive. Methodologies typically revolve around either proxy metrics which tell us little about the utility of an explanation, or human evaluation which---while generally considered the gold standard---is expensive, difficult to scale, and challenging to reproduce. In this paper, we investigate if Multimodal Large Language Models (MLLMs) can act as surrogate human participants to automate the evaluation of explainable AI. To do this, we first propose a ``purpose-grounded'' evaluation framework with specific metrics applicable to both MLLMs and human users alike. Then, we offer some conceptual insights as to why and when we can trust these models to automate this process. Lastly, we conduct extensive empirical testing across five purpose-grounded areas and nine popular MLLM APIs. Results indicate that all models trend positively, but OpenAI models (particularly GPT-5 mini) are the best aligned at replicating user responses, and for complex visual reasoning tasks more advanced modes (like GPT-5) perform best, which would be the recommended choices in practice. We expect our insights not to replace user testing, but rather to help XAI researchers complement their current evaluation methodologies and move the field away from heavy reliance on proxy metrics for measuring explanation utility.