Contextualized Evaluation of Vision Language Models through Dynamic Interviews
Abstract
Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static settings and the dynamic, interactive, and contextualized nature of real-world applications. To bridge this gap, we propose CEDI (Contextualized Evaluations of MLLMs through a Dynamic multi-round Interview, a framework that recasts evaluation as a three-party interaction between an evaluatee model, an automated examiner, and a grader. The examiner conducts multi-turn, semi-structured interviews guided by a graph-based representation of the task. By navigating state-space transitions, CEDI deploys diverse strategies—from clarification requests to adversarial probes—to elicit performance evidence.