MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
Abstract
Recent progress in deep research systems has been impressive, yet their evaluation remains limited not only by benchmark coverage but also by methodology. Existing benchmarks predominantly assess final reports using fixed rubrics, failing to evaluate the underlying research process. Most also offer limited multimodal coverage, rely on synthetic tasks that do not reflect real-world query complexity, and cannot be refreshed as knowledge evolves. To address these gaps, we introduce MiroEval, an evaluation framework for deep research systems, accompanied by a 100-task benchmark (70 text-only, 30 multimodal) constructed via a dual-path pipeline that supports periodic updates, enabling a live and evolving setting. The proposed evaluation suite assesses deep research systems along three complementary dimensions: adaptive synthesis quality evaluation with task-specific rubrics, agentic factuality evaluation via active retrieval and reasoning over both web sources and multimodal attachments, and process-centric evaluation audits how the system searches, reasons, and refines throughout its investigation. Evaluation across 11 systems yields three principal findings: the three evaluation dimensions capture complementary aspects of system capability, with each revealing distinct strengths and weaknesses across systems; process quality serves as a reliable predictor of overall outcome while revealing weaknesses invisible to output-level metrics; and multimodal tasks pose substantially greater challenges, with declines concentrated in synthesis and process rather than factuality. Robustness experiments, a human ranking study, and three-annotator verification confirm the reliability of both the evaluation framework and the benchmark. MiroEval provides a holistic diagnostic tool for the next generation of deep research agents.