FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation
Abstract
The financial domain poses substantial challenges for Vision-Language Models (VLMs) due to specialized chart formats and knowledge-intensive reasoning requirements. However, existing financial benchmarks are largely single-turn and rely on limited question formats, which hinders comprehensive evaluation in realistic application scenarios. To address this gap, we propose FinMTM, a multi-turn multimodal benchmark that expands diversity along both data and task dimensions. On the data side, we curate and annotate 11,133 bilingual (Chinese and English) financial QA pairs grounded in commonly used financial visuals, including candlestick charts, statistical plots, and report figures. On the task side, FinMTM covers single and multiple choice questions, multi-turn open-ended dialogues, and agent-based tasks. We further design task-specific evaluation protocols, including a set-overlap scoring rule for multiple choice questions, a weighted combination of turn-level and session-level scores for multi-turn dialogues, and a composite metric that integrates planning quality with final outcomes for agent tasks. Extensive experimental evaluation of 22 VLMs reveal their limitations in fine-grained visual perception, long-context reasoning, and complex agent workflows.