AgentFinVQA: An Auditable, On-Premise Multi-Agent Pipeline for Financial Chart QA
Abstract
Financial chart question answering (QA) in regulated settings requires more than accuracy: practitioners must assess whether an answer is trustworthy, while many institutions cannot send sensitive client data to external model providers. Existing chart-QA agents provide limited auditability and often rely on proprietary APIs. We present AgentFinVQA, a multi-agent pipeline that decomposes each query into planning, OCR, legend grounding, visual inspection, and verification while recording intermediate outputs and verifier decisions in a traceable Model Evaluation Packet (MEP). On FinMME, AgentFinVQA improves over primary-backbone-matched zero-shot baselines by 7.68 percentage points with Gemini-3 Flash (71.24% vs. 63.56%, McNemar p ≈ 1.1 × 10⁻¹⁶) and 4.84 percentage points with open-weights Qwen3.6-27B-FP8 served locally. Error analysis shows that question misunderstanding, legend confusion, and extraction errors account for nearly two-thirds of failures and are among the failure types least reliably identified by the verifier. These results show that auditable financial chart QA can retain substantial accuracy gains while enabling per-answer inspection and fully on-premise deployment. We release our code to support reproducible evaluation: https://anonymous.4open.science/r/agentFin-anon/.