CARVE: Caption-Anchored, Rubric-Verified Extraction for Training-Scale Figure-Required Scientific VQA
Abstract
Fine-tuning vision-language models (VLMs) to comprehend complex scientific literature is bottlenecked by a trade-off between dataset scale and data quality. Human-curated benchmarks preserve quality but are far too small (1--3k items) to serve as fine-tuning data. Conversely, large-scale automated corpora achieve the volume necessary for fine-tuning but are plagued by repetitive language and a heavy reliance on multiple-choice questions (MCQs). To bridge this gap, we introduce CARVE (Caption-Anchored, Rubric-Verified Extraction), a massive dataset of 392,957 items designed to fill this gap in scientific VQA: it provides the large scale required for robust model training while strictly retaining the high-quality, prior-free, and completely free-form characteristics of expert human benchmarks. CARVE is constructed using a scalable text-first, vision-verified pipeline featuring multi-stage quality filtering, ensuring that every generated question is grounded and unanswerable without the underlying figure. We evaluate CARVE by fine-tuning open-source VLMs across the Qwen and InternVL families. Our fine-tuned models exhibit notable cross-family transfer gains on complex, open-ended scientific reasoning benchmarks, including a +3.9 lift on CharXiv-reasoning at 7B scale. Ultimately, our work demonstrates that shifting away from multiple-choice constraints toward training-scale, free-form visual data is a powerful paradigm for building VLMs capable of true scientific comprehension.