RAVEN-Bench: A Paired EO-IR Video QA Benchmark for Aerial Multimodal Understanding
Abstract
Multimodal large language models (MLLMs) have advanced rapidly on image and video understanding, but their ability to reason over paired aerial electro-optical and infrared (EO-IR) videos remains underexplored. Such videos contain small targets, sparse temporal evidence, changing viewpoints, and modality-dependent cues that are poorly captured by RGB-centric benchmarks. We introduce RAVEN-Bench (Reasoning over Aerial Videos with EO-IR aligNment), a paired EO-IR video QA benchmark for low-altitude aerial multimodal understanding. RAVEN-Bench contains 48 temporally aligned EO-IR video pairs and 576 human-verified four-way multiple-choice questions. It supports EO-only, IR-only, and paired EO+IR evaluation, and uses hierarchical capability levels and structured question groups to diagnose cross-modal gain, modality reliance, consistency, and reasoning coherence. We evaluate representative frontier and open-weight MLLMs, showing that aggregate performance can overestimate reliable EO-IR reasoning. RAVEN-Bench provides a focused testbed for assessing whether progress on RGB video benchmarks transfers to aerial cross-spectral understanding.