EDMA: Entropy-Driven Multimodal Answering
Emanuele Mezzi ⋅ Gertjan Burghouts ⋅ Fabio Massacci ⋅ Mengyuan Zhang
Abstract
Multimodal question answering (MMQA) is based on integrating heterogeneous data sources, selectively leveraging relevant modalities, and ignoring distractors. Existing approaches based on Multimodal Large Language Models (MLLMs) improve performance by splitting the task into intermediate steps, but do not quantify how each modality contributes to reaching the final answer. We argue that quantifying the effect that each modality has on the answering process improves process traceability and final performance. Thus, we propose Entropy-Driven Multimodal Answering (EDMA), which uses logical entropy to formalise multimodal QA as an iterative process of entropy reduction and to quantify the contribution of each modality in distinguishing correct from incorrect candidate answers. This enables the construction of quantifiable answering trajectories, where each step is associated with a measurable entropy reduction. EDMA introduces (i) modality-conditioned partitioning to estimate unimodal entropy reduction, (ii) entropy-driven multimodal fusion to capture complementary information across modalities, and (iii) entropy-based selection of the answering trajectory that minimises the most logical entropy. Experiments on multimodal QA benchmarks show that EDMA consistently outperforms prompting-based baselines and state-of-the-art methods. Averaged across datasets, EDMA achieves gains of $12.8$ and $19.8$ F1 points over the strongest prompting baseline and state-of-the-art system, respectively. Moreover, we show that higher entropy correlates with more false positives and, through controlled interventions on the choice of the answering trajectory, provide evidence that lower-entropy partitions causally reduce false positives, validating entropy level as a reliable proxy for performance. We share our code at https://anonymous.4open.science/r/EDMA/README.md.
Chat is not available.
Successful Page Load