AIM: Adaptive Interaction in Multi-Agent Debate for Multimodal LLM Inference
Abstract
Multi-Agent Debate (MAD) improves reasoning on multimodal tasks, but its accuracy and token efficiency depend on how agents interact during debate. A largely overlooked design choice is which interaction modalities agents use to exchange evidence. Existing methods rely on fixed text- or graph-based interaction, which fails to convey fine-grained multimodal evidence, degrading accuracy while inflating token cost due to verbose descriptions. To address this, we propose Adaptive Interaction in Multi-Agent Debate (AIM), an adaptive framework that dynamically selects the best combination of interaction modalities for each instance. Beyond text and graph, AIM introduces task-dependent modality interaction, where agents exchange targeted regions of interest from the task-specific multimodal input, such as image regions, audio segments, and spatio-temporal clips. To determine which combination is the best for each instance, AIM first generates a structured baseline response, extracts interpretable routing features, and finally uses a lightweight router to select a modality combination that avoids both under-fusion (i.e., agents miss decisive evidence) and over-fusion (i.e., redundant modalities inflate token cost) risks. Across eight multimodal question answering (QA) benchmarks of Visual-QA, Audio-QA, and Video-QA, AIM achieves the highest accuracy on all datasets, improving upon the best single-modality baseline by up to 8.9% while reducing token cost by up to 53.4% relative to the best fixed-modality baseline.