Beyond Accuracy: Formalization for Understanding Information Exchange in Multi-Agent Debate
Abstract
Multi-agent debate allows language models to exchange reasoning and revise their answers, but final accuracy does not reveal which properties of that exchange distinguish successful debates. We introduce a claim-level observability framework that translates explanations into attributed, temporally ordered propositions and explicit reasoning relations. Using the controlled Knight-Knave-Spy domain, we computationally evaluate claim truth and local logical support, quantify information content, and track propositions across agents. Across 120 debates, we examine information quantity, quality, and uptake. Within experimental conditions, net initial information distinguishes correct from incorrect debates with an AUC of 0.862, compared with 0.901 for the presence of a correct initial solution. False-claim resistance also discriminates outcomes (0.860), while selective uptake remains inconclusive (0.787, 95% CI [0.474, 0.984]) and the number of distinct initial propositions is less informative (0.532). These findings associate collective success with the quality of initial information and resistance to the spread of false propositions.