When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning
Abstract
Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. Prior martingale-null theory explains such failures as debate without truth-directed signal. We develop a more complete account of when practical LLM debate succeeds: it needs both correct minority proposals and a readout that can recover them. We formalize this view through latent headroom, the gap between majority vote and proposal-oracle performance, which measures unused correct expertise in an expert society. For readout, we developLatent Verification Debate (LVD), a theory that models candidate proposals as receiving latent verified evidence before influencing final generation. For supply, we study neural-thicket expert societies, where nearby model perturbations provide heterogeneous specialists, and introduce a coverage-based construction that increases complementary proposal supply. Across matched-budget reasoning and multi-discipline benchmarks, our construction increases recoverable headroom, and controlled readout experiments show that gains arise from recovering surfaced correct minorities rather than merely adding interaction. Together, these results identify diverse proposal supply and verification-aware utilization as the two mechanisms that determine when debate helps.