MACA: Tracing and Correcting Bias Evolution in Multi-Agent LLM Deliberation
Avighna G Babu ⋅ Maulik Rastogi ⋅ Aashrit Anand ⋅ Carol Chen ⋅ Sarvesh Gharat
Abstract
Committees of large language model agents are increasingly used for decisions involving multiple roles and sources of evidence. Communication within these committees can, however, cause a local allocation preference to emerge, propagate, and amplify, while existing evaluations typically identify such deviations only after deliberation has ended. We introduce \textsc{Maca}, an in-flight auditing and correction framework for multi-agent resource allocation. At each turn, a Judge evaluates the panelist's reasoning and allocation against the task evidence and constraints, combining qualitative review with a task-specific Gini signal. When the audit identifies a potential deviation, an Improver supplies targeted feedback for subsequent deliberation. We also study asymmetric oversight, in which 4B models serve as panelists while a 9B model performs the Judge and Improver roles. Across two allocation studies and three communication topologies, asymmetric oversight achieves the lowest mean deviation among the tested model assignments. In the climate resilience study, it reduces normalized deviation to $0.043$ ($12.9$ million), compared with $0.077$ ($23.2$ million) under same-model oversight. In the Fresno case study, it reduces target deviation to $16.1$ million, compared with $23.0$ million under same-model oversight and $23.9$ million without oversight. The asymmetric configuration nevertheless exceeds the Fresno budget by $7.5$ million on average, showing that closer target alignment does not guarantee complete feasibility. These results suggest that in-flight oversight can improve constraint adherence and that its effectiveness depends on how model size is assigned within the committee.
Chat is not available.
Successful Page Load