Phantom Consensus: Measuring Social Reasoning in Multi-Agent LLM Decision-Making
Abstract
Multi-agent LLM systems increasingly aggregate judgments across model instances, but reasoning about other agents can itself introduce failure. We study a setting in which an LLM changes its public decision after incorrectly inferring what other agents privately judged, even when the group's measured private judgments do not support that inferred consensus. We call this failure phantom consensus and study it as a pluralistic-ignorance-style pattern in controlled LLM collectives. To measure this phenomenon, we introduce PHANTOM-Bench, a benchmark that separates isolated private judgments, public decisions after ambiguous social signals, elicited estimates of the group's private judgment, transparency-based recovery, and controls for neutral public text and explicit vote-following. Across nine closed- and open-weight models, ambiguous public signals consistently degrade decision quality and can turn correct private group judgments into incorrect public collective outcomes. Public accuracy drops by as much as 47.8 percentage points, and collective rationality failures reach 47.3\% in the most susceptible setting. These failures are associated with erroneous second-order judgments, are not explained by public text alone, and are distinct from direct vote-following. Revealing the agents' recorded private decisions repairs many failures without revealing the gold label. Our results show that evaluating multi-agent LLM systems requires measuring not only final agreement, but whether that agreement reflects independently measured private judgments or an imagined consensus.