False Convergence: Representation-Level Signals of Incorrect Consensus in Multi-Agent LLM Systems
Abstract
Multi-agent LLM systems seek to improve reliability by distributing generation, critique, and evaluation across multiple model instances. A common design employs inter-agent agreement as a termination criterion: once the agents converge, the system accepts the consensus output. Such systems may terminate in true convergence, where the accepted answer is correct, or false convergence, where the accepted answer is incorrect. Yet based on the convergence verdict alone, the two outcomes are indistinguishable. This motivates our central question: can internal representations distinguish true convergence from false convergence? We train linear probes on internal activations from a proposer-critic-judge system built with Gemma 4-31B-IT and evaluated on the text-only subset of Humanity's Last Exam (HLE). With probe locations selected using only training data, the probes distinguish true from false convergence with a mean held-out AUROC of 0.702 across 1,432 converged trajectories. A terminal proposer probe outperforms a probe of the initial proposal and adds predictive information beyond the initial probe and recorded interaction metadata. In counterfactual replays, correcting an incorrect terminal proposal increases a fixed critic probe's true-convergence score more than a style-only rewrite or an alternative wrong answer. Independently trained critic and judge probes also distinguish correct from incorrect consensus in GLM, Llama, and Mistral checkpoints. In those three systems, corrected proposals prompt renewed deliberation more often than both controls. Together, our results indicate that convergence behavior and correctness-related internal signals can dissociate: the system's external verdict may collapse distinctions that remain partially recoverable from internal activations.