Knowing When to Speak: State-Aware Participation in Multi-Agent LLM Systems
Abstract
Multi-agent LLM systems increasingly consist of many long-lived agents, each with its own files, permissions, tools, and memory, sharing a channel. In production such systems are coordinated by stacking context into a lead agent and assembling the team at runtime, which works until the agents' contexts stop fitting in one window; files and permissions are per-agent by design, and often the agent that touched a file knows more about the task than the coordinator does. Existing methods nevertheless make the decision of who should act next in a central agent or a graph, from a compressed representation of the agents. We propose an alternative in which each agent makes that decision itself, from the shared record and its own state, and study whether current LLM agents can. We elicit from each agent, before its turn, a probability that a message from it now would advance the group, and score it against an exact per-agent oracle in a controlled testbed and against disclosure oracles on public benchmarks. We find that agents asked directly do poorly (AUROC 0.55–0.63 for four of five model families), but that most of the failure is in locating where the group is rather than in judging their own relevance: when the group's current state is written into the prompt, the judgment is near-perfect (0.993), and a wrong state is followed just as faithfully (0.976 against the oracle it implies). We propose Orient-then-Speak, a three-line, task-agnostic prompt that adds the missing step, and show that it recovers most of the gap (+0.07 to +0.24 AUROC across five model families; the strongest model gains least). Used as a participation rule, self-selection lets agents coordinate without accumulating context: at 𝑁=10, it reaches the utility of broadcasting with 36% of the tokens and outperforms round-robin, random, and a central matcher that decides from disclosed summaries; from N=20 to 50 it outperforms broadcasting at 7–19% of its cost. On Silo-Bench and MuSiQue, whether the group's state can be recovered from the record predicts where the prompt transfers and where only state supplied by the harness helps.