DiscernBench: Evaluating Cooperative Discernment of LLM Agents in Institutional Settings
Tanishka Magar ⋅ Disha Sheshanarayana
Abstract
The deployment of large language models in operational settings demands evaluation of judgment under institutional constraint. We introduce DiscernBench, a benchmark for evaluating cooperative discernment: the ability of an agent to determine whether to comply with, refuse, or escalate a request in an institutional setting. DiscernBench comprises scenarios derived from real-world incident reports across domains such as aviation, law enforcement, and emergency management. We evaluate eight LLMs across three decision-making configurations and analyze both final decisions and the intermediate judgments produced during multi-agent deliberation. Under disagreement, Final Responders exhibit model- and role-dependent aggregation patterns in which a lone specialist recommendation can disproportionately influence the collective decision, a phenomenon we term Minority-Signal Override. Role-specialized multi-agent deliberation raises average over-escalation from $3.79\%$ under direct single-agent judgment to $28.27\%$, with the direction and magnitude of minority-signal influence varying substantially across models and specialist roles. These results reveal distinct collective-decision behaviors that are not captured by final-answer accuracy alone. Trustworthy multi-agent systems therefore require evaluating not only individual specialist reasoning, but also how conflicting judgments are transformed into collective decisions. Code and data are available at \url{https://anonymous.4open.science/r/DiscernBench-19C9}.
Chat is not available.
Successful Page Load