Be Cautious when Blaming Persuaders: LLM Committees Reach Defensible Targets on Their Own
Paulina Tomaszewska ⋅ David Williams-King ⋅ Christian Schroeder de Witt ⋅ Sahar Abdelnabi
Abstract
A committee of LLM agents shifts its position over the course of a debate. Was it adversarially persuaded, or did it simply come to this conclusion through well-intentioned debate? When the option an adversary pushes is itself defensible, it is hard to spot a difference between these cases. We present HoldSway, a committee testbed that runs every scenario twice: once with a persuader and once with an identical persuader-free committee, so that adversarial influence is measured as a residual (persuasion--control gap) rather than interpreted from the outcome. Persuader-free committees reach the persuader's target on their own in $0.467$ of runs on Qwen3-32B and $0.673$ on Llama-3.1-70B, against $0.547$ and $0.827$ with a best-resourced persuader. This means a marginal contribution of persuaders ($+0.08$ to $+0.15$). What affects conversion is the target's quality more than the persuader's effort. To test this, we construct a cleanly-wrong set whose target is inferior on every stated criterion, where we found that persuader-free committees almost never convert. Probes on internal representations confirm the observation that it is hard to distinguish natural drift from persuasion. We release the parametrized benchmark with a communication routing knob and the white-box probe suite.
Chat is not available.
Successful Page Load