Humans Correct Their Judgements Through Debate, but Weaker AI Models Not
Abstract
How can we effectively oversee AI systems that surpass human intelligence? One proposed answer is AI Safety via Debate, an approach in which two AI systems argue opposing positions and then a judge, who can be either a human or a weaker model, decides which one is correct. By facilitating supervisors in providing high-quality training signals, debate aims to encourage truthful and safe behavior from AI systems whose capabilities exceed our own. However, human judgment is not neutral: people often rely on biases which advanced AI may learn to exploit. Drawing on insights from collective intelligence and social cognition, this paper investigates whether debate can guide human and LLM supervisors toward the truth despite their prior beliefs. We introduce and evaluate several protocol configurations, including (1) the use of open-source, proprietary models and (2) human participants as judges (N=507), (3) different methods for eliciting prior beliefs, and (4) variations in claims within the same topic to test whether belief updates generalize across related subtopics. We also explore setups inspired by the wisdom of the crowds, such as the number and diversity of judges, as well as interaction mechanisms including deliberation and voting, designed to better align judgments with the truth. Overall, we find that debate increases accuracy relative to prior beliefs for human judges (from 41.3\% to 57.6\%), whereas the consultancy baseline, in which a single expert argues for one answer, does not. For weaker language models, however, the results are mixed. Debate does not provide similar benefits in most settings and can even degrade decision accuracy. Regarding multi-judge protocols, human deliberation yields the highest accuracy for human supervisors (62.2\%), while for LLMs it mainly benefits those starting from incorrect beliefs, without producing significant improvements in aggregate performance. We provide empirical evidence supporting Debate as a viable path toward scalable human oversight, while casting doubt on the assumption that such benefits translate to weaker AI supervisors, raising concerns about the reliability of fully automated training and control pipelines in high-stakes AI safety contexts.