Aggregating Open-Ended Answers Across Populations of LLM Agents
Sneheel Sarangi ⋅ Kevin Chou ⋅ Haocheng Guo ⋅ Ahan Karnik ⋅ Vihaan Kabra
Abstract
Methods for aggregating multiple large language models are well developed for closed-ended tasks, where agents choose from a shared set of answers. Open-ended generation is harder because different agents may produce only partially overlapping content, and there is no fixed set of choices to vote over. We study a setting where open-ended answers can be decomposed into claims and ask how ideas from closed-ended aggregation carry over to these claims. We first adapt counting and reliability weighting, where each claim is scored using the number and reliability of the agents that assert it. We then study an additional source of evidence that only appears in generation: how agents produced or omitted a claim. On list-answering tasks, correct claims tend to appear earlier in an agent's output, while omission is stronger evidence against a claim when a silent agent has produced many other answers. We use these signals in a behavioural claim scorer and choose how many claims to return separately for each question. On QAMPARI, behavioural aggregation reaches $31.70$ set F1, compared with $29.21$ for counting and $27.40$ for the best single agent. On CWQ, it reaches $23.65$, compared with $21.81$ for counting. The gains hold across agent subsets and under question-type distribution shift on QAMPARI. ASQA shows a different regime: its responses are short and similar in length, so the behavioural signals are weak, and reliability weighting performs best at $27.45$ $\mathrm{DR}^*$. These results show that voting and reliability weighting can be extended to decomposable open-ended answers, but that open-ended generation also provides task-dependent evidence beyond agent agreement alone.
Chat is not available.
Successful Page Load