What Makes LLMs Effective Information Seekers? Diagnosing Question Asking and Belief Tracking in LLMs
Daniel Machado Pedrozo ⋅ Bryan Lincoln Marques de Oliveira ⋅ Murilo da Luz ⋅ Telma Lima ⋅ Luckeciano Carvalho Melo
Abstract
Efficient question-asking is a crucial capability for multi-turn agentic systems, yet it remains difficult to quantify at the level of individual queries. Outcome-level metrics such as win rate compress a dialogue into a single success signal, obscuring whether a model uses each turn efficiently or merely reaches the target eventually. We introduce the Information Gain Game (InfoGainme), a multi-turn analysis framework that embeds guessing games in closed hypothesis spaces and uses information gain (IG) to measure the uncertainty reduction produced by every question. We evaluate $14$ models across three domains and three observability modes, with and without Chain-of-Thought (CoT), and parse reasoning traces to inspect the candidate questions models consider before asking. We find that turn-level IG remains informative in lost games and separates models even when win rate saturates. Three observability modes decompose information seeking into question formulation, belief maintenance, and belief construction; CoT gains concentrate in the implicit-belief regimes and trace analysis shows that CoT acts as a scaffold for explicit candidate tracking. The remaining bottleneck is not only choosing among candidate questions: models also fail to generate near-balanced partitioning questions, and often do not select the best question from their own deliberation pool. We release trajectories, per-turn IG values, and parsed reasoning traces to support training agents that ask better questions at each step.
Chat is not available.
Successful Page Load