GISA: A Benchmark for General Information-Seeking Assistant
Abstract
The advancement of large language models~(LLMs) has significantly accelerated the development of search agents capable of autonomously gathering information through multi-turn web interactions. However, existing benchmarks often construct queries backward from answers, producing tasks misaligned with real-world needs. Moreover, these benchmarks tend to focus on either locating specific information or aggregating information from multiple sources, while relying on static answer sets prone to data contamination. To bridge these gaps, we introduce GISA, a benchmark for General Information-Seeking Assistants comprising 373 human-crafted queries reflecting authentic information-seeking scenarios. GISA features four structured answer formats (item, set, list, and table), enabling deterministic evaluation. It integrates both deep reasoning and broad aggregation within unified tasks, and includes a live subset with periodically updated answers to resist memorization. Notably, GISA provides complete human search trajectories for every query, serving as valuable references for analyzing and improving search agent behavior. Experiments on mainstream LLMs and commercial search products reveal that even the best model achieves only 19.30% exact match, with performance notably degrading on tasks requiring complex planning and comprehensive information gathering.