ListQA: A Benchmark for Evaluating List-Formatted Factual Knowledge Retrieval in Large Language Models
Abstract
Many real-world questions demand not a single fact but an organized collection of them, yet existing factual knowledge benchmarks almost exclusively target single-answer retrieval. We introduce ListQA, a benchmark of 9,045 human-curated, cross-validated questions (estimated error rate below 2%) that require LLMs to recall multiple facts and compose them into structured lists of 3--10 elements. Questions range from flat lists Which NCAA teams went undefeated between 2000 and 2024? to hierarchical lists with sub-attributes ... and what was their record and result?, spanning eight categories across 60+ countries. An evergreen design (explicit temporal constraints or historically immutable facts) keeps ground truth valid without periodic updates. For evaluation, we frame element matching as a linear sum assignment problem: optimal bipartite matching paired with LLM-based semantic grading jointly captures factual accuracy, hallucination, and format compliance. We benchmark 34 models (8 families, 42 configurations across standard and thinking-enabled inference) and find: the best standard-inference model scores below 37% LLM-Judge-F1, multi-faceted questions are harder across the board, and providing the source Wikipedia page lifts scores dramatically, confirming that models can organize facts but struggle to retrieve them from parameters alone. Extended analyses cover base vs. instruction-tuned checkpoints, oracle RAG, and token efficiency. We release the dataset and evaluation code.