From Intent to Evidence: A Categorical Approach for Structural Evaluation of Deep Research Agents
Abstract
Deep Research Agents (DRAs) answer complex questions by searching the web, checking evidence, and synthesizing conclusions across heterogeneous sources. We introduce a category-theoretic framework for evaluating such agents. The framework treats deep research as a structured mapping from user intent to evidence-grounded conclusions, making retrieval traces, cross-source alignment, and final synthesis explicit. Guided by this view, we build a mechanism-aware benchmark of 296 bilingual questions covering four structural skills: following multi-hop evidence chains, verifying claims across sources, re-ordering fragmented information, and rejecting unsupported assumptions. We evaluate 16 systems with human verification and find that these tasks remain difficult: the best system reaches 19.9\% average accuracy. The results reveal complementary strengths across systems, but also persistent weaknesses in long-horizon retrieval and intersection-heavy verification. We further instantiate two theory-guided interventions, tracked search and category tools, in API-based agents. These variants improve over their corresponding baselines, suggesting that the framework is useful not only for diagnosis but also for modest, targeted system design.