MCPHallu: Benchmarking Reasoning, Execution, and Memory Hallucinations in MCP Agents
Yizhen Jiang ⋅ 浩伟 郭 ⋅ Tianyi Bai ⋅ Zengjie Hu ⋅ Lu Ma ⋅ Zhengyang Zhao ⋅ Haoze Sun ⋅ Peng Pei ⋅ Wentao Zhang
Abstract
Large language model agents are increasingly deployed through the Model Context Protocol (MCP), which allows them to discover and invoke tools across multiple servers at runtime. Existing MCP benchmarks primarily measure task completion capability but provide limited insight into why agents fail, reporting aggregate success rates without attributing failures to specific reasoning, execution, or memory breakdowns. We introduce MCPHallu, a benchmark for diagnosing hallucination-prone agent failures in MCP environments along three axes: reasoning, execution, and memory. MCPHallu contains 358 tasks across five capability domains and runs on 36 production MCP servers in isolated containers. Each task is designed with a structural trigger that makes one of four subtypes the dominant failure risk: Branch Collapse, Unreachable Goal, Tool Misuse, or Context Forgetting. This design enables subtype-level analysis rather than aggregate task success alone. We evaluate fourteen frontier LLM agents from eight vendors on 5,012 trajectories. Current MCP agents remain vulnerable to hallucination-inducing tool environments, achieving an average score of 0.642. Failures concentrate on Unreachable Goal tasks, where the average score drops to 0.458 and 22.1\% of runs pass, indicating that agents often continue issuing tool calls or claim successful completion rather than recognizing infeasible goals. Performance across the four subtypes is largely independent: per-model ranks span on average 6.9 of 13 positions, and Branch Collapse is uncorrelated with Unreachable Goal across models ($r = -0.12$). MCPHallu is released at https://anonymous.4open.science/r/mcp_hallu.
Chat is not available.
Successful Page Load