Beyond Task Success: Probing Cognitive Primitives in Web Agents
Abstract
Cognitive primitives such as planning, exploration, and backtracking are widely regarded as core to competent web agents. Yet existing studies are largely retrospective: agents are evaluated on domain-based tasks, and trajectories are inspected post hoc for evidence of these behaviors. Because such tasks are not designed to require specific primitives, observed behaviors do not reliably reflect true capability demands. This raises a key question: Can we design realistic tasks where success provably depends on a chosen cognitive primitive? WEBSTRESS addresses this by casting capability evaluation as a controlled comparison. For each task, we construct a paired variant in which a target capability becomes necessary (e.g., transforming a direct path into an obstructed one requiring exploration). The performance gap between the pair measures the agent’s reliance on that capability. Across seven environments and seven capabilities, this yields 519 paired tasks. Evaluating six frontier agents reveals consistent weaknesses: all suffer 18-28\% drops on capability-demanding variants, with 75\% of failures being belief failures, where agents incorrectly declare success. In contrast, humans exhibit only a 5.7\% performance drop under the same paired evaluation. These results highlight that, despite strong overall performance, substantial gaps remain between human and agent cognitive behavior, and that our approach provides a principled way to measure them.