CogArena: Benchmarking Multimodal Agents on Interactive Behavioral Experiments
Abstract
Claims about the cognitive capabilities of large language models are typically tested by translating behavioral paradigms into natural language prompts, bypassing the visual interfaces and time-pressured interactions through which human participants actually engage with experiments. In parallel, recent demonstrations that AI agents can produce human-like response data on live behavioral tasks raise concerns that online datasets used across the social and behavioral sciences are increasingly susceptible to contamination. We introduce CogArena, a benchmark of interactive experimental paradigms drawn from the social and behavioral sciences. Agents interact with the experiments through a standard browser, processing visual stimuli and responding under configurable deadlines. We designed a three-level scoring pipeline that evaluates task completion, performance accuracy normalized against literature-derived human baselines, and alignment with canonical task-specific signatures. Across frontier multimodal agents, we find that agents complete the tasks, but their behavior diverges from humans on the signatures these paradigms reliably elicit in people. CogArena provides the tasks, API, scoring pipeline, and a public leaderboard for ongoing community submissions.