Before Agents Act: Interpreting Failures under Socially Distributed Access
Chia-Chun Hsu ⋅ Luke Lu
Abstract
Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. A failed task then leaves a basic question: did the agent obtain what it needed before trying to act? Incognita asks what happens when the task and success criterion stay fixed but access does not. We transform eighteen customer-service tasks into three settings: direct access, one known intermediary, and six role-isolated participants whose capabilities must be discovered. Across 864 trials with four models, social access reduced success for every model; the pre-specified intervals excluded zero for two. The latest tested model, gpt-5.6-sol, achieved the highest social-access success at 0.65, a 0.11 decrease from centralized indirect access with an interval that included zero. Exploratory comparisons separated five of six model pairs under social access, while neither centralized setting separated any at this sample size. In post-hoc task-blocked tests, three pairwise interaction $p$-values remained significant after multiplicity adjustment. A reference-relative analysis of interaction traces links the wider model gaps to failures to obtain needed information, locating where unsuccessful trajectories stopped without assigning their cause.
Chat is not available.
Successful Page Load