Same Behavior, Different Feedback: Interpreting Multi-Agent Coordination Beyond Performance Metrics
Kyungyoon Jung ⋅ Donggun Lee ⋅ Hyun Seung Moon ⋅ Eunji Shin ⋅ Seon G Kim ⋅ SeongWon Hong ⋅ Juho Kim ⋅ Tak Yeon Lee
Abstract
Human feedback is used to evaluate and steer multi-agent systems, yet little is known about how observers construct feedback from unfolding coordination. We study this process in a simulated smart factory, where 216 participants monitored 108 video clips of different agent behaviors and produced 1,746 feedback items. A mixed-methods analysis yields a five-dimensional taxonomy capturing the unit of observation, observed behavior, monitoring strategy, interpretation, and feedback type. Human collaboration ratings tracked task performance, but weakly separated low- and high-performance (AUC $= .60$), showing that system performance metrics and human judgment capture related but distinct aspects of coordination. Observers viewing the same behavior also selected different moments as noteworthy ($11.2\%$ temporal overlap) and, even when attending to the same moment, shared only $32.8\%$ of their taxonomy codes. Nevertheless, pooling observers expanded the temporal coverage from $26.7\%$ with one observer to $89.5\%$ with eight. Together, these findings suggest that evaluating multi-agent coordination requires complementary evidence: performance metrics capture task outcomes, while aggregated human feedback reveals diverse and otherwise overlooked aspects of agent behavior.
Chat is not available.
Successful Page Load