Skip to yearly menu bar Skip to main content


Before Measuring Human-Agent Teams: Validating Human References in Agent Benchmarks

Anqi Peter Li ⋅ Ethan Yip ⋅ Kundana C Kommini

Abstract

Chat is not available.