ClawBenchPro: Benchmarking How Well Agent Harnesses Work
Abstract
The evolution of Large Language Models (LLMs) from conversational assistants to autonomous agents is increasingly driven by agent harnesses (e.g., OpenClaw), which provide the operational substrate for executing complex, multi-step workflows in real-world environments. However, existing claw-supported benchmarks often suffer from narrow domain coverage and reliance on single-harness evaluations, leading to evaluation distortion and limited diagnostic diversity. To bridge this gap, we present ClawBenchPro, a comprehensive benchmark comprising 1.0k samples across 43 distinct domains. By integrating skill augmentation, environmental heterogeneity, and multi-session dependencies, we replicate complexity workflow at scale. Extensive experiments across 14 frontier LLMs and four representative harnesses, show model's average scores clustering at 60-74\%. We identify a critical domain disparity: models excel in cyber-native tasks but struggle with real-world operational workflows. Moreover, a pervasive reliability bottleneck emerges, where incremental progress rarely translates to zero-defect execution. Comparative results show that while OpenClaw incurs higher computational costs, NanoClaw offers superior cost-efficiency (despite a 3.2–6.5\% accuracy trade-off), and Hermes-Agent emerges as the most stable framework for claw-style interactions. ClawBenchPro establishes a rigorous, reproducible standard for diagnosing agent limitations and advancing next-generation harness architectures.