Harbor Adapters and Harbor-Mix: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Abstract
Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of frontier models across 54 benchmarks, enabling a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Mix, a curated set of 100 difficult, diverse, and high-quality tasks refined from the adapted benchmarks. Harbor-Mix preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; the strongest model resolves only 15.6\% of its tasks. We release the adapters, evaluation results, in-depth analysis, and Harbor-Mix as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.