SLATE: A Large-Scale Benchmark for Long-Horizon API Selection in Enterprise Tool Use
Abstract
Enterprises increasingly want LLM agents to automate standard operating procedures over their software stacks—mapping each step of a procedure to the right call among thousands of internal APIs, without hand-authored step-to-API mappings. The core difficulty is API selection at scale: choosing and composing the correct tools over a long horizon, judged only by whether the overall task succeeds. Current tool-use benchmarks do not test this regime—they expose only a handful of tools, evaluate single-step invocation, or grade multi-step behavior with subjective LLM-as-a-judge scores that penalize valid alternative trajectories. We introduce SLATE (Synthetic Large-scale API Toolkit for E-commerce), a large-scale benchmark that tests large-toolset, long-horizon API selection in a realistic enterprise domain. SLATE couples (i) queries decomposed into hierarchical, tool-grounded procedures over a ~1,000-tool e-commerce library, (ii) a deterministic, context-aware simulator that returns grounded outputs for valid calls and failure signals otherwise, and (iii) an end-to-end task-success metric—exact match against the grounded final outcome—that credits diverse-but-valid trajectories while remaining fully automatic and reproducible. Because only task-level (binary) success feedback is available, a central question is how efficiently an agent can roll out to a correct solution. A study on an open-source model shows that even self-reflective agents solve well under half of the tasks, establishing SLATE as a challenging testbed for reliable enterprise tool use. We release SLATE, its simulator, and the evaluation harness.