AgentScript-Eval: Benchmarking LLMs and Agents for Code Generation in an Enterprise DSL
Dipin Khati ⋅ Shubham Mehrotra ⋅ Bin Bi ⋅ Zhou Yu ⋅ Denys Poshyvanyk ⋅ James Zhu ⋅ Sitaram Asur ⋅ Phil Mui
Abstract
Domain-specific languages (DSLs) are widely used in enterprise software, giving a compact interface to complex tasks, from data access to infrastructure automation. Platforms increasingly extend them to agents, where a developer declares the topics, permitted actions, and guardrails of an agent rather than writing general-purpose code. Large language Models(LLMs) write general-purpose code well, but that ability has been measured almost entirely on languages abundant in public pretraining data, and enterprise DSLs are neither abundant nor public. We introduce \textbf{AgentScript-Eval}, an execution-graded benchmark of $100$ real tasks for AgentScript, a declarative DSL used within a major enterprise software platform to build conversational agents. Tasks span authoring an agent from a specification, debugging a broken definition, and optimizing a working one. Evaluating five models zero-shot, we find that the regime is hard: the best model reaches only $25.0\%$ pass@3, and every model scores $0\%$ on authoring from scratch. We further find that model scale and price are weak predictors of success---a $31$B open-weight model attains the highest compilation rate of any model tested ($48.0\%$) and matches the strongest frontier model on partial-credit accuracy at the lowest token cost. Adding an agentic compile-and-edit loop is the single largest lever, roughly doubling partial-credit accuracy. Beyond that gain, the wording of domain guidance matters more than its packaging. The same skill pack scores $28\%$ as plain documentation but only $14\%$ as a live plugin. Rewriting one skill inside that plugin then raises it to $37\%$, our best result, with catalog and tooling untouched. A per-category breakdown shows the difficulty is concentrated in authoring from scratch, where no zero-shot model succeeds, and even the best agentic configuration stays under $17\%$. Larger models did not reliably outperform smaller ones; what mattered was the harness surrounding the model and the guidance supplied to it.
Chat is not available.
Successful Page Load