FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration
Abstract
Repository-level benchmarks such as SWE-bench have advanced the evaluation of LLM agents for software engineering. Successful deployment, however, depends not only on repository code but also on the deployment environment: dependencies, container images, orchestration, service connectivity, and health monitoring. Existing evaluations provide limited coverage of whether agents can configure and diagnose these components. We introduce FDE-Bench, a benchmark of 136 tasks evaluating LLM agents' ability to configure deployment environments across single container images, multi-container Compose stacks, and Kubernetes, in greenfield and diagnose-and-repair modes. Each task defines a natural-language deployment ticket and a machine-readable specification, and the agent must configure the environment to satisfy that specification. Grading is by replay. The agent's declarative artifacts are re-deployed in a pristine sandbox and scored by four gated binary check layers: build, readiness, behavior, and spec conformance. Every check is programmatic and no LLM judge takes part. Zero-intelligence baselines certify that the tasks resist gaming. A do-nothing agent, a spec-transcriber, and a generic stub resolve none of the 136 tasks and reach mean Deployment Score at most 0.39, while every task's reference solution resolves in the release gate. Seven agents from four providers, run under an identical minimal harness, resolve between 55 and 77 percent of tasks. Repair tasks resolve 28 points higher than greenfield authoring, with the gap positive for every agent; nine tasks resist all seven agents, and a practicing engineer directing Claude-Sonnet-5 resolves 92 percent of a 25-task subset that the model alone resolves at 68.