DIS-Bench: Evaluating LLMs on System Testing via Directed Input Synthesis
Abstract
LLM-based coding agents have shown strong capabilities on real-world software testing tasks, yet existing evaluations concentrate on unit test generation, leaving system testing unexplored. System testing is both practically important for software quality assurance and inherently difficult, as it demands repository-level holistic understanding and long-horizon reasoning over program behavior. In this paper, we propose to measure LLM agents' system testing capabilities through its core problem, Directed Input Synthesis (DIS): generating system-level inputs that exercise specified code branches. We introduce DIS-Bench, a benchmark built via an automated pipeline based on random testing that selects target branches that are both reachable and challenging. Because DIS-Bench is constructed from execution behavior rather than human-authored issues or patches, it is scalable and less susceptible to contamination from existing training corpora. DIS-Bench contains 8267 tasks drawn from 37 real-world repositories. We evaluate 7 mainstream open-sourced LLMs on DIS-Bench-Lite under both BM25 retrieval and state-of-the-art coding agents, including OpenHands, Mini-SWE-Agent, and Claude-Code. The results show that DIS is highly challenging for current models: the best-performing configuration resolves only 28.1\% of targets at pass@1. We further perform a manual inspection to characterize the dominant failure modes, revealing reasoning bottlenecks of LLMs on repository-level system testing tasks.