EnvFaultBench: Benchmarking LLM Agents on Software Environment-Fault Troubleshooting
Abstract
Large language model (LLM) coding agents are increasingly used as general-purpose software engineering assistants, yet existing benchmarks focus on source-code patches or environment setup from scratch. A common class of failures remains unaddressed: the code is correct but the environment is faulty due to dependency conflicts, configuration drift, stale caches, or resource contention. We present EnvFaultBench, to the best of our knowledge the first benchmark targeting environment-fault troubleshooting. It contains 348 Dockerized instances derived from real GitHub issues across 75 open-source projects in three ecosystems (Python, TypeScript/JavaScript, JVM), covering 23 fault types in three root-cause layers. Hidden functional oracles accept any command sequence that restores target behavior. Evaluating 10 LLMs (4 proprietary, 6 open-weight) under a unified agent framework, we find that the best model (GPT-5.5) resolves 65.8% while a zero-shot baseline resolves only 10.9%. The gap between strong and weak models is strongly associated with diagnostic efficiency under our bounded protocol: stronger models commit to a repair within 2–3 steps (Spearman ρ = −0.85 between first-fix latency and solve rate), while weaker models exhaust the budget on undirected exploration. Data and code are publicly available.