DockRepair: Evaluating Safe Abstention in Agentic Docker Compose Remediation
Abstract
Tool-using coding agents can recover routine Docker Compose faults, yet operational safety also requires stopping when the evidence does not justify another change. We study that boundary with DockRepair, a deterministic recovery system for declared TCP dependencies in single-host Compose projects. DockRepair represents declared and observed states as symbolic facts, uses read-only DNS, TCP, and listener probes when runtime state is ambiguous, and searches a bounded catalog of project-scoped repairs. It applies one action at a time and externally verifies the original dependency after every change. In a frozen benchmark of ten application-failure pairs, each repeated three times for four systems, all systems repaired all 15 supported trials. On 15 unsupported trials, DockRepair achieved 15 correct abstentions; the three evaluated coding-agent baselines achieved 6, 4, and 4. On eight held-out pairs, removing active probes reduced supported-repair success from 8/8 to 4/8, while removing graph-based diagnosis selection tied the full system. Within this bounded scope, active diagnosis, constrained execution, and external verification provide a useful safety boundary for automated recovery.