From Local Agent Success to Ecosystem Convergence: Evaluating Human-Agent Teams in Distributed Software Evolution
Abstract
Repository-local success can miss failures when software features cross independently deployed services. We formulate distributed software evolution as an evaluation problem for human-agent teams and define ecosystem convergence: feature acceptance, regression safety, contract compatibility, mixed-version behavior, and rollout must all pass. We introduce the Agentic Development Plane as a falsifiable coordination hypothesis and DevPlaneBench as its evaluation methodology. A bounded protocol prototype establishes executability, not benefit. An exploratory pilot compares two fresh teams using information-equivalent prose with two using typed evolution records, under common local verification gates, on one price-history task. All four pass the exercised behavioral and rollout checks. Raw conjunctive success is 0/2 for prose and 2/2 for typed records, but the difference arises solely from a schema-publication rule not explicitly disclosed in the common task material. We retain these scores without interpreting them as protocol superiority. The contribution is a validity-centered evaluation design and a concrete demonstration of why evaluator policy, transition coverage, and coordination treatment must be audited separately. Evaluating benefits for human-agent teams and the comparative effectiveness of the full protocol remains future work.