ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Abstract
Coding agents are typically evaluated with desired behavior already specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce ProgramDistill, a benchmark that evaluates the capabilities of coding agents to implement features discovered through interaction with fully functional reference applications. We build ProgramDistill by factorizing reference applications into features of different granularities, each associated with replayable behaviors executable via the feature gold code patch. Our pipeline, mine-craft-patch, discovers 1,975 replay-verified behaviors across 26 applications and constructs 4,063 tasks without human intervention. Across eight frontier coding agents, Claude Opus 5 falls from 96.0% at depth 1 to 46.0% at depth 8, while full-application reconstruction reaches only 26.1%. ProgramDistill provides a scalable benchmark with a controlled progression for evaluating and diagnosing coding agents, as well as a natural basis for future curriculum-based training.