Same Outcome, Different Behavior: Interpreting Procedural Transfer Across Trust Boundaries in Coding Agents
Anjum Asiya ⋅ Shashwat Suthar ⋅ Nayyar Zaidi
Abstract
A coding procedure can be correct in one setting but become unsafe when it is reused in a new setting where an important assumption has changed. We study this problem in coding agents that are given correct procedural memory from an earlier task and asked to solve a similar target task with a changed security-relevant assumption. Our controlled experiment spans 13 synthetic task families, four memory conditions, two repetitions, and six completed model and agent configurations: 624 recorded runs and 595 technically valid outcomes. We measure functionality pass ($F$) and focal security witness pass ($S$) separately; $U = F \land \neg S$ records functional completion with a focal witness failure. Correct source memory versus no memory (C$-$N) changes $U$ by $-11.5$ to $+13.6$ percentage points across configurations. Estimates are positive, zero, and negative; every corresponding exploratory bootstrap interval includes zero. Lower $U$ can accompany lower functionality. In a 52-run sample, with the coding protocol and sampling rule frozen after the experiment and initial audits but before annotation, two independent model-assisted coding passes agree on 7 runs with explicit assumption recognition before editing, 20 with recognizable source procedure retention without the target check, 22 with concrete target protection, and 7 with no meaningful implementation. Fifteen sampled runs have agreed protection without agreed prior explicit recognition. Matched traces and patches show unchecked reuse, successful adaptation, and different implementation paths reaching the same focal endpoint. These observations do not establish whether memory caused an edit. Adding a generic applicability reminder (B$-$C) yields negative $U$ estimates in all three Codex configurations, including under fixed-cohort missingness bounds; MiniSWE does not share that pattern. This does not establish model effects. Within these synthetic shifted targets, agent evaluation should combine endpoint outcomes with trajectory and implementation evidence: behaviorally distinct routes can reach the same measured outcome.
Chat is not available.
Successful Page Load