Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair
Abstract
Fault localization can focus a code model's repair on the statements a failing test implicates, but a targeted edit may succeed merely because it is small, and a second model call may succeed without using the failure at all. We separate these explanations with three arms applied to the same failed candidate (blind whole-solution resampling, spectrum-based localization followed by suspect-span infilling, and same-length infilling at a disjoint random code span) across three frozen 26-32B models, three benchmarks and 488 failing candidates, plus a separately declared 24B fourth model from a third family. Three results follow. First, localization is rarely available: only 9.0% of failing candidates expose a failing public test with a usable spectrum, because the largest group of failures passes every test the user can see. Second, among the 177 candidates localizable from a strong suite, localized infilling loses decisively to blind resampling at a matched number of attempts (3:40, p = 3.0 x 10^-9), in the direction opposite to our hypothesis. That loss replicates in a third family: on 62 candidates from Mistral-Small-24B, run on two of the three benchmarks, it is -11.3 points (0.6% against 11.9%, 95% CI [-16.6, -6.8], discordance 0:25), the largest margin we measure, so our rule's two-family replication clause is satisfied only in the inverted direction, by the method losing to its own comparator. Third, giving the span arm more room does not rescue it: widening the edit until the median span grows from 1 line to 5 lines leaves it still losing, and an adaptive policy adds at most +0.6 points over always editing wider. Against the random-span placebo localized infilling leads pooled (11:1, Holm-adjusted p = .019), but that lead resolves in no individual model under the per-model attempt-level analysis our shipped analysis plan designates primary (best Holm p = .087), so we report the location effect as suggestive rather than established. Re-pricing attempts as tokens narrows but does not overturn this: a span attempt spends 21.7 generated tokens against 371.1, yet 16 localized attempts reach 6.8% while one blind attempt already reaches 10.1%. The mechanism is visible in the generations: infilling reproduces the removed span verbatim in 48.9% of attempts and yields 0.23 distinct programs per attempt against 0.83 for resampling. A capability probe down to 0.6B cannot carry the location comparison at all, because 65.5% of the span arms' spliced programs do not parse at that scale, so we restrict every localization conclusion to the 24-32B models tested.