SpanBench: When Does Bidirectional Generation Help? Controlled Multi-Span Infilling and Local Editing
Abstract
Masked-diffusion language models (dLLMs) are often credited with two advantages over autoregressive (AR) generation: coordinating edits across multiple spans, and editing locally without collateral damage. We test both with SpanBench, which varies the number of dependent spans (k up to 8 in the controlled code grid; prose adds intermediate k), span distance, and mask budget over 1,132 development instances verified by execution or matching, and evaluates a 7B masked-diffusion model against three AR protocols with measured native-interface NFE and same-hardware wall time. We find that high-k coordination failure is not inevitable for AR: sequential fill-in-the-middle succeeds on 1/29 k=8 instances, while masked diffusion succeeds on 21/29 and a single joint AR proposal on 18/29, before any verifier feedback. Verifier-conditioned revision makes the AR system the strongest overall at comparable measured NFE and batch-1 latency (analytic FLOPs favor AR; Appendix D), whereas fresh resampling and verifier-conditioned remasking add at most 1.2 points to diffusion, whose outcomes are highly stable across three runs (98.3% identical in all). The zero-unintended-change guarantee turns out to be a property of the editing interface: an AR span-patch control (same instruct checkpoint as repair) inherits it, since bytes outside the licensed spans are unchanged by construction. Success under that matched scope still favors diffusion (89.4% vs. 77.3%; exact McNemar p=0.009; clustered p=0.008), because the AR model reconstructs the original text instead of executing the edit; its stronger full-document repair costs 5.2x the measured forward evaluations (332 vs. 64) and carries no guarantee. Multi-span behavior, we conclude, depends on generation protocol and edit scope at least as much as on the AR-versus-diffusion choice. On a frozen, source-disjoint held-out set (857 instances), the system ranking and direction of the k=4 divergence reproduce; the locality contrast is direction-consistent. The generator, content-hash-pinned datasets, per-instance outputs, and all evaluation code are released anonymously (Appendix D). Anonymized artifact : https://osf.io/3utx4/overview?view_only=6167b3681178418dab7e5d230649912e