Causal Edit Serialization
Abstract
We study \emph{Causal Edit Serialization} (CES), a decoder-only transformer extension that emits explicit edit programs while preserving a standard causal mask. CES supports \emph{insert}, \emph{delete}, \emph{replace}, \emph{move}, and \emph{copy} operations through a serialized edit domain-specific language (DSL), a lightweight localization head, and dual-index rotary position embeddings that separate generation order from source-token addressing. We train on both self-supervised random corruptions and a large GPT-5-mini edit corpus built from ClimbMix documents by generating approximate human-like revisions and parsing them into the same edit DSL; copy examples are added from a separate repeated-span mining pipeline. We compare CES against standard autoregressive training and an autoregressive DSL baseline that emits textualized positions directly. Across model scales and data regimes, CES learns strong source-position localization, improves human-target similarity on WikiAtomic/CoEdIT-style source-target pairs, and produces higher valid-edit rates than the AR-DSL baseline, while incurring only a modest degradation in ordinary language-modeling quality in the best regimes. Arithmetic experiments further show that edit-aware training can improve standard autoregressive generation quality and validity across sampling temperatures, even relative to a clean-data AR baseline. These results suggest that explicit causal editing is a viable way to give decoder-only models compositional control over earlier tokens, both for revision and, in some settings, for generation itself.