Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models
John Sweeney
Abstract
Swapping two LLM post-training stages can change the final model, but an aggregate loss or benchmark delta only says that something moved, not where the order-dependent residue landed. We ask whether this residue leaves a memory trace: a structured signal that is localized in output space, changes held-out order-gap NLL under targeted interventions, and is assignable from paired endpoint weights. For a specified two-stage path $(A,B)$, we define commutator memory by projecting the Lie bracket $b_{AB} := H_B g_A - H_A g_B$ at the base model through the logits into token scores $\tau_k=\mathbb{E}[e_k\,\delta z_k]$, whose sum is the leading bracket-predicted order effect. The token readout is localized: endpoint and batch-resampled checks preserve bracket top-token supports far more than norm-matched random directions (82--99% vs. 35--49% top-20 overlap), and a separate endpoint-support prediction check recovers 53--57% of Qwen top-1% active-token supports. It has intervention leverage in the tested protocols: top-positive-$\tau_k$ token interventions close 32% (median) of the held-out Qwen-3-4B SFT ordering gap with tau-neutral controls near zero, with the same bracket-vs-control separation on a pre-specified above-floor Qwen-2.5-1.5B fp32 subset. It is assignable from paired endpoints: given a base model, candidate stages, and paired alternate-order endpoints, the statistic $\langle\theta_{AB}-\theta_{BA},\,b_{AB}\rangle$ distinguishes $k=1$ SGD training order in 66/72 pair-seed units across four LLMs. Matched-batch DPO and frozen-rollout reward-surrogate experiments are matched-path stress tests; AdamW is reported as an analogous lifted-state pilot. Commutator memory turns training order from a scalar nuisance into a localized, intervention-tested, and paired-endpoint-readable residue.
Chat is not available.
Successful Page Load