Architecture Outweighs Budget: a Controlled Comparison of Low-Budget AR-to-Diffusion Conversion in MoE LLMs
Abstract
Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion recipes differ by roughly three orders of magnitude in data budget and have not been compared under a common evaluation protocol. We compare two conversions of the same 30B Mixture-of-Experts (MoE) parent using the same recipe, corpus, budget, trainable parameter set, and evaluation harness. The in-place model retrains the parent weights. The frozen-tower model reads a frozen causal copy of the parent through cross-attention. This change raises HumanEval pass@10 from 6.19 to 71.60, an 11.6× improvement. With 1 B training tokens, the frozen-tower model reaches 79% of the score of a published 500 B-token in-place conversion under that model's inference code. This is a cross-study reference at one protocol, and it is a training-budget comparison: batch-size-1 serving remains slower than the AR parent. It also exceeds that model under greedy decoding at every measured budget. Across the masked-diffusion recipes and the 30B in-place arm, the converted models retain 77–82% of parent MMLU-Pro performance, but generation degrades as output length increases. We show that the frozen-tower class contains an exact parent sampler and that its gradient omits the context term present in in-place training. A matched dense-parent experiment reproduces the gap: the frozen-tower and in-place models retain 61.5% and 7.6% of parent HumanEval pass@10, respectively. Evaluation protocol also changes the scores of the 500 B-token model non-monotonically, whereas the AR parent's scores vary by at most two points. Within the tested low-budget regime, conversion architecture is the main constraint. The construction applies to any causal transformer with softmax attention.