FLARE: Diffusion for Hybrid Language Model
Yuchen Zhu ⋅ Jing Shi ⋅ Chongjian GE ⋅ Hao Tan ⋅ Yiran Xu ⋅ Wanrong Zhu ⋅ Jason Kuen ⋅ Ryan Rossi ⋅ Koustava Goswami ⋅ Rajiv Jain ⋅ Yongxin Chen ⋅ Molei Tao ⋅ Jiuxiang Gu
Abstract
Autoregressive (AR) large language models have achieved broad practical success, but their sequential decoding remains a major bottleneck for low-latency deployment. Efficiency efforts have advanced along two largely orthogonal axes: hybrid attention architectures that reduce the cost of each forward pass, and diffusion language models (dLLMs) that enable parallel token generation. Existing dLLMs, however, have yet to translate their theoretical decoding parallelism into commensurate real-world throughput, have not incorporated modern hybrid attention backbones, and still trail scale-matched AR models on quality. In this work, we present FLARE, a systematic recipe for converting hybrid AR LLMs into capable, real-time-fast dLLMs under a practical training budget. Through a controlled study, we identify transfer data quality as the dominant driver of AR-to-dLLM performance loss, outweighing loss formulation and attention-mask design of which are emphasized by prior works. To enable efficient training and deployment of a hybrid softmax-plus-linear-attention backbone, we develop hardware-aware Triton kernels for diffusion-style linear attention and an SGLang-based inference system that exposes both AR-Trust and Diffusion-Trust decoding from the same checkpoint. Starting from Qwen3.5 checkpoints with only 10B tokens from public datasets, FLARE-9B matches the top-tier open-source dLLM LLaDA-2.1 Flash (100B-A5B) at 1/10 of the parameters, exceeding it on GPQA-Diamond ($71.2$ vs. $66.7$), and FLARE-2B reaches $2.2$ times the throughput of LLaDA-2.1 mini (16B-A1B) on GSM8K at eight-way concurrency on a single A100.
Chat is not available.
Successful Page Load