BitDance: Scaling Autoregressive Generative Models with Binary Tokens
Yuang Ai ⋅ Jiaming Han ⋅ Shaobin Zhuang ⋅ Weijia Mao ⋅ Xuefeng Hu ⋅ Ziyan Yang ⋅ Zhenheng Yang ⋅ Yali Wang ⋅ Xiangyu Yue ⋅ Hao Chen ⋅ Huaibo Huang
Abstract
We present BitDance, a scalable autoregressive (AR) image generator that predicts binary visual tokens instead of codebook indices. With high-entropy binary latents, BitDance lets each token represent up to $\mathbf{2^{256}}$ states, yielding a compact yet highly expressive discrete representation. Sampling from such a huge token space is difficult with standard classification. To resolve this, BitDance uses a binary diffusion head: instead of predicting an index with softmax, it employs continuous-space diffusion to generate the binary tokens. Furthermore, we propose next-patch diffusion, a new decoding method that predicts multiple tokens in parallel with high accuracy, greatly speeding up inference. On ImageNet 256$\times$256, BitDance achieves an FID of **1.24**, the best among AR models. With next-patch diffusion, BitDance beats state-of-the-art parallel AR models that use 1.4B parameters, while using **5.4$\times$** fewer parameters (260M) and achieving **8.7$\times$** speedup. For text-to-image generation and image editing, BitDance trains on large-scale multimodal tokens and generates high-resolution, photorealistic images efficiently, showing strong performance and favorable scaling. When generating 1024$\times$1024 images, BitDance achieves a speedup of over **30$\times$** compared to prior AR models. We release code and models to facilitate further research on AR foundation models.
Chat is not available.
Successful Page Load