Logit-Conditioned Diffusion Decoding for Frozen Discrete-Token VLMs
Abstract
Pre-trained discrete-token vision-language models (VLMs) generate images by sampling sequences of codebook indices, but their visual fidelity is bounded by the VQ-VAE decoding stage. Prior work addresses this bottleneck by replacing the discrete tokenizer with continuous or hybrid visual token representations, or by integrating diffusion into the generation pipeline; both routes require foundation-scale retraining or realignment of VLM, which is undesirable when the VLM's broader capabilities should be preserved or when retraining is infeasible. We show that such retraining is unnecessary: before each token is sampled, discrete-token VLMs already compute a full logit distribution over the codebook, but standard decoding discards it after selecting a single index. Under Otsu's threshold, the pre-sampling logits of capable VLMs assign statistically meaningful probability to at least two codebook entries per token, and recovering this distribution can substantially close the fidelity gap. We propose DCDD, a post-hoc framework that conditions a diffusion decoder on this discarded distribution while leaving the VLM frozen. At inference, DCDD converts the VLM's pre-sampling logits into distribution-weighted code vectors; during training, it uses VQ-VAE encoder-derived proxy logits calibrated to match the distributional statistics of the VLM's inference-time logits. Our diffusion decoder then maps these representations to high-fidelity images as an optional alternative to the native VQ-VAE decoder. Trained on ImageNet-1K for 50K steps, DCDD reduces reconstruction FID by 74\% and Janus-Pro 7B's generation FID on MJHQ-30K by 34\%, while semantic-alignment across GenEval, DPG-Bench, and WISE remain consistent. Comparison against a diffusion model from the same backbone family applied as an image-to-image refiner confirms that the gain comes not from the diffusion model alone, but from conditioning on the VLM's pre-sampling logit distribution.