A Boltzmann Kernel for Discrete Diffusion: How Locality and Equivariance Explain Creativity
Abstract
Despite the strong generative capabilities of discrete diffusion models, the mechanisms that enable them to generalize beyond their training data remain poorly understood. We study this question through the architectural inductive biases of locality and translation equivariance. For Uniform Discrete Diffusion Models (UDMs), we derive an analytical characterization of the optimal denoiser under these constraints. The resulting predictor takes the form of a Boltzmann-like kernel over training patches, weighting each patch exponentially by its Hamming distance to the noisy input. Experiments on MNIST and CIFAR-10 show close agreement between trained models and the analytical prediction, with discrepancies increasing as receptive-field and vocabulary sizes reduce local data coverage. We further extend our analysis to Masked Diffusion Models (MDMs), where we introduce a relaxed kernel formulation that captures their empirical generalization behaviour, and validate it experimentally.