Diffusion-Time Concept Manifolds: Sparse Autoencoder Groups for Interpreting Denoising Language Models
Abstract
Diffusion language models generate text through iterative denoising, but current interpretability tools mostly analyze static hidden states and do not explain how concept information evolves across denoising time. We introduce DynaManifold-SAE, a method for discovering sparse autoencoder feature groups whose activations form time-dependent concept manifolds in masked and diffusion language models. The method builds sparse latent codes across mask ratios, constructs candidate feature groups, and evaluates them with held-out coordinate prediction, geodesic consistency, persistence across time, and causal interventions. Across BERT and LLaDA, DynaManifold-SAE reliably identifies sparse sentiment geometry that is stable across seeds and stronger than individual-feature, PCA, random-group, and graph-structured baselines. The discovered groups transfer from controlled templates to natural SST-2 examples, indicating that they capture semantic structure rather than template artifacts. We further show that these groups explain a denoising-specific mechanism: they predict token reveal order and confidence growth during LLaDA unmasking, and targeted interventions alter reveal confidence and token recovery more than matched controls. These results suggest that sparse feature manifolds provide a practical bridge between static representation geometry and the dynamics of diffusion language generation.