Dynamic Causal Structure Discovery for Autoregressive Visual Generation
Songliang Guo ⋅ Qi Yan ⋅ Jianzhou Wang ⋅ Xinfu Liu ⋅ Cheng Zhen ⋅ Yirui Wu ⋅ Lixin Yuan ⋅ Wenxiao Zhang ⋅ Jun Liu
Abstract
Autoregressive (AR) models achieve strong performance in visual generation, yet their prefix-based factorization imposes a rigid one-dimensional dependency structure. We show that this structure is fundamentally suboptimal: any predefined ordering necessarily induces a trade-off between contextual sufficiency and necessity, resulting in both missing and redundant dependencies. To address this limitation, we propose Dynamic Causal Structure Discovery, an inference-time framework that replaces static conditioning with dynamic, content-dependent causal parent sets. By interpreting the Transformer as a causal graph, we reformulate generation as the problem of constructing a minimal parent set for each token. To make this objective tractable, we approximate $\mathrm{do}$-interventions as vector-space ablations and derive a Fisher-based criterion that characterizes and corrects structural deviations during decoding. Our framework transforms AR factorization into an order-agnostic dependency structure, yielding improved generation quality and global coherence without retraining.
Chat is not available.
Successful Page Load