PixelART: Image-to-Layer Decomposition without Latents or Text-to-Image Pretraining
Abstract
Image-to-layer decomposition converts a flattened image into editable RGBA layers, enabling element-level editing in design workflows. Existing diffusion-based systems typically adapt large pretrained text-to-image models and introduce RGBA autoencoders or variable-layer architectural modules. We revisit this design choice and ask whether layer decomposition actually requires latent autoencoding or T2I pretraining. We introduce PixelART, a pixel-space rectified-flow Transformer trained from scratch for image-to-layer decomposition. PixelART directly denoises regional RGBA pixel patches with a single-stream multi-modal diffusion Transformer, avoiding RGBA-VAEs, pretrained T2I backbones, and layer-specific decoders. We identify a key property of the task: high-noise timesteps determine layer assignment and coarse layer organization, while low-noise timesteps mainly refine color, alpha, texture, and boundaries. Based on this observation, we propose a \emph{terminal-boosted timestep sampling} to increase training coverage in the high-noise assignment regime. Trained on 4M multi-layer design templates up to (1024\times1024), PixelART achieves state-of-the-art layer and composite reconstruction on \designbenchmark while using over (10\times) fewer parameters and substantially lower latency and memory than VAE-based diffusion baselines. Ablations show that pixel-space \mathcao{x}-prediction, high-noise timestep coverage, and data/model scaling are critical, while T2I initialization provides no measurable final gain in the data-rich I2L setting.