Dynamic Quadtree Tokenization for Autoregressive PDE Forecasting
Abstract
The quadratic attention cost of Vision Transformers (ViTs) forces a sharp compromise between discretization and rollout horizon, and binds tightest in the fine-scale PDE regime where shocks, reaction fronts, and material interfaces occupy a small, time-varying fraction of the domain. Further, conventional neural surrogates do not adapt to localized features dynamically. We propose WAMRViT, a ViT that tokenizes its input as a balanced quadtree under a wavelet-inspired refinement criterion, encodes position and refinement level jointly via a 3D rotary positional embedding with a soft level axis, and regrids in cell space at inference to remain stable over long rollouts. A multi-scale variant additionally keeps each leaf at its native source resolution and defers the per-level resolution differential to the model. Unlike prior adaptive-tokenization work, the tokenizer imposes no a priori token count; we evaluate under fully adaptive topology over long autoregressive rollouts; and WAMRViT is the first machine-learning surrogate to natively tokenize multi-level Adaptive Mesh Refinement (AMR) data. We demonstrate that on uniform-grid datasets, WAMRViT exceeds a finest-patch uniform backbone on the region of interest while using a fraction of the tokens, with a structural interpolation cost on the global metric that the multi-scale variant closes at the initial rollout steps. On a complex AMR combustion problem whose extremely fine scale features uniform-grid baselines cannot represent natively and must project onto a coarser grid, WAMRViT operates directly on the adaptive cells and substantially reduces finest-level error at matched parameters. Code: https://anonymous.4open.science/r/wamrvit-review-74E1