Simplifying Transformer-Based U-Net Neural Physics Simulators
Pietro Sittoni ⋅ Francesco Tudisco
Abstract
U-Net--style architectures are widely adopted for modeling physical systems, as their multiscale structure enables efficient processing of high-resolution data while reflecting the hierarchical structure of many physical phenomena. The most successful U-Net-based neural physics simulators often combine convolutions, Fourier layers, or specialized transformer mechanisms such as windowed or axial attention; however, these structured components can limit adaptability across different spatial and spatiotemporal dimensions. In this work, we revisit transformer-based U-Nets with the goal of simplifying attention-based neural physics simulators without sacrificing multiscale efficiency or predictive accuracy. We introduce \texttt{UFlex}, a simple attention-based U-Net whose blocks are obtained through only minimal modifications to standard self-attention, allowing the same architecture to be applied across one-, two-, three-dimensional, and spatiotemporal regular-grid problems without major design changes. Evaluated on seven challenging benchmarks (four 2D and three 3D), \texttt{UFlex} scales to resolutions of up to $512 \times 512$ in 2D and $256 \times 128 \times 256$ in 3D, while reducing training memory and accelerating training compared to state-of-the-art transformer baselines, all while achieving state-of-the-art predictive accuracy.
Chat is not available.
Successful Page Load