Rethinking Attention in Depth for Operator Learning
Abstract
Transformers have rapidly emerged as powerful surrogate models for learning solution operators for partial differential equations (PDEs). Standard attention tends to exhibit depth-wise representational bottlenecks across layers, with the role of depth remaining underexplored in modeling complex PDEs. In this paper, we rethink the role of depth by promoting it to an explicit computational dimension of attention kernel construction, rather than a passive architectural parameter. With this strategy, each attention block's native representation is enriched via a gated mixture of inter-layer differences, enabling information to be coupled and propagated across layers. The depth-wise hierarchy admits a theoretically motivated multi-kernel interpretation with weakly correlated kernels and improved generalization bounds. Beyond dense feature aggregation, our approach yields a more expressive form of depth-wise attention, mitigating spectral collapse, preserving interaction-scale diversity, and improving intra- and inter-layer route utilization. Extensive empirical evaluations demonstrate its broad compatibility across state-of-the-art solvers, presenting consistent performance gains on eight PDE benchmarks.