The Alternation Depth Principle for Neural Operator Design
Haoze Song ⋅ Zhilu Lai ⋅ Wei Wang
Abstract
Despite the success of Neural Operators (NOs), architectural hyperparameters such as depth are still typically selected by empirical grid search, with limited connection to the governing PDE. We introduce the \textit{alternation depth}, $\alpha_F$, a structural quantity of the PDE right-hand side that counts the minimum number of non-pointwise/pointwise coupling stages needed to build its leading nonlinear spatial structure. We use $L_{\min} = \alpha_F + 1$ as a principled \emph{depth-efficiency lower-bound guideline} for IVP solution maps: a starting depth, not a predictor of the empirical optimum $L^*$. We provide partial formal support: (*sufficiency*) an $L$-block architecture can approximate any operator in a formally defined $L$-alternation class under compactness and regularity assumptions; (*separation*) for a constructed hard family at finite resolution and polynomial activations, using fewer blocks forces width to grow as a power of the discretization size. Empirically, a 736-run study over nine in-scope IVP datasets plus one Darcy stress test, using FNO-family variants and one kernel-integral baseline (KIN), is consistent with depth improvements appearing at or above $L_{\min}$ in the tested settings. The observed optimum can be larger because of spectral approximation, pointwise approximation difficulty, data regime, and optimization. We summarize the resulting practitioner guidance in four capacity-allocation rules.
Chat is not available.
Successful Page Load