KerONet: A Single Softmax Readout Suffices for Physics-Informed Operator Learning
Seth Dale ⋅ Carolyn Koh ⋅ Dinesh Mehta
Abstract
Operator learning has experienced a shift toward attention-based architectures. This trend has more recently extended to physics-informed operator learning, where stacked cross-attention blocks add considerable computational cost relative to the foundational PI-DeepONet. We ask whether this complexity is necessary. Starting from the DeepONet branch-trunk decomposition, we replace the bilinear readout with a single softmax-weighted readout in which the trunk emits a query vector and the branch emits a set of key-value pairs—a structure we show admits universal approximation via reduction to a simplex-bilinear form. The resulting architecture, KerONet, uses no iterated attention, layer normalization, or explicit learnable projection matrices. Under a unified hyperparameter optimization protocol with a strict $200{,}000$-parameter budget, we evaluate KerONet alongside iterated-attention architectures (PIT, PINTO) and PI-DeepONet on four canonical 1D PDEs whose solutions range from globally smooth to sharply localized. KerONet outperforms PIT and PINTO on every PDE, across all five seeds, in accuracy (approximately $2$–$4\times$ lower rel-$L^2$ error), reproducibility ($2$–$31\times$ lower standard deviation), and speed ($1.8$–$4.4\times$ faster training time). Against PI-DeepONet, KerONet shows substantial gains where input-dependent localized features dominate ($2.5$–$8\times$ lower error on Burgers and Allen-Cahn) while exhibiting comparable accuracy on smoother operators (advection and diffusion-reaction)–a balance the iterated-attention architectures do not achieve. Our findings demonstrate that iterated cross-attention is not necessary to realize the benefits of query-dependent feature mixing in physics-informed operator learning.
Chat is not available.
Successful Page Load