Cauchy Scientific Networks: Loss–Architecture Alignment and Its Limit
Abstract
Cauchy activations and Cauchy-adaptive layers can give large gains on scientific learning tasks, but the gains are selective and not uniquely Cauchy-specific. We organize this selectivity with an empirical loss-architecture alignment taxonomy: an architecture can help when the objective exposes structures it can use, and can lose under a mismatched objective. In score learning, Cauchy score networks improve over SiLU under a Fokker-Planck PDE loss, but the same advantage disappears under denoising score matching and a Gaussian activation is comparable to Cauchy under the PDE loss. In loss-activation factorials, robust residual losses explain more of the heavy-tailed and spectral gains than the activation alone; an Allen-Cahn trap initially escaped by CauchyAct+Cauchy loss is also escaped, sometimes more strongly, by CauchyAct+L1, CauchyAct+Huber, and ReLU+Cauchy loss. We also analyze ResidualCAN, a CauchyAct main path plus a softmax-weighted Cauchy Adaptive Node residual. It achieves 1520x lower stationary Allen-Cahn error than a SiLU MLP at epsilon = 1.0 and remains strongest or second strongest against KAN-style, tuned SIREN/Fourier, and adaptive PINN baselines across an Allen-Cahn epsilon-suite. Ablations show that most of this gain comes from a dense local-basis residual rather than from Cauchy activation alone. A Poisson/Burgers boundary check prevents a broad PDE dominance claim: tuned SIREN/Fourier and KAN-style baselines can dominate away from transition-layer structure. The resulting claim is deliberately bounded: Cauchy components are one useful instance of robust/local-basis co-design in low-dimensional scientific tasks, not a universal activation or a uniquely rational mechanism.