Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency
Abstract
Derivative-informed training improves neural operators by directly supervising their input--output sensitivities, but existing methods rely on offline tangent solves and stored derivative labels. We propose sketched tangent consistency loss (sTCL), an on-the-fly drop-in derivative regularizer for neural-operator training that improves learned Jacobians without changing the architecture or generating offline sensitivity data. At each training step, sTCL samples a small number of perturbation directions in the input function space, computes the corresponding surrogate Jacobian--vector products using forward-mode automatic differentiation, and penalizes the residual of the forward sensitivity equation. This equation is obtained by differentiating the governing PDE with respect to the input perturbation direction, so derivative supervision is obtained directly from the physics rather than from stored tangent labels. We further show that sketching makes online derivative-informed training computationally feasible, while conditioning is the key to making it effective. Random directional sketching keeps the derivative penalty at the same order of cost as the standard data-fitting loss, making sTCL compatible with stochastic neural-network training as a lightweight loss term. However, raw forward-sensitivity residual penalties can fail for stiff, ill-conditioned, or indefinite tangent operators. To address this, we introduce lightweight operator-aware preconditioners selected by a simple tangent-operator decision rule. Across four PDE benchmarks---Helmholtz, nonlinear diffusion--reaction, Burgers, and two-dimensional Navier--Stokes---operator-conditioned sTCL achieves accuracy comparable to offline DIFNO while eliminating the offline derivative-data generation stage. These results show that on-the-fly derivative-informed training need not merely amortize offline tangent-solve cost into training; with appropriate sketching and conditioning, sTCL provides an attractive drop-in path to derivative-informed neural operators.