Amplitude Decoupling in Gaussian Process Training: Exact Decomposition, Pole Cancellation, and Evaluation-Efficient Optimization
Piyush Sao ⋅ Keita Teranishi ⋅ Sudip K Seal ⋅ Pedro Valero-Lara ⋅ Narasinga R Miniskar
Abstract
Gaussian process hyperparameter optimization via marginal likelihood is notoriously brittle. We identify a precise, removable cause of a common source of line-search inefficiency in unprofiled exact GP training: mismatch between the current scalar log-amplitude and its profiled optimum. The GP negative log-likelihood decomposes exactly as $F(a, x) = f(x) + \frac{n}{2}(e^{-r} + r - 1)$, where $r = a - a^*(x)$ is the mismatch between the current log-amplitude $a = \log \sigma_f^2$ and its analytical optimum $a^*(x)$ for the shape parameters $x = (\log \ell, \log \tau)$. The residual $\Psi_n(r) = \frac{n}{2}(e^{-r} + r - 1)$ is a convex, nonnegative penalty that (i) introduces pole singularities at analytic continuations of the kernel matrix past its singular boundary, and (ii) classifies 80% of rejected Armijo trials overall, and up to 100% in high-SNR regimes, as amplitude-caused in our L-BFGS experiments. Profiling out $\sigma_f^2$ eliminates this term exactly, reducing poles to logarithmic branch points and converting stiff line profiles into nearly flat ones. We introduce an Armijo failure diagnostic that exactly decomposes each rejected unprofiled trial into profiled-shape and amplitude-mismatch contributions, determining whether a particular rejection was caused by scalar amplitude mismatch or by the profiled shape objective. Experiments on synthetic regression tasks show that, in the tested L-BFGS/Armijo settings, amplitude profiling reduces the number of objective evaluations needed to reach a target marginal likelihood by 2-5x, with no change to the per-evaluation cost or the global optimum.
Chat is not available.
Successful Page Load