A Revisit of Hamiltonian Monte Carlo Efficiency on Bayesian Neural Networks
Abstract
Bayesian neural networks (BNNs) offer principled uncertainty quantification, yet sampling from their posteriors via Hamiltonian Monte Carlo (HMC) remains challenging. While recent work theoretically identified the efficiency degradation of ReLU-based networks, smooth unbounded activations (e.g., GELU, Swish, Mish), which are prevalent in modern architectures, were assumed to be efficient due to their smoothness. In this work, we challenge this existing justification by deriving explicit local error formulas for the leapfrog integrator. Our analyses reveal that, in addition to differentiability, third-order curvatures of the potential energy also have a significant influence on HMC efficiency of BNNs inference. Empirically, we found that even smooth unbounded activations may degrade sampling performance as severely as piecewise linear activations. Thus, optimally tuning the step size for these networks may require a comparable empirical step size scaling as their piecewise counterparts. This result sheds light on the implications for architectural design in Bayesian deep learning.