Gradient Regularized Newton Boosting Trees with Global Convergence
Nikita Zozoulenko ⋅ Daniel Falkowski ⋅ Thomas Cass ⋅ Lukas Gonon
Abstract
Gradient Boosting Decision Trees (GBDTs) dominate tabular machine learning, with modern implementations like XGBoost, LightGBM, and CatBoost being based on Newton boosting: a second-order descent step in the space of decision trees. Despite its empirical success, the global convergence of Newton boosting is poorly understood compared to first-order boosting. In this paper, we introduce Restricted Newton Descent, a framework for convex optimization with Newton's method on Hilbert spaces with inexact iterates, based on the concepts of cosine angle and weak gradient edge. Within this framework, we recover Newton boosting with GBDTs and classical finite-dimensional theory as special cases. We first prove that vanilla Newton boosting achieves a linear rate of convergence for smooth, strongly convex losses that satisfy a Hessian-dominance condition. To handle general convex losses with Lipschitz Hessians, we extend a recent gradient regularized Newton scheme to the restricted weak learner setting, establishing the first global $\mathcal{O}(\frac{1}{k^2})$ convergence rate for second-order GBDTs. This scheme minimally modifies the classical algorithm by introducing an adaptive $\ell_2$-regularization term proportional to the square root of the gradient norm at each iteration. In numerical experiments, we show that this scheme converges while vanilla Newton boosting may diverge.
Chat is not available.
Successful Page Load