SearcharxivSearch

arXiv subjects

Daniel Falkowski

Publications and source records attributed to Daniel Falkowski.

3 recordsLinked to original sources

Gradient Regularized Newton Boosting Trees with Global Convergence

Gradient Boosting Decision Trees (GBDTs) dominate tabular machine learning, with modern implementations like XGBoost, LightGBM, and CatBoost being based on Newton boosting: a second-order descent step in the space of decision trees. Despite its empirical success, the global convergence of Newton boosting is poorly understood compared to first-order boosting. In this paper, we introduce Restricted Newton Descent, which studies convex optimization with Newton's method on Hilbert spaces with inexact iterates, based on the concepts of cosine angle and weak gradient edge. Within this framework, we recover Newton boosting with GBDTs and classical finite-dimensional theory as special cases. We first prove that vanilla Newton boosting achieves a linear rate of convergence for smooth, strongly convex losses that satisfy a Hessian-dominance condition. To handle general convex losses with Lipschitz Hessians, we extend a recent gradient regularized Newton scheme to the restricted weak learner setting. This scheme minimally modifies the classical algorithm by introducing an adaptive $\ell_2$-regularization term proportional to the square root of the gradient norm at each iteration. We establish a $\mathcal{O}(\frac{1}{k^2})$ rate for this scheme, thereby obtaining a globally convergent second-order GBDT algorithm with a rate matching that of first-order boosting with Nesterov momentum. In numerical experiments, we show that our scheme converges while vanilla Newton boosting may diverge.

stat.ML

Conjugating by singular operators: On the boundedness of similarity transforms near singular points

We consider the question of, given operators $A$, $Z$ and a sequence of invertible operators $U_n\to Z$, whether the sequence $U_nAU_n^{-1}$ is bounded in norm, as well as generalizations of this where $U_nAU_n^{-1}$ is modified by some bounded linear map on bounded linear operators. In the setting of Hilbert spaces, we provide a complete classification in terms of algebraic criteria of those $A$ for which such a sequence exists, as long as $Z$ is of generalized index zero, which always holds in finite-dimensional contexts. In the process, we prove that particular coefficients arising in inverses of certain good paths going to $Z$ can also be classified in terms of an entirely algebraic criterion.

math.FA

On joint eigen-decomposition of matrices

The problem of approximate joint diagonalization of a collection of matrices arises in a number of diverse engineering and signal processing problems. This problem is usually cast as an optimization problem, and it is the main goal of this publication to provide a theoretical study of the corresponding cost-functional. As our main result, we prove that this functional tends to infinity in the vicinity of rank-deficient matrices with probability one, thereby proving that the optimization problem is well posed. Secondly, we provide unified expressions for its higher-order derivatives in multilinear form, and explicit expressions for the gradient and the Hessian of the functional in standard form, thereby opening for new improved numerical schemes for the solution of the joint diagonalization problem. A special section is devoted to the important case of self-adjoint matrices.

math.NA