SearcharxivSearch

arXiv subjects

Alexander Lim

Publications and source records attributed to Alexander Lim.

3 recordsLinked to original sources

Faithful-Newton Framework: Bridging Inner and Outer Solvers for Enhanced Optimization

Newton-type methods enjoy fast local convergence and strong empirical performance, but achieving global guarantees comparable to first-order methods remains challenging. Even for simple strongly convex problems, no straightforward variant of Newton's method matches the global complexity of gradient descent. While more sophisticated variants can improve iteration complexity, they typically require solving difficult subproblems with high per-iteration costs, leading to worse overall complexity. These limitations stem from treating the subproblem as an afterthought, either as a black box, yielding overly complex and impractical formulations, or in isolation, without regard to its role in advancing the optimization of the main objective. By tightening the integration between the inner iterations of the subproblem solvers and the outer iterations of the optimization algorithm, we introduce simple Newton-type variants, called Faithful-Newton framework, which, in a sense, remain faithful to the overall simplicity of classical Newton's method by retaining simple linear system subproblems. The key conceptual difference, however, is that the quality of the subproblem solution is directly assessed based on its effectiveness in reducing optimality, which in turn enables desirable convergence complexities across a variety of settings. Under standard assumptions, we show that our variants, depending on parameter choices, achieve global superlinear convergence, condition-number-independent linear convergence, and/or local quadratic convergence, even when using inexact Newton steps, for strongly convex problems; and competitive iteration complexity for general convex problems. Numerical experiments further demonstrate that our proposed methods perform competitively compared with several alternative Newton-type approaches.

math.NA

Conjugate Direction Methods Under Inconsistent Systems

Since the development of the conjugate gradient (CG) method in 1952 by Hestenes and Stiefel, CG, has become an indispensable tool in computational mathematics for solving positive definite linear systems. On the other hand, the conjugate residual (CR) method, closely related CG and introduced by Stiefel in 1955 for the same settings, remains relatively less known outside the numerical linear algebra community. Since their inception, these methods -- henceforth collectively referred to as conjugate direction methods -- have been extended beyond positive definite to indefinite, albeit consistent, settings. Going one step further, in this paper, we investigate the theoretical and empirical properties of these methods under inconsistent systems. Among other things, we show that small modifications to the original algorithms allow for the pseudo-inverse solution. Furthermore, we show that CR is essentially equivalent to the minimum residual method, proposed by Paige and Saunders in 1975, in such contexts. Lastly, we conduct a series of numerical experiments to shed lights on their numerical stability (or lack thereof) and their performance for inconsistent systems. Surprisingly, we will demonstrate that, unlike CR and contrary to popular belief, CG can exhibit significant numerical instability, bordering on catastrophe in some instances.

math.NA

Complexity Guarantees for Nonconvex Newton-MR Under Inexact Hessian Information

We consider an extension of the Newton-MR algorithm for nonconvex unconstrained optimization to the settings where Hessian information is approximated. Under a particular noise model on the Hessian matrix, we investigate the iteration and operation complexities of this variant to achieve appropriate sub-optimality criteria in several nonconvex settings. We do this by first considering functions that satisfy the (generalized) Polyak-\L ojasiewicz condition, a special sub-class of nonconvex functions. We show that, under certain conditions, our algorithm achieves global linear convergence rate. We then consider more general nonconvex settings where the rate to obtain first order sub-optimality is shown to be sub-linear. In all these settings, we show that our algorithm converges regardless of the degree of approximation of the Hessian as well as the accuracy of the solution to the sub-problem. Finally, we compare the performance of our algorithm with several alternatives on a few machine learning problems.

math.OC