SearcharxivSearch

arXiv subjects

Alexander Bodard

Publications and source records attributed to Alexander Bodard.

12 recordsLinked to original sources

Newton methods beyond Hessian Lipschitz continuity: A nonlinear preconditioning approach

Newton-type methods are typically analyzed under Lipschitz continuity of the Hessian, an assumption that can fail for objectives with higher-order or polynomial growth. We introduce a class of nonlinearly preconditioned Newton methods that apply Newton's root-finding scheme to a transformed optimality mapping, thereby extending recent nonlinear preconditioning ideas from first-order methods to the second-order setting. The resulting methods are naturally analyzed under Lipschitz continuity of a preconditioned Hessian, a condition that significantly relaxes the classical Hessian Lipschitz continuity assumption. Under this generalized smoothness model, we establish local superlinear and quadratic convergence guarantees, and develop a globalization strategy for the nonregularized method despite the fact that the preconditioned Newton direction need not be a descent direction. We further propose a regularized variant for isotropic preconditioners, and show that it attains an $O(\varepsilon^{-3/2})$ iteration complexity. An adaptive version removes the need to know the smoothness constant and allows inexact subproblem solutions while preserving the same complexity order.

math.OC

PANOC-lite: A simpler and more efficient algorithm for composite minimization

This work introduces a simple and efficient linesearch method for composite minimization that accelerates proximal-gradient iterations with fast Newton-type directions. Our algorithm is based on simple operations and only requires the standard proximal-gradient oracle, similar to PANOC and ZeroFPR, provided that the nonsmooth term is convex. Noteworthy improvements include a cheaper backtracking procedure, in the sense that no additional gradients need to be evaluated, and an enlarged range of permitted stepsizes. Global subsequential convergence and local superlinear convergence are established under conventional assumptions by considering a novel merit function which is less expensive to evaluate than alternatives like the forward-backward envelope. Finally, the proposed approach is validated on model predictive control problems with collision avoidance constraints, as well as on the LIBSVM and CUTEst benchmarks.

math.OC

Nonlinearly preconditioned gradient flows

We study a continuous-time dynamical system which arises as the limit of a broad class of nonlinearly preconditioned gradient methods. Under mild assumptions, we establish existence of global solutions and derive Lyapunov-based convergence guarantees. For convex costs, we prove a sublinear decay in a geometry induced by some reference function, and under a generalized gradient-dominance condition we obtain exponential convergence. We further uncover a duality connection with mirror descent, and use it to establish that the flow of interest solves an infinite-horizon optimal-control problem of which the value function is the Bregman divergence generated by the cost. These results clarify the structure and optimization behavior of nonlinearly preconditioned gradient flows and connect them to known continuous-time models in non-Euclidean optimization.

math.OC

Scaled relative graphs for pairs of operators beyond classical monotonicity

We introduce a generalization of the scaled relative graph (SRG) to pairs of operators, enabling the visualization of their relative incremental properties. This novel SRG framework provides the geometric counterpart for the study of nonlinear resolvents based on paired monotonicity conditions. We demonstrate that these conditions apply to linear operators composed with monotone mappings, a class that notably includes NPN transistors, allowing us to compute the response of multivalued, nonsmooth, and highly nonmonotone electrical circuits.

math.OC

Escaping saddle points without Lipschitz smoothness: the power of nonlinear preconditioning

We study generalized smoothness in nonconvex optimization, focusing on $(L_0, L_1)$-smoothness and anisotropic smoothness. The former was empirically derived from practical neural network training examples, while the latter arises naturally in the analysis of nonlinearly preconditioned gradient methods. We introduce a new sufficient condition that encompasses both notions, reveals their close connection, and holds in key applications such as phase retrieval and matrix factorization. Leveraging tools from dynamical systems theory, we then show that nonlinear preconditioning -- including gradient clipping -- preserves the saddle point avoidance property of classical gradient descent. Crucially, the assumptions required for this analysis are actually satisfied in these applications, unlike in classical results that rely on restrictive Lipschitz smoothness conditions. We further analyze a perturbed variant that efficiently attains second-order stationarity with only logarithmic dependence on dimension, matching similar guarantees of classical gradient methods.

math.OC

Second-order methods for provably escaping strict saddle points in composite nonconvex and nonsmooth optimization

This study introduces two second-order methods designed to provably avoid saddle points in composite nonconvex optimization problems: (i) a nonsmooth trust-region method and (ii) a curvilinear linesearch method. These developments are grounded in the forward-backward envelope (FBE), for which we analyze the local second-order differentiability around critical points and establish a novel equivalence between its second-order stationary points and those of the original objective. We show that the proposed algorithms converge to second-order stationary points of the FBE under a mild local smoothness condition on the proximal mapping of the nonsmooth term. Notably, for \( \C^2 \)-partly smooth functions, this condition holds under a standard strict complementarity assumption. To the best of our knowledge, these are the first second-order algorithms that provably escape nonsmooth strict saddle points of composite nonconvex optimization, regardless of the initialization. Our preliminary numerical experiments show promising performance of the developed methods, validating our theoretical foundations.

math.OC

A localized consensus-based sampling algorithm

We propose a localized consensus-based method for sampling from non-Gaussian distributions, a task that frequently arises when solving Bayesian inverse problems. Our method arises from an alternative derivation of consensus-based sampling (CBS). Starting from ensemble-preconditioned Langevin dynamics, we replace the potential by its Moreau envelope -- a smoother approximation -- in order to replace the gradient in the Langevin equation with a proximal operator. We then approximate this operator by a weighted mean. In the limit of infinitely smoothing the potential to a quadratic function, this procedure recovers the standard CBS dynamics. In addition, outside this limit, we retrieve a refined variant of polarized CBS. We call the resulting algorithm localized consensus-based sampling, since particles interact more with nearby particles than with faraway ones. Our method is affine-invariant, exact for Gaussian targets in the mean-field limit, and demonstrates improved robustness over polarized CBS in numerical experiments. Like other consensus-based methods, localized CBS is gradient-free and easily parallelizable.

math.NA

The inexact power augmented Lagrangian method for constrained nonconvex optimization

This work introduces an unconventional inexact augmented Lagrangian method where the augmenting term is a Euclidean norm raised to a power between one and two. The proposed algorithm is applicable to a broad class of constrained nonconvex minimization problems that involve nonlinear equality constraints. In a first part of this work, we conduct a full complexity analysis of the method under a mild regularity condition, leveraging an accelerated first-order algorithm for solving the H\"older-smooth subproblems. Interestingly, this worst-case result indicates that using lower powers for the augmenting term leads to faster constraint satisfaction, albeit with a slower decrease of the dual residual. Notably, our analysis does not assume boundedness of the iterates. Thereafter, we present an inexact proximal point method for solving the weakly-convex and H\"older-smooth subproblems, and demonstrate that the combined scheme attains an improved rate that reduces to the best-known convergence rate whenever the augmenting term is a classical squared Euclidean norm. Different augmenting terms, involving a lower power, further improve the primal complexity at the cost of the dual complexity. Finally, numerical experiments validate the practical performance of unconventional augmenting terms.

math.OC

EM++: A parameter learning framework for stochastic switching systems

This paper proposes a general switching dynamical system model, and a custom majorization-minimization-based algorithm EM++ for identifying its parameters. For certain families of distributions, such as Gaussian distributions, this algorithm reduces to the well-known expectation-maximization method. We prove global convergence of the algorithm under suitable assumptions, thus addressing an important open issue in the switching system identification literature. The effectiveness of both the proposed model and algorithm is validated through extensive numerical experiments.

math.OC

Global Convergence Analysis of the Power Proximal Point and Augmented Lagrangian Method

In this paper we study an unconventional inexact Augmented Lagrangian Method (ALM) for convex optimization problems, as first proposed by Bertsekas, wherein the penalty term is a potentially non-Euclidean norm raised to a power between one and two. We analyze the algorithm through the lens of a nonlinear Proximal Point Method (PPM), as originally introduced by Luque, applied to the dual problem. While Luque analyzes the order of local convergence of the scheme with Euclidean norms our focus is on the non-Euclidean case which prevents us from using standard tools for the analysis such as the nonexpansiveness of the proximal mapping. To allow for errors in the primal update, we derive two implementable stopping criteria under which we analyze both the global and the local convergence rates of the algorithm. More specifically, we show that the method enjoys a fast sublinear global rate in general and a local superlinear rate under suitable growth assumptions. We also highlight that the power ALM can be interpreted as classical ALM with an implicitly defined penalty-parameter schedule, reducing its parameter dependence. Our experiments on a number of relevant problems suggest that for certain powers the method performs similarly to a classical ALM with fine-tuned adaptive penalty rule, despite involving fewer parameters.

math.OC

PANTR: A proximal algorithm with trust-region updates for nonconvex constrained optimization

This work presents PANTR, an efficient solver for nonconvex constrained optimization problems, that is well-suited as an inner solver for an augmented Lagrangian method. The proposed scheme combines forward-backward iterations with solutions to trust-region subproblems: the former ensures global convergence, whereas the latter enables fast update directions. We discuss how the algorithm is able to exploit exact Hessian information of the smooth objective term through a linear Newton approximation, while benefiting from the structure of box-constraints or l1-regularization. An open-source C++ implementation of PANTR is made available as part of the NLP solver library ALPAQA. Finally, the effectiveness of the proposed method is demonstrated in nonlinear model predictive control applications.

math.OC

SPOCK: A proximal method for multistage risk-averse optimal control problems

Risk-averse optimal control problems have gained a lot of attention in the last decade, mostly due to their attractive mathematical properties and practical importance. They can be seen as an interpolation between stochastic and robust optimal control approaches, allowing the designer to trade-off performance for robustness and vice-versa. Due to their stochastic nature, risk-averse problems are of a very large scale, involving millions of decision variables, which poses a challenge in terms of efficient computation. In this work, we propose a splitting for general risk-averse problems and show how to efficiently compute iterates on a GPU-enabled hardware. Moreover, we propose Spock - a new algorithm that utilizes the proposed splitting and takes advantage of the SuperMann scheme combined with fast directions from Anderson's acceleration method for enhanced convergence speed. We implement Spock in Julia as an open-source solver, which is amenable to warm-starting and massive parallelization.

math.OC