SearcharxivSearch

arXiv subjects

Karl Welzel

Publications and source records attributed to Karl Welzel.

4 recordsLinked to original sources

Local Convergence of Adaptively Regularized Tensor Methods

Optimization methods that make use of derivatives of the objective function up to order $p > 2$ are called tensor methods. Among them, ones that minimize a regularized $p$th-order Taylor expansion at each step have been shown to possess optimal global complexity, which improves as $p$ increases. The local convergence of such optimization algorithms on functions that have Lipschitz continuous $p$th derivatives and are uniformly convex of order $q$ has been studied by Doikov and Nesterov [Math. Program., 193 (2022), pp. 315--336]. We extend these local convergence results to locally uniformly convex functions and fully adaptive methods, which do not need knowledge of the Lipschitz constant, thus providing the first sharp local rates for AR$p$. We discuss the surprising new challenges encountered by nonconvex local models and non-unique model minimizers. For $p > 2$, our examples show that in particular when using the global minimizer of the subproblem, even asymptotically not all iterations need to be successful. Only if the "right" local model minimizer is used, the $p/(q-1)$th-order local convergence from the non-adaptive case is preserved for $p > q-1$, otherwise the superlinear rate can degrade. We thus confirm that adaptive higher-order methods achieve superlinear convergence for certain degenerate problems as long as $p$ is large enough and provide sharp bounds on the order of convergence one can expect in the limit.

math.OC

On Global Rates for Regularization Methods based on Secant Derivative Approximations

An inexact and globally convergent framework for high-order adaptive regularization methods is presented, in which approximations may be used for the $p$th-order tensor, based on lower-order derivatives. Between each recalculation of the $p$th-order derivative approximation, a high-order secant equation can be used to update the $p$th-order tensor as proposed in (Karl Welzel and Raphael A Hauser, Approximating higher-order derivative tensors using secant updates, SIAM J.Optim, 34(1), 2024) or the approximation can be kept constant in a lazy manner. When refreshing the $p$th-order tensor approximation after $m$ steps, an exact evaluation of the tensor or a finite difference approximation can be used with an explicit discretization stepsize. For all the newly adaptive regularization variants, we prove an $\mathcal{O}\left( \max[ \epsilon_1^{-(p+1)/p}, \, {\epsilon_2^{-(p+1)/(p-1)}} ] \right)$ bound on the number of iterations needed to reach an $(\epsilon_1, \, \epsilon_2)$ second-order stationary points. Discussions on the number of oracle calls for each introduced variant are also provided. When $p=2$, we obtain a second-order method that uses quasi-Newton approximations with an $\mathcal{O}\left(\max[\epsilon_1^{-3/2}, \, \, \epsilon_2^{-3}]\right)$ iteration bound to achieve approximate second-order stationarity. Numerical illustrations for the case $p=3$ are provided in both the deterministic and noisy settings showcasing the merits of secant updates for approximating third-order information, as well as the robustness of our proposed method even in noisy cases.

math.OC

Efficient Implementation of Third-Order Tensor Methods with Adaptive Regularization for Unconstrained Optimization

High-order tensor methods that employ local Taylor models of degree $p$ within adaptive regularization frameworks (AR$p$) have recently received significant attention, due to their optimal/improved global and local rates of convergence, for both convex and nonconvex optimization problems. In this paper, we showcase the numerical performance of standard second- and third-order variants ($p=2,3$) and propose novel techniques for key algorithmic aspects when $p\geq 3$. In particular, we extend the interpolation-based updating strategy for the regularization parameter introduced in [Gould, Porcelli and Toint, Comput Optim Appl (2012) 53:1--22] for $p=2$, to the case when $p \geq 3$. We identify fundamental differences between the different local minima of the regularised subproblems for $p=2$ and $p \geq 3$ and their effect on algorithm performance. For $p\geq 3$, we introduce a novel pre-rejection technique that rejects poor/unsuccessful subproblem minimizers prior to any function evaluation. Numerical studies showcase the efficiency improvements generated by our proposed modifications of the AR$3$ algorithm. We also assess numerically, the effect of different subproblem termination conditions and the choice of the initial regularization parameter on the overall algorithm performance. Finally, we benchmark our best-performing AR$3$ variants, as well as those in [Birgin et al., Optim Lett (2020) 14:815--838], against second-order ones (AR$2$). Encouraging results on standard test problems are obtained, confirming that AR$3$ variants can be made to outperform second-order variants in terms of objective evaluations, derivative evaluations, and number of subproblem solves. We provide an efficient, extensive and modular software package in MATLAB that includes many AR$2$ and AR$3$ variants, including Hessian- and tensor-free ones, allowing ease of use and experimentation for interested users.

math.OC

Approximating Higher-Order Derivative Tensors Using Secant Updates

Quasi-Newton methods employ an update rule that gradually improves the Hessian approximation using the already available gradient evaluations. We propose higher-order secant updates which generalize this idea to higher-order derivatives, approximating for example third derivatives (which are tensors) from given Hessian evaluations. Our generalization is based on the observation that quasi-Newton updates are least-change updates satisfying the secant equation, with different methods using different norms to measure the size of the change. We present a full characterization for least-change updates in weighted Frobenius norms (satisfying an analogue of the secant equation) for derivatives of arbitrary order. Moreover, we establish convergence of the approximations to the true derivative under standard assumptions and explore the quality of the generated approximations in numerical experiments.

math.OC