Searcharxiv⌕ Search

arXiv subjects

Yurii Nesterov

Publications and source records attributed to Yurii Nesterov.

At least 37 records · Page 2Linked to original sources

High-order methods beyond the classical complexity bounds, II: inexact high-order proximal-point methods with segment search

A bi-level optimization framework (BiOPT) was proposed in [3] for convex composite optimization, which is a generalization of bi-level unconstrained minimization framework (BLUM) given in [20]. In this continuation paper, we introduce a $p$th-order proximal-point segment search operator which is used to develop novel accelerated methods. Our first algorithm combines the exact element of this operator with the estimating sequence technique to derive new iteration points, which are shown to attain the convergence rate $\mathcal{O}(k^{-(3p+1)/2})$, for the iteration counter $k$. We next consider inexact elements of the high-order proximal-point segment search operator to be employed in the BiOPT framework. We particularly apply the accelerated high-order proximal-point method at the upper level, while we find approximate solutions of the proximal-point segment search auxiliary problem by a combination of non-Euclidean composite gradient and bisection methods. For $q=\lfloor p/2 \rfloor$, this amounts to a $2q$th-order method with the convergence rate $\mathcal{O}(k^{-(6q+1)/2})$ (the same as the optimal bound of $2q$th-order methods) for even $p$ ($p=2q$) and the superfast convergence rate $\mathcal{O}(k^{-(3q+1)})$ for odd $p$ ($p=2q+1$).

math.OC↗

High-order methods beyond the classical complexity bounds, I: inexact high-order proximal-point methods

In this paper, we introduce a \textit{Bi-level OPTimization} (BiOPT) framework for minimizing the sum of two convex functions, where both can be nonsmooth. The BiOPT framework involves two levels of methodologies. At the upper level of BiOPT, we first regularize the objective by a $(p+1)$th-order proximal term and then develop the generic inexact high-order proximal-point scheme and its acceleration using the standard estimation sequence technique. At the lower level, we solve the corresponding $p$th-order proximal auxiliary problem inexactly either by one iteration of the $p$th-order tensor method or by a lower-order non-Euclidean composite gradient scheme with the complexity $\mathcal{O}(\log \tfrac{1}{\varepsilon})$, for the accuracy parameter $\varepsilon>0$. Ultimately, if the accelerated proximal-point method is applied at the upper level, and the auxiliary problem is handled by a non-Euclidean composite gradient scheme, then we end up with a $2q$-order method with the convergence rate $\mathcal{O}(k^{-(p+1)})$, for $q=\lfloor p/2 \rfloor$, where $k$ is the iteration counter.

math.OC↗

Subgradient Ellipsoid Method for Nonsmooth Convex Problems

In this paper, we present a new ellipsoid-type algorithm for solving nonsmooth problems with convex structure. Examples of such problems include nonsmooth convex minimization problems, convex-concave saddle-point problems and variational inequalities with monotone operator. Our algorithm can be seen as a combination of the standard Subgradient and Ellipsoid methods. However, in contrast to the latter one, the proposed method has a reasonable convergence rate even when the dimensionality of the problem is sufficiently large. For generating accuracy certificates in our algorithm, we propose an efficient technique, which ameliorates the previously known recipes.

math.OC↗

On Inexact Solution of Auxiliary Problems in Tensor Methods for Convex Optimization

In this paper we study the auxiliary problems that appear in $p$-order tensor methods for unconstrained minimization of convex functions with $ν$-Hölder continuous $p$th derivatives. This type of auxiliary problems corresponds to the minimization of a $(p+ν)$-order regularization of the $p$th order Taylor approximation of the objective. For the case $p=3$, we consider the use of Gradient Methods with Bregman distance. When the regularization parameter is sufficiently large, we prove that the referred methods take at most $\mathcal{O}(\log(ε^{-1}))$ iterations to find either a suitable approximate stationary point of the tensor model or an $ε$-approximate stationary point of the original objective function.

math.OC↗

Tensor Methods for Finding Approximate Stationary Points of Convex Functions

In this paper we consider the problem of finding $ε$-approximate stationary points of convex functions that are $p$-times differentiable with $ν$-Hölder continuous $p$th derivatives. We present tensor methods with and without acceleration. Specifically, we show that the non-accelerated schemes take at most $\mathcal{O}\left(ε^{-1/(p+ν-1)}\right)$ iterations to reduce the norm of the gradient of the objective below a given $ε\in (0,1)$. For accelerated tensor schemes we establish improved complexity bounds of $\mathcal{O}\left(ε^{-(p+ν)/[(p+ν-1)(p+ν+1)]}\right)$ and $\mathcal{O}\left(|\log(ε)|ε^{-1/(p+ν)}\right)$, when the Hölder parameter $ν\in [0,1]$ is known. For the case in which $ν$ is unknown, we obtain a bound of $\mathcal{O}\left(ε^{-(p+1)/[(p+ν-1)(p+2)]}\right)$ for a universal accelerated scheme. Finally, we also obtain a lower complexity bound of $\mathcal{O}\left(ε^{-2/[3(p+ν)-2]}\right)$ for finding $ε$-approximate stationary points using $p$-order tensor methods.

math.OC↗

Tensor Methods for Minimizing Convex Functions with Hölder Continuous Higher-Order Derivatives

In this paper we study $p$-order methods for unconstrained minimization of convex functions that are $p$-times differentiable ($p\geq 2$) with $ν$-Hölder continuous $p$th derivatives. We propose tensor schemes with and without acceleration. For the schemes without acceleration, we establish iteration complexity bounds of $\mathcal{O}\left(ε^{-1/(p+ν-1)}\right)$ for reducing the functional residual below a given $ε\in (0,1)$. Assuming that $ν$ is known, we obtain an improved complexity bound of $\mathcal{O}\left(ε^{-1/(p+ν)}\right)$ for the corresponding accelerated scheme. For the case in which $ν$ is unknown, we present a universal accelerated tensor scheme with iteration complexity of $\mathcal{O}\left(ε^{-p/[(p+1)(p+ν-1)]}\right)$. A lower complexity bound of $\mathcal{O}\left(ε^{-2/[3(p+ν)-2]}\right)$ is also obtained for this problem class.

math.OC↗

Rates of superlinear convergence for classical quasi-Newton methods

We study the local convergence of classical quasi-Newton methods for nonlinear optimization. Although it was well established a long time ago that asymptotically these methods converge superlinearly, the corresponding rates of convergence still remain unknown. In this paper, we address this problem. We obtain first explicit non-asymptotic rates of superlinear convergence for the standard quasi-Newton methods, which are based on the updating formulas from the convex Broyden class. In particular, for the well-known DFP and BFGS methods, we obtain the rates of the form $(\frac{n L^2}{μ^2 k})^{k/2}$ and $(\frac{n L}{μk})^{k/2}$ respectively, where $k$ is the iteration counter, $n$ is the dimension of the problem, $μ$ is the strong convexity parameter, and $L$ is the Lipschitz constant of the gradient.

math.OC↗

Smoothness parameter of power of Euclidean norm

In this paper, we study derivatives of powers of Euclidean norm. We prove their Hölder continuity and establish explicit expressions for the corresponding constants. We show that these constants are optimal for odd derivatives and at most two times suboptimal for the even ones. In the particular case of integer powers, when the Hölder continuity transforms into the Lipschitz continuity, we improve this result and obtain the optimal constants.

math.OC↗

New Results on Superlinear Convergence of Classical Quasi-Newton Methods

We present a new theoretical analysis of local superlinear convergence of classical quasi-Newton methods from the convex Broyden class. As a result, we obtain a significant improvement in the currently known estimates of the convergence rates for these methods. In particular, we show that the corresponding rate of the Broyden-Fletcher-Goldfarb-Shanno method depends only on the product of the dimensionality of the problem and the logarithm of its condition number.

math.OC↗

Greedy Quasi-Newton Methods with Explicit Superlinear Convergence

In this paper, we study greedy variants of quasi-Newton methods. They are based on the updating formulas from a certain subclass of the Broyden family. In particular, this subclass includes the well-known DFP, BFGS and SR1 updates. However, in contrast to the classical quasi-Newton methods, which use the difference of successive iterates for updating the Hessian approximations, our methods apply basis vectors, greedily selected so as to maximize a certain measure of progress. For greedy quasi-Newton methods, we establish an explicit non-asymptotic bound on their rate of local superlinear convergence, which contains a contraction factor, depending on the square of the iteration counter. We also show that these methods produce Hessian approximations whose deviation from the exact Hessians linearly convergences to zero.

math.OC↗

Contracting Proximal Methods for Smooth Convex Optimization

In this paper, we propose new accelerated methods for smooth convex optimization, called contracting proximal methods. At every step of these methods, we need to minimize a contracted version of the objective function augmented by a regularization term in the form of Bregman divergence. We provide global convergence analysis for a general scheme admitting inexactness in solving the auxiliary subproblem. In the case of using for this purpose high-order tensor methods, we demonstrate an acceleration effect for both convex and uniformly convex composite objective functions. Thus, our construction explains acceleration for methods of any order starting from one. The augmentation of the number of calls of oracle due to computing the contracted proximal steps is limited by the logarithmic factor in the worst-case complexity bound.

math.OC↗

Local convergence of tensor methods

In this paper, we study local convergence of high-order Tensor Methods for solving convex optimization problems with composite objective. We justify local superlinear convergence under the assumption of uniform convexity of the smooth component, having Lipschitz-continuous high-order derivative. The convergence both in function value and in the norm of minimal subgradient is established. Global complexity bounds for the Composite Tensor Method in convex and uniformly convex cases are also discussed. Lastly, we show how local convergence of the methods can be globalized using the inexact proximal iterations.

math.OC↗

Minimizing Uniformly Convex Functions by Cubic Regularization of Newton Method

In this paper, we study the iteration complexity of cubic regularization of Newton method for solving composite minimization problems with uniformly convex objective. We introduce the notion of second-order condition number of a certain degree and justify the linear rate of convergence in a nondegenerate case for the method with an adaptive estimate of the regularization parameter. The algorithm automatically achieves the best possible global complexity bound among different problem classes of uniformly convex objective functions with Hölder continuous Hessian of the smooth part of the objective. As a byproduct of our developments, we justify an intuitively plausible result that the global iteration complexity of the Newton method is always better than that of the gradient method on the class of strongly convex functions with uniformly bounded second derivative.

math.OC↗

Gradient Methods with Memory

In this paper, we consider gradient methods for minimizing smooth convex functions, which employ the information obtained at the previous iterations in order to accelerate the convergence towards the optimal solution. This information is used in the form of a piece-wise linear model of the objective function, which provides us with much better prediction abilities as compared with the standard linear model. To the best of our knowledge, this approach was never really applied in Convex Minimization to differentiable functions in view of the high complexity of the corresponding auxiliary problems. However, we show that all necessary computations can be done very efficiently. Consequently, we get new optimization methods, which are better than the usual Gradient Methods both in the number of oracle calls and in the computational time. Our theoretical conclusions are confirmed by preliminary computational experiments.

math.OC↗

Optimization Methods for Fully Composite Problems

In this paper, we propose a new Fully Composite Formulation of convex optimization problems. It includes, as a particular case, the problems with functional constraints, max-type minimization problems, and problems of Composite Minimization, where the objective can have simple nondifferentiable components. We treat all these formulations in a unified way, highlighting the existence of very natural optimization schemes of different order. We prove the global convergence rates for our methods under the most general conditions. Assuming that the upper-level component of our objective function is subhomogeneous, we develop efficient modification of the basic Fully Composite first-order and second-order Methods, and propose their accelerated variants.

math.OC↗

Dynamic pricing under nested logit demand

Recently, there is growing interest and need for dynamic pricing algorithms, especially, in the field of online marketplaces by offering smart pricing options for big online stores. We present an approach to adjust prices based on the observed online market data. The key idea is to characterize optimal prices as minimizers of a total expected revenue function, which turns out to be convex. We assume that consumers face information processing costs, hence, follow a discrete choice demand model, and suppliers are equipped with quantity adjustment costs. We prove the strong smoothness of the total expected revenue function by deriving the strong convexity modulus of its dual. Our gradient-based pricing schemes outbalance supply and demand at the convergence rates of $\mathcal{O}(\frac{1}{t})$ and $\mathcal{O}(\frac{1}{t^2})$, respectively. This suggests that the imperfect behavior of consumers and suppliers helps to stabilize the market.

math.OC↗

Inexact Tensor Methods with Dynamic Accuracies

In this paper, we study inexact high-order Tensor Methods for solving convex optimization problems with composite objective. At every step of such methods, we use approximate solution of the auxiliary problem, defined by the bound for the residual in function value. We propose two dynamic strategies for choosing the inner accuracy: the first one is decreasing as $1/k^{p + 1}$, where $p \geq 1$ is the order of the method and $k$ is the iteration counter, and the second approach is using for the inner accuracy the last progress in the target objective. We show that inexact Tensor Methods with these strategies achieve the same global convergence rate as in the error-free case. For the second approach we also establish local superlinear rates (for $p \geq 2$), and propose the accelerated scheme. Lastly, we present computational results on a variety of machine learning problems for several methods and different accuracy policies.

math.OC↗

Convex optimization based on global lower second-order models

In this paper, we present new second-order algorithms for composite convex optimization, called Contracting-domain Newton methods. These algorithms are affine-invariant and based on global second-order lower approximation for the smooth component of the objective. Our approach has an interpretation both as a second-order generalization of the conditional gradient method, or as a variant of trust-region scheme. Under the assumption, that the problem domain is bounded, we prove $\mathcal{O}(1/k^{2})$ global rate of convergence in functional residual, where $k$ is the iteration counter, minimizing convex functions with Lipschitz continuous Hessian. This significantly improves the previously known bound $\mathcal{O}(1/k)$ for this type of algorithms. Additionally, we propose a stochastic extension of our method, and present computational results for solving empirical risk minimization problem.

math.OC↗