SearcharxivSearch

arXiv subjects

Yurii Nesterov

Publications and source records attributed to Yurii Nesterov.

At least 19 recordsLinked to original sources

Multiconic Optimization for Symmetric Cones and Hyperbolic Coupling

We develop a new interior-point algorithm for solving multiconic optimization problems using the parabolic target space approach. The feasible cone in these problems is composed as a direct product of many small-dimensional cones. Our approach is based on a new concept, called the hyperbolic coupling. This provides a new framework that has an advantage of interdependent pairs of primal-dual variables. In this way, their behaviour is much more controllable. We justify all main steps in the complexity analysis of the algorithm and prove that the overall complexity of solving this type of large-scale nonlinear problems by our algorithm is comparable with the best known complexity for solving linear programming problems of the same dimension.

math.OC

Theorem of Alternative for Extended Homogeneous Linear System and its Application in Conic Optimization

In this paper, we develop a new framework for constructing infeasible-start primal-dual methods for Conic Optimization. Our approach can be seen as a straightforward consequence of Gordan Theorem of Alternative. Given by the target upper bound $\epsilon > 0$ for the duality gap as the only input parameter, we form an auxiliary convex problem of minimizing barrier function with linear equality constraints. Its solution can be easily transformed to the requested output. This function can be minimized by different schemes of Unconstrained Optimization, with possible quadratic convergence in the end of the process. In our paper, we analyze the Damped Newton Method and a short-step path-following scheme. For both of them, we prove polynomial-time complexity results. Our methods are able to benefit from the hot-start opportunities. We can ensure the residual of the linear equality constraints in the primal and dual problems to be at the level of machine accuracy, independently on the accuracy parameter $\epsilon$.

math.OC

Universal Reduced-Operator Method and High-Order Global Curvature Bounds

In this paper, we develop a new concept of Global Curvature Bound (GCB) for an arbitrary nonlinear operator between abstract metric spaces. We use this notion to characterize the global complexity of high-order algorithms solving composite variational problems, which include convex minimization and min-max problems. We develop the new universal Reduced-Operator Method, which automatically achieves the fastest universal rate within our class, while our analysis does not need any specific assumptions about smoothness of the target nonlinear operator. Every step of our universal method of order $p \geq 1$ requires access to the $p$-th order derivative of the operator and the solution of a strictly monotone doubly regularized subproblem. For $p = 1$, this corresponds to computing the standard Jacobian matrix of the operator and solving a simple monotone subproblem, which can be handled using different methods of Convex Optimization. All our results are consequences of the new theorem on the quadratic growth of GCB for general nonlinear operators.

math.OC

Universal Complexity Bounds for Universal Gradient Methods in Nonlinear Optimization

In this paper, we provide the universal first-order methods of Composite Optimization with new complexity analysis. It delivers some universal convergence guarantees, which are not linked directly to any parametric problem class. However, they can be easily transformed into the rates of convergence for the particular problem classes by substituting the corresponding upper estimates for the Global Curvature Bound of the objective function. We analyze in this way the simple gradient method for nonconvex minimization, gradient methods for convex composite optimization, and their accelerated variant. For them, the only input parameter is the required accuracy of the approximate solution. The accelerated variant of our scheme automatically ensures the best possible rate of convergence simultaneously for all parametric problem classes containing the smooth part of the objective function.

math.OC

Interior-Point Algorithms for Monotone Linear Complementarity Problem Based on Different Predictor Directions

In this paper, we introduce two parabolic target-space interior-point algorithms for solving monotone linear complementarity problems. The first algorithm is based on a universal tangent direction, which has been recently proposed for linear optimization problems. We prove that this method has the best known worst-case complexity bound. We extend onto LCP its auto-correcting version, and prove its local quadratic convergence under a non-degeneracy assumption. In our numerical experiments, we compare the new algorithms with a general method, recently developed for weighted monotone linear complementarity problems.

math.OC

Asymmetric Long-Step Primal-Dual Interior-Point Methods with Dual Centering

In this paper, we develop a new asymmetric framework for solving primal-dual problems of Conic Optimization by Interior-Point Methods (IPMs). It allows development of efficient methods for problems, where the dual formulation is simpler than the primal one. The problems of this type arise, in particular, in Semidefinite Optimization (SDO), for which we propose a new method with very attractive computational cost. Our long-step predictor-corrector scheme is based on centering in the dual space. It computes the affine-scaling predicting direction by the use of the dual barrier function, controlling the tangent step size by a functional proximity measure. We show that for symmetric cones, the search procedure at the predictor step is very cheap. In general, we do not need sophisticated Linear Algebra, restricting ourselves only by Cholesky factorization. However, our complexity bounds correspond to the best known polynomial-time results. Moreover, for symmetric cones the bounds automatically depend on the minimal barrier parameter between the primal or the dual feasible sets. We show by SDO-examples that the corresponding gain can be very big. We argue that the dual framework is more suitable for adjustment to the actual complexity of the problem. As an example, we discuss some classes of SDO-problems, where the number of iterations is proportional to the square root of the number of linear equality constraints. Moreover, the computational cost of one iteration there is similar to that one for Linear Optimization. We support our theoretical developments by preliminary but encouraging numerical results with randomly generated SDO-problems of different size.

math.OC

Local and Global Convergence of Greedy Parabolic Target-Following Methods for Linear Programming

In the first part of this paper, we prove that, under some natural non-degeneracy assumptions, the Greedy Parabolic Target-Following Method, based on {\em universal tangent direction} has a favorable local behavior. In view of its global complexity bound of the order $O(\sqrt{n} \ln {1 \over \epsilon})$, this fact proves that the functional proximity measure, used for controlling the closeness to Greedy Central Path, is large enough for ensuring a local super-linear rate of convergence, provided that the proximity to the path is gradually reduced. This requirement is eliminated in our second algorithm based on a new auto-correcting predictor direction. This method, besides the best-known polynomial-time complexity bound, ensures an automatic switching onto the local quadratic convergence in a small neighborhood of solution. Our third algorithm approximates the path by quadratic curves. On the top of the best-known global complexity bound, this method benefits from an unusual local cubic rate of convergence. This amelioration needs no serious increase in the cost of one iteration. We compare the advantages of these local accelerations with possibilities of finite termination. The conditions allowing the optimal basis detection sometimes are even weaker than those required for the local superlinear convergence. Hence, it is important to endow the practical optimization schemes with both abilities. The proposed methods have a very interesting combination of favorable properties, which can be hardly found in the most of existing Interior-Point schemes. As all other parabolic target-following schemes, the new methods can start from an arbitrary strictly feasible primal-dual pair and go directly towards the optimal solution of the problem in a single phase. The preliminary computational experiments confirm the advantage of the second-order prediction.

math.OC

Improved global performance guarantees of second-order methods in convex minimization

In this paper, we attempt to compare two distinct branches of research on second-order optimization methods. The first one studies self-concordant functions and barriers, the main assumption being that the third derivative of the objective is bounded by the second derivative. The second branch studies cubic regularized Newton methods (CRNMs) with the main assumption that the second derivative is Lipschitz continuous. We develop a new theoretical analysis for a path-following scheme (PFS) for general self-concordant functions, as opposed to the classical path-following scheme developed for self-concordant barriers. We show that the complexity bound for this scheme is better than that of the Damped Newton Method (DNM) and show that our method has global superlinear convergence. We propose also a new predictor-corrector path-following scheme (PCPFS) that leads to further improvement of constant factors in the complexity guarantees for minimizing general self-concordant functions. We also apply path-following schemes to different classes of constrained optimization problems and obtain the resulting complexity bounds. Finally, we analyze an important subclass of general self-concordant functions, namely a class of strongly convex functions with Lipschitz continuous second derivative, and show that for this subclass CRNMs give even better complexity bounds.

math.OC

An optimal lower bound for smooth convex functions

First order methods endowed with global convergence guarantees operate using global lower bounds on the objective. The tightening of the bounds has been shown to increase both the theoretical guarantees and the practical performance. In this work, we define a global lower bound for smooth differentiable objectives that is optimal with respect to the collected oracle information. The bound can be readily employed by the Gradient Method with Memory to improve its performance. Further using the machinery underlying the optimal bounds, we introduce a modified version of the estimate sequence that we use to construct an Optimized Gradient Method with Memory possessing the best known convergence guarantees for its class of algorithms, even in terms of the proportionality constant. We additionally equip the method with an adaptive convergence guarantee adjustment procedure that is an effective replacement for line-search. Simulation results on synthetic but otherwise difficult smooth problems validate the theoretical properties of the bound and proposed methods.

math.OC

High-Order Reduced-Gradient Methods for Composite Variational Inequalities

This paper can be seen as an attempt of rethinking the {\em Extra-Gradient Philosophy} for solving Variational Inequality Problems. We show that the properly defined {\em Reduced Gradients} can be used instead for finding approximate solutions to Composite Variational Inequalities by the higher-order schemes. Our methods are optimal since their performance is proportional to the lower worst-case complexity bounds for corresponding problem classes. They enjoy the provable hot-start capabilities even being applied to minimization problems. The primal version of our schemes demonstrates a linear rate of convergence under an appropriate uniform monotonicity assumption.

math.OC

Primal subgradient methods with predefined stepsizes

In this paper, we suggest a new framework for analyzing primal subgradient methods for nonsmooth convex optimization problems. We show that the classical step-size rules, based on normalization of subgradient, or on the knowledge of optimal value of the objective function, need corrections when they are applied to optimization problems with constraints. Their proper modifications allow a significant acceleration of these schemes when the objective function has favorable properties (smoothness, strong convexity). We show how the new methods can be used for solving optimization problems with functional constraints with possibility to approximate the optimal Lagrange multipliers. One of our primal-dual methods works also for unbounded feasible set.

math.OC

Convex quartic problems: homogenized gradient method and preconditioning

We consider a convex minimization problem for which the objective is the sum of a homogeneous polynomial of degree four and a linear term. Such task arises as a subproblem in algorithms for quadratic inverse problems with a difference-of-convex structure. We design a first-order method called Homogenized Gradient, along with an accelerated version, which enjoy fast convergence rates of respectively $\mathcal{O}(\kappa^2/K^2)$ and $\mathcal{O}(\kappa^2/K^4)$ in relative accuracy, where $K$ is the iteration counter. The constant $\kappa$ is the quartic condition number of the problem. Then, we show that for a certain class of problems, it is possible to compute a preconditioner for which this condition number is $\sqrt{n}$, where $n$ is the problem dimension. To establish this, we study the more general problem of finding the best quadratic approximation of an $\ell_p$ norm composed with a quadratic map. Our construction involves a generalization of the so-called Lewis weights.

math.OC

Gradient Methods for Stochastic Optimization in Relative Scale

We propose a new concept of a relatively inexact stochastic subgradient and present novel first-order methods that can use such objects to approximately solve convex optimization problems in relative scale. An important example where relatively inexact subgradients naturally arise is given by the Power or Lanczos algorithms for computing an approximate leading eigenvector of a symmetric positive semidefinite matrix. Using these algorithms as subroutines in our methods, we get new optimization schemes that can provably solve certain large-scale Semidefinite Programming problems with relative accuracy guarantees by using only matrix-vector products.

math.OC

Super-Universal Regularized Newton Method

We analyze the performance of a variant of Newton method with quadratic regularization for solving composite convex minimization problems. At each step of our method, we choose regularization parameter proportional to a certain power of the gradient norm at the current point. We introduce a family of problem classes characterized by H\"older continuity of either the second or third derivative. Then we present the method with a simple adaptive search procedure allowing an automatic adjustment to the problem class with the best global complexity bounds, without knowing specific parameters of the problem. In particular, for the class of functions with Lipschitz continuous third derivative, we get the global $O(1/k^3)$ rate, which was previously attributed to third-order tensor methods. When the objective function is uniformly convex, we justify an automatic acceleration of our scheme, resulting in a faster global rate and local superlinear convergence. The switching between the different rates (sublinear, linear, and superlinear) is automatic. Again, for that, no a priori knowledge of parameters is needed.

math.OC

Adaptive Third-Order Methods for Composite Convex Optimization

In this paper we propose third-order methods for composite convex optimization problems in which the smooth part is a three-times continuously differentiable function with Lipschitz continuous third-order derivatives. The methods are adaptive in the sense that they do not require the knowledge of the Lipschitz constant. Trial points are computed by the inexact minimization of models that consist in the nonsmooth part of the objective plus a quartic regularization of third-order Taylor polynomial of the smooth part. Specifically, approximate solutions of the auxiliary problems are obtained by using a Bregman gradient method as inner solver. Different from existing adaptive approaches for high-order methods, in our new schemes the regularization parameters are adjusted taking into account the progress of the inner solver. With this technique, we show that the basic method finds an $\epsilon$-approximate minimizer of the objective function performing at most $\mathcal{O}\left(|\log(\epsilon)|\epsilon^{-\frac{1}{3}}\right)$ iterations of the inner solver. An accelerated adaptive third-order method is also presented with total inner iteration complexity of $\mathcal{O}\left(|\log(\epsilon)|\epsilon^{-\frac{1}{4}}\right)$.

math.OC

Quartic Regularity

In this paper, we propose new linearly convergent second-order methods for minimizing convex quartic polynomials. This framework is applied for designing optimization schemes, which can solve general convex problems satisfying a new condition of quartic regularity. It assumes positive definiteness and boundedness of the fourth derivative of the objective function. For such problems, an appropriate quartic regularization of Damped Newton Method has global linear rate of convergence. We discuss several important consequences of this result. In particular, it can be used for constructing new second-order methods in the framework of high-order proximal-point schemes. These methods have convergence rate $\tilde O(k^{-p})$, where $k$ is the iteration counter, $p$ is equal to 3, 4, or 5, and tilde indicates the presence of logarithmic factors in the complexity bounds for the auxiliary problems, which are solved at each iteration of the schemes.

math.OC

Gradient Regularization of Newton Method with Bregman Distances

In this paper, we propose a first second-order scheme based on arbitrary non-Euclidean norms, incorporated by Bregman distances. They are introduced directly in the Newton iterate with regularization parameter proportional to the square root of the norm of the current gradient. For the basic scheme, as applied to the composite optimization problem, we establish the global convergence rate of the order $O(k^{-2})$ both in terms of the functional residual and in the norm of subgradients. Our main assumption on the smooth part of the objective is Lipschitz continuity of its Hessian. For uniformly convex functions of degree three, we justify global linear rate, and for strongly convex function we prove the local superlinear rate of convergence. Our approach can be seen as a relaxation of the Cubic Regularization of the Newton method, which preserves its convergence properties, while the auxiliary subproblem at each iteration is simpler. We equip our method with adaptive line search procedure for choosing the regularization parameter. We propose also an accelerated scheme with convergence rate $O(k^{-3})$, where $k$ is the iteration counter.

math.OC

High-order methods beyond the classical complexity bounds, II: inexact high-order proximal-point methods with segment search

A bi-level optimization framework (BiOPT) was proposed in [3] for convex composite optimization, which is a generalization of bi-level unconstrained minimization framework (BLUM) given in [20]. In this continuation paper, we introduce a $p$th-order proximal-point segment search operator which is used to develop novel accelerated methods. Our first algorithm combines the exact element of this operator with the estimating sequence technique to derive new iteration points, which are shown to attain the convergence rate $\mathcal{O}(k^{-(3p+1)/2})$, for the iteration counter $k$. We next consider inexact elements of the high-order proximal-point segment search operator to be employed in the BiOPT framework. We particularly apply the accelerated high-order proximal-point method at the upper level, while we find approximate solutions of the proximal-point segment search auxiliary problem by a combination of non-Euclidean composite gradient and bisection methods. For $q=\lfloor p/2 \rfloor$, this amounts to a $2q$th-order method with the convergence rate $\mathcal{O}(k^{-(6q+1)/2})$ (the same as the optimal bound of $2q$th-order methods) for even $p$ ($p=2q$) and the superfast convergence rate $\mathcal{O}(k^{-(3q+1)})$ for odd $p$ ($p=2q+1$).

math.OC