SearcharxivSearch

arXiv subjects

Fedor Stonyakin

Publications and source records attributed to Fedor Stonyakin.

At least 19 recordsLinked to original sources

An Inexact Augmented Lagrangian Method for $(L_0, L_1)$-Smooth Convex Optimization

Augmented Lagrangian methods are among the most effective approaches for solving constrained convex optimization problems. However, classical complexity analyses of first-order methods applied within the augmented Lagrangian framework usually rely on the assumption that the objective function has a Lipschitz continuous gradient. This assumption excludes an important class of generalized smooth functions whose gradients may grow unboundedly. In this paper, we study an inexact augmented Lagrangian method for solving linearly constrained convex optimization problems with $(L_0,L_1)$-smooth objective functions. We show that the augmented Lagrangian subproblems preserve the $(L_0,L_1)$-smooth structure, with parameters depending on the penalty coefficient. This property allows us to employ recent accelerated first-order schemes designed for generalized smooth optimization instead of classical smooth optimization methods. In particular, we combine the inexact augmented Lagrangian framework with a two-stage acceleration procedure based on clipped gradient descent and accelerated optimization.

math.OC

On Some Versions of Subspace Optimization Methods with Inexact Gradient Information

It is well-known that accelerated gradient methods possess optimal complexity estimates for the class of convex smooth minimization problems. In many practical situations, it makes sense to work with inexact gradients. However, this can lead to the accumulation of corresponding inexactness in the theoretical estimates of the rate of convergence. We propose some modifications of first-order methods for convex optimization with an inexact gradient based on subspace optimization, such as Nemirovski's Conjugate Gradient method and the Sequential Subspace Optimization method. We study their convergence under different conditions on the inexactness both in the gradient value and in the accuracy of the solution of the subspace optimization subproblems. Besides this, we investigate a generalization of these results to the class of quasar-convex (weakly-quasi-convex) functions.

math.OC

Normalized First-Order Methods for Convex (L0, L1)-Smooth Optimization with Inexact Gradients

Generalized smoothness, such as (L0, L1)-smoothness, have recently attracted considerable attention due to their ability to model optimization problems arising in modern machine and deep learning, where the classical Lipschitz assumptions of the gradient is often violated. At the same time, computing exact gradients may be impractical or computationally expensive in many applications. In this work, we study convex (L0, L1)-smooth optimization (for normalized gradient method we consider quasi-convex problems too) under access only to a normalized approximation recently proposed Comparison Oracle, which returns an inexact normalized gradient in linear time with a bounded absolute error. Within this framework, we develop comparison-oracle variants of Normalized Gradient Descent and Gradient Descent with Polyak stepsizes. We establish explicit upper bounds on the approximation error that guarantee convergence and derive convergence rates for all proposed methods. Unlike existing analyses, our results require neither classical smoothness assumptions nor access to exact gradients or their exact normalized counterparts. Finally, numerical experiments corroborate the theoretical findings.

math.OC

Adaptive Variant of Frank-Wolfe Method for Relative Smooth Convex Optimization Problems

The paper introduces a new adaptive version of the Frank-Wolfe algorithm for relatively smooth convex functions. It is proposed to use the Bregman divergence other than half the square of the Euclidean norm in the formula for step-size. Algorithm convergence estimates for minimization problems of relatively smooth convex functions with the triangle scaling property are proved. Computational experiments are performed, and conditions are shown in which the obvious advantage of the proposed algorithm over its Euclidean norm analogue is shown. We also found examples of problems for which the proposed variation of the Frank-Wolfe method works better than known accelerated gradient-type methods for relatively smooth convex functions with the triangle scaling property.

math.OC

Optimal Convergence Rate for Mirror Descent Methods with special Time-Varying Step Sizes Rules

In this paper, the optimal convergence rate $O\left(N^{-1/2}\right)$ (where $N$ is the total number of iterations performed by the algorithm), without the presence of a logarithmic factor, is proved for mirror descent algorithms with special time-varying step sizes, for solving classical constrained non-smooth problems, problems with the composite model and problems with non-smooth functional (inequality types) constraints. The proven result is an improvement on the well-known rate $O\left(\log (N) N^{-1/2}\right)$ for the mirror descent algorithms with the time-varying step sizes under consideration. It was studied a new weighting scheme assigns smaller weights to the initial points and larger weights to the most recent points. This scheme improves the convergence rate of the considered mirror descent methods, which in the conducted numerical experiments outperform the other methods providing a better solution in all the considered test problems.

math.OC

About some works of Boris Polyak on convergence of gradient methods and their development

The paper presents a review of the state-of-the-art of subgradient and accelerated methods of convex optimization, including in the presence of disturbances and access to various information about the objective function (function value, gradient, stochastic gradient, higher derivatives). For nonconvex problems, the Polak-Lojasiewicz condition is considered and a review of the main results is given. The behavior of numerical methods in the presence of sharp minima is considered. The purpose of this survey is to show the influence of the works of B.T. Polyak (1935 -- 2023) on gradient optimization methods and their neighborhoods on the modern development of numerical optimization methods.

math.OC

Higher Degree Inexact Model for Optimization problems

In this paper, it was proposed a new concept of the inexact higher degree $(δ, L, q)$-model of a function that is a generalization of the inexact $(δ, L)$-model, $(δ, L)$-oracle and $(δ, L)$-oracle of degree $q \in [0,2)$. Some examples were provided to illustrate the proposed new model. Adaptive inexact gradient and fast gradient methods for convex and strongly convex functions were constructed and analyzed using the new proposed inexact model. A universal fast gradient method that allows solving optimization problems with a weaker level of smoothness, among them non-smooth problems was proposed. For convex optimization problems it was proved that the proposed gradient and fast gradient methods could be converged with rates $O\left(\frac{1}{k} + \fracδ{k^{q/2}}\right)$ and $O\left(\frac{1}{k^2} + \fracδ{k^{(3q-2)/2}}\right)$, respectively. For the gradient method, the coefficient of $δ$ diminishes with $k$, and for the fast gradient method, there is no error accumulation for $q \geq 2/3$. It proposed a definition of an inexact higher degree oracle for strongly convex functions and a projected gradient method using this inexact oracle. For variational inequalities and saddle point problems, a higher degree inexact model and an adaptive method called Generalized Mirror Prox to solve such class of problems using the proposed inexact model were proposed. Some numerical experiments were conducted to demonstrate the effectiveness of the proposed inexact model, we test the universal fast gradient method to solve some non-smooth problems with a geometrical nature.

math.OC

Universal methods for variational inequalities: deterministic and stochastic cases

In this paper, we propose universal proximal mirror methods to solve the variational inequality problem with Holder continuous operators in both deterministic and stochastic settings. The proposed methods automatically adapt not only to the oracle's noise (in the stochastic setting of the problem) but also to the Holder continuity of the operator without having prior knowledge of either the problem class or the nature of the operator information. We analyzed the proposed algorithms in both deterministic and stochastic settings and obtained estimates for the required number of iterations to achieve a given quality of a solution to the variational inequality. We showed that, without knowing the Holder exponent and Holder constant of the operators, the proposed algorithms have the least possible in the worst case sense complexity for the considered class of variational inequalities. We also compared the resulting stochastic algorithm with other popular optimizers for the task of image classification.

math.OC

Gradient-Type Methods For Decentralized Optimization Problems With Polyak-Łojasiewicz Condition Over Time-Varying Networks

This paper focuses on the decentralized optimization (minimization and saddle point) problems with objective functions that satisfy Polyak-Łojasiewicz condition (PL-condition). The first part of the paper is devoted to the minimization problem of the sum-type cost functions. In order to solve a such class of problems, we propose a gradient descent type method with a consensus projection procedure and the inexact gradient of the objectives. Next, in the second part, we study the saddle-point problem (SPP) with a structure of the sum, with objectives satisfying the two-sided PL-condition. To solve such SPP, we propose a generalization of the Multi-step Gradient Descent Ascent method with a consensus procedure, and inexact gradients of the objective function with respect to both variables. Finally, we present some of the numerical experiments, to show the efficiency of the proposed algorithm for the robust least squares problem.

math.OC

Adaptive Algorithms for Relatively Lipschitz Continuous Convex Optimization Problems

Recently there were proposed some innovative convex optimization concepts, namely, relative smoothness [1] and relative strong convexity [2,3]. These approaches have significantly expanded the class of applicability of gradient-type methods with optimal estimates of the convergence rate, which are invariant regardless of the dimensionality of the problem. Later Yu. Nesterov and H. Lu introduced some modifications of the Mirror Descent method for convex minimization problems with the corresponding analogue of the Lipschitz condition (so-called relative Lipschitz continuity). By introducing an artificial inaccuracy to the optimization model, we propose adaptive methods for minimizing a convex Lipschitz continuous function, as well as for the corresponding class of variational inequalities. We also consider an adaptive "universal" method, applicable to convex minimization problems both on the class of relatively smooth and relatively Lipschitz continuous functionals with optimal estimates of the convergence rate. The universality of the method makes it possible to justify the applicability of the obtained theoretical results to a wider class of convex optimization problems. We also present the results of numerical experiments.

math.OC

Highly Smoothness Zero-Order Methods for Solving Optimization Problems under PL Condition

In this paper, we study the black box optimization problem under the Polyak--Lojasiewicz (PL) condition, assuming that the objective function is not just smooth, but has higher smoothness. By using "kernel-based" approximation instead of the exact gradient in Stochastic Gradient Descent method, we improve the best known results of convergence in the class of gradient-free algorithms solving problem under PL condition. We generalize our results to the case where a zero-order oracle returns a function value at a point with some adversarial noise. We verify our theoretical results on the example of solving a system of nonlinear equations.

math.OC

Algorithms for solving variational inequalities and saddle point problems with some generalizations of Lipschitz property for operators

The article is devoted to the development of numerical methods for solving saddle point problems and variational inequalities with simplified requirements for the smoothness conditions of functionals. Recently there were proposed some notable methods for optimization problems with strongly monotone operators. Our focus here is on newly proposed techniques for solving strongly convex-concave saddle point problems. One of the goals of the article is to improve the obtained estimates of the complexity of introduced algorithms by using accelerated methods for solving auxiliary problems. The second focus of the article is introducing an analogue of the boundedness condition for the operator in the case of arbitrary (not necessarily Euclidean) prox structure. We propose an analogue of the mirror descent method for solving variational inequalities with such operators, which is optimal in the considered class of problems.

math.OC

Intermediate Gradient Methods with Relative Inexactness

This paper is devoted to first-order algorithms for smooth convex optimization with inexact gradients. Unlike the majority of the literature on this topic, we consider the setting of relative rather than absolute inexactness. More precisely, we assume that an additive error in the gradient is proportional to the gradient norm, rather than being globally bounded by some small quantity. We propose a novel analysis of the accelerated gradient method under relative inexactness and strong convexity and improve the bound on the maximum admissible error that preserves the linear convergence of the algorithm. In other words, we analyze how robust is the accelerated gradient method to the relative inexactness of the gradient information. Moreover, based on the Performance Estimation Problem (PEP) technique, we show that the obtained result is optimal for the family of accelerated algorithms we consider. Motivated by the existing intermediate methods with absolute error, i.e., the methods with convergence rates that interpolate between slower but more robust non-accelerated algorithms and faster, but less robust accelerated algorithms, we propose an adaptive variant of the intermediate gradient method with relative error in the gradient.

math.OC

Online Optimization Problems with Functional Constraints under Relative Lipschitz Continuity and Relative Strong Convexity Conditions

Recently, there were introduced important classes of relatively smooth, relatively continuous, and relatively strongly convex optimization problems. These concepts have significantly expanded the class of problems for which optimal complexity estimates of gradient-type methods in high-dimensional spaces take place. Basing on some recent works devoted to online optimization (regret minimization) problems with both relatively Lipschitz continuous and relatively strongly convex objective function, we introduce algorithms for solving the strongly convex optimization problem with inequality constraints in the online setting. We propose a scheme with switching between productive and nonproductive steps for such types of problems and prove its convergence rate for the class of relatively Lipschitz and strongly convex minimization problems. We also provide an experimental comparison between the proposed method and AdaMirr, recently proposed for relatively Lipschitz convex problems.

math.OC

Solving strongly convex-concave composite saddle point problems with a small dimension of one of the variables

The article is devoted to the development of algorithmic methods ensuring efficient complexity bounds for strongly convex-concave saddle point problems in the case when one of the groups of variables is high-dimensional, and the other is relatively low-dimensional (up to a hundred). The proposed technique is based on reducing problems of this type to a problem of minimizing a convex (maximizing a concave) functional in one of the variables, for which it is possible to find an approximate gradient at an arbitrary point with the required accuracy using an auxiliary optimization subproblem with another variable. In this case, the ellipsoid method is used for low-dimensional problems (if necessary, with an inexact $δ$-subgradient), and accelerated gradient methods are used for high-dimensional problems. For the case of a very small dimension of one of the groups of variables (up to 5), an approach based on a new version of the multidimensional analog of the Yu. E. Nesterov's method on the square (multidimensional dichotomy) is proposed with the possibility of using inexact values of the gradient of the objective functional.

math.OC

Generalized Mirror Prox for Monotone Variational Inequalities: Universality and Inexact Oracle

We introduce an inexact oracle model for variational inequalities (VI) with monotone operator, propose a numerical method which solves such VI's and analyze its convergence rate. As a particular case, we consider VI's with Hölder-continuous operator and show that our algorithm is universal. This means that without knowing the Hölder parameter $ν$ and Hölder constant $L_ν$ it has the best possible complexity for this class of VI's, namely our algorithm has complexity $O\left( \inf_{ν\in[0,1]}\left(\frac{L_ν}{\varepsilon} \right)^{\frac{2}{1+ν}}R^2 \right)$, where $R$ is the size of the feasible set and $\varepsilon$ is the desired accuracy of the solution. We also consider the case of VI's with strongly monotone operator and generalize our method for VI's with inexact oracle and our universal method for this class of problems. Finally, we show, how our method can be applied to convex-concave saddle point problems with Hölder-continuous partial subgradients.

math.OC

Mirror Descent and Constrained Online Optimization Problems

We consider the following class of online optimization problems with functional constraints. Assume, that a finite set of convex Lipschitz-continuous non-smooth functionals are given on a closed set of $n$-dimensional vector space. The problem is to minimize the arithmetic mean of functionals with a convex Lipschitz-continuous non-smooth constraint. In addition, it is allowed to calculate the (sub)gradient of each functional only once. Using some recently proposed adaptive methods of Mirror Descent the method is suggested to solve the mentioned constrained online optimization problem with optimal estimate of accuracy. For the corresponding non-Euclidean prox-structure the case of a set of $n$-dimensional vectors lying on the standard $n$-dimensional simplex is considered.

math.OC

Mirror Descent for Constrained Optimization Problems with Large Subgradient Values

Based on the ideas of arXiv:1710.06612, we consider the problem of minimization of the Holder-continuous non-smooth functional $f$ with non-positive convex (generally, non-smooth) Lipschitz-continuous functional constraint. We propose some novel strategies of step-sizes and adaptive stopping rules in Mirror Descent algorithms for the considered class of problems. It is shown that the methods are applicable to the objective functionals of various levels of smoothness. Applying the restart technique to the Mirror Descent Algorithm there was proposed an optimal method to solve optimization problems with strongly convex objective functionals. Estimates of the rate of convergence of the considered algorithms are obtained depending on the level of smoothness of the objective functional. These estimates indicate the optimality of considered methods from the point of view of the theory of lower oracle bounds. In addition, the case of a quasi-convex objective functional and constraint was considered.

math.OC