Searcharxiv⌕ Search

arXiv subjects

Ilya Kuruzov

Publications and source records attributed to Ilya Kuruzov.

12 recordsLinked to original sources

On Some Versions of Subspace Optimization Methods with Inexact Gradient Information

It is well-known that accelerated gradient methods possess optimal complexity estimates for the class of convex smooth minimization problems. In many practical situations, it makes sense to work with inexact gradients. However, this can lead to the accumulation of corresponding inexactness in the theoretical estimates of the rate of convergence. We propose some modifications of first-order methods for convex optimization with an inexact gradient based on subspace optimization, such as Nemirovski's Conjugate Gradient method and the Sequential Subspace Optimization method. We study their convergence under different conditions on the inexactness both in the gradient value and in the accuracy of the solution of the subspace optimization subproblems. Besides this, we investigate a generalization of these results to the class of quasar-convex (weakly-quasi-convex) functions.

math.OC↗

A Line-search-free Method for Adaptive Decentralized Optimization

We study decentralized optimization over networks where agents cooperatively minimize a smooth (strongly) convex sum of local losses while communicating only with immediate neighbors. Prevailing decentralized methods require either centralized knowledge of global problem and network parameters for stepsize tuning--hence impractical, or costly per-iteration line-searches that demand access to local function values. We propose line-search-free, fully decentralized algorithms in which each agent adapts its stepsize using only past local iterates and gradients--with no extra function evaluations and no global tuning. The key technical ingredient is a new Lyapunov function, from which a natural adaptive stepsize rule emerges: at each iteration, each agent selects the largest stepsize that guarantees descent, based solely on a local curvature estimate built from successive gradients. The proposed algorithms enjoy strong theoretical guarantees: sublinear convergence rates for merely convex objectives and linear rates under strong convexity. Numerical experiments on standard benchmarks show consistent improvements over the state of the art, both adaptive and non-adaptive.

math.OC↗

Robust Decentralized Optimization under Node Failures via Adaptive Regularization

We study decentralized minimization of a sum of functions over a network where nodes may only leave, the remaining nodes stay connected, and the topology freezes between departures. Standard methods forget departed functions, causing a permanent bias proportional to the heterogeneity of the data. We propose Legacy Gradient Tracking (Legacy-GT): before leaving, a node compresses its function into a gradient-anchored quadratic legacy, bequeaths it to a neighbor, and provides a correction that exactly preserves the gradient-tracking invariant. We prove that the optimal legacy curvature is the average of the strong-convexity and smoothness constants, that legacies compose losslessly across chained departures, and that an adaptive anchor rule yields error bounds that decay geometrically after the last departure to a small residual--the minimum of the network's optimization error at departure and its heterogeneity radius. In contrast, the classic drop-and-forget baseline suffers a bias that never decays. Numerical experiments confirm the theory.

math.OC↗

Adaptive Decentralized Composite Optimization via Three-Operator Splitting

The paper studies decentralized optimization over networks, where agents minimize a sum of {\it locally} smooth (strongly) convex losses and plus a nonsmooth convex extended value term. We propose decentralized methods wherein agents {\it adaptively} adjust their stepsize via local backtracking procedures coupled with lightweight min-consensus protocols. Our design stems from a three-operator splitting factorization applied to an equivalent reformulation of the problem. The reformulation is endowed with a new BCV preconditioning metric (Bertsekas-O'Connor-Vandenberghe), which enables efficient decentralized implementation and local stepsize adjustments. We establish robust convergence guarantees. Under mere convexity, the proposed methods converge with a sublinear rate. Under strong convexity of the sum-function, and assuming the nonsmooth component is partly smooth, we further prove linear convergence. Numerical experiments corroborate the theory and highlight the effectiveness of the proposed adaptive stepsize strategy.

math.OC↗

A Parameter-free Decentralized Algorithm for Composite Convex Optimization

The paper studies decentralized optimization over networks, where agents minimize a composite objective consisting of the sum of smooth convex functions--the agents' losses--and an additional nonsmooth convex extended value function. We propose a decentralized algorithm wherein agents ${\it adaptively}$ adjust their stepsize using local backtracking procedures that require ${\it no}$ ${\it global}$ (network) information or extensive inter-agent communications. Our adaptive decentralized method enjoys robust convergence guarantees, outperforming existing decentralized methods, which are not adaptive. Our design is centered on a three-operator splitting, applied to a reformulation of the optimization problem. This reformulation utilizes a proposed BCV metric, which facilitates decentralized implementation and local stepsize adjustments while guarantying convergence.

math.OC↗

Adaptive Stepsize Selection in Decentralized Convex Optimization

We study decentralized optimization where multiple agents minimize the average of their (strongly) convex, smooth losses over a communication graph. Convergence of the existing decentralized methods generally hinges on an apriori, proper selection of the stepsize. Choosing this value is notoriously delicate: (i) it demands global knowledge from all the agents of the graph's connectivity and every local smoothness/strong-convexity constants--information they rarely have; (ii) even with perfect information, the worst-case tuning forces an overly small stepsize, slowing convergence in practice; and (iii) large-scale trial-and-error tuning is prohibitive. This work introduces a decentralized algorithm that is fully adaptive in the choice of the agents' stepsizes, without any global information and using only neighbor-to-neighbor communications--agents need not even know whether the problem is strongly convex. The algorithm retains strong guarantees: it converges at \emph{linear} rate when the losses are strongly convex and at \emph{sublinear} rate otherwise, matching the best-known rates of (nonadaptive) parameter-dependent methods.

math.OC↗

Mixed Newton Method for Optimization in Complex Spaces

In this paper, we modify and apply the recently introduced Mixed Newton Method, which is originally designed for minimizing real-valued functions of complex variables, to the minimization of real-valued functions of real variables by extending the functions to complex space. We show that arbitrary regularizations preserve the favorable local convergence properties of the method, and construct a special type of regularization used to prevent convergence to complex minima. We compare several variants of the method applied to training neural networks with real and complex parameters.

math.OC↗

Gradient-Type Methods For Decentralized Optimization Problems With Polyak-Łojasiewicz Condition Over Time-Varying Networks

This paper focuses on the decentralized optimization (minimization and saddle point) problems with objective functions that satisfy Polyak-Łojasiewicz condition (PL-condition). The first part of the paper is devoted to the minimization problem of the sum-type cost functions. In order to solve a such class of problems, we propose a gradient descent type method with a consensus projection procedure and the inexact gradient of the objectives. Next, in the second part, we study the saddle-point problem (SPP) with a structure of the sum, with objectives satisfying the two-sided PL-condition. To solve such SPP, we propose a generalization of the Multi-step Gradient Descent Ascent method with a consensus procedure, and inexact gradients of the objective function with respect to both variables. Finally, we present some of the numerical experiments, to show the efficiency of the proposed algorithm for the robust least squares problem.

math.OC↗

The Mirror-Prox Sliding Method for Non-smooth decentralized saddle-point problems

The saddle-point optimization problems have a lot of practical applications. This paper focuses on such non-smooth problems in decentralized case. This work contains generalization of recently proposed sliding for centralized problem. Through specific penalization method and this sliding we obtain algorithm for non-smooth decentralized saddle-point problems. Note, the proposed method approaches lower bounds both for number of communication rounds and calls of (sub-)gradient per node.

math.OC↗

Accelerated Methods for $α$-Weakly-Quasi-Convex Problems

We provide a quick overview of the class of $α$-weakly-quasi-convex problems and its relationships with other problem classes. We show that the previously known Sequential Subspace Optimization method retains its optimal convergence rate when applied to minimization problems with smooth $α$-weakly-quasi-convex objectives. We also show that Nemirovski's conjugate gradients method of strongly convex minimization achieves its optimal convergence rate under weaker conditions of $α$-weak-quasi-convexity and quad\-ratic growth. Previously known results only capture the special case of 1-weak-quasi-convexity or give convergence rates with worse dependence on the parameter $α$.

math.OC↗

Solving strongly convex-concave composite saddle point problems with a small dimension of one of the variables

The article is devoted to the development of algorithmic methods ensuring efficient complexity bounds for strongly convex-concave saddle point problems in the case when one of the groups of variables is high-dimensional, and the other is relatively low-dimensional (up to a hundred). The proposed technique is based on reducing problems of this type to a problem of minimizing a convex (maximizing a concave) functional in one of the variables, for which it is possible to find an approximate gradient at an arbitrary point with the required accuracy using an auxiliary optimization subproblem with another variable. In this case, the ellipsoid method is used for low-dimensional problems (if necessary, with an inexact $δ$-subgradient), and accelerated gradient methods are used for high-dimensional problems. For the case of a very small dimension of one of the groups of variables (up to 5), an approach based on a new version of the multidimensional analog of the Yu. E. Nesterov's method on the square (multidimensional dichotomy) is proposed with the possibility of using inexact values of the gradient of the objective functional.

math.OC↗

Sequential Subspace Optimization for Quasar-Convex Optimization Problems with Inexact Gradient

It is well-known that accelerated gradient first-order methods possess optimal complexity estimates for the class of convex smooth minimization problems. In many practical situations it makes sense to work with inexact gradient information. However, this can lead to an accumulation of corresponding inexactness in the theoretical estimates of the rate of convergence. We propose some modification of the Sequential Subspace Optimization Method (SESOP) for minimization problems with quasar-convex functions with inexact gradient. A theoretical result is obtained indicating the absence of accumulation of gradient inexactness. A numerical implementation of the proposed version of the SESOP method and its comparison with the known Similar Triangle Method with an inexact gradient is carried out.

math.OC↗