Searcharxiv⌕ Search

arXiv subjects

Necdet Serhat Aybat

Publications and source records attributed to Necdet Serhat Aybat.

At least 19 recordsLinked to original sources

Restarted Accelerated Primal-Dual Algorithms with Adaptive Stepsizes for Nonlinear Conic Constrained Convex Optimization

We propose restarted accelerated primal-dual algorithms with (non-monotone) backtracking (rAPDB) for convex nonlinear conic programs, with quadratically constrained quadratic programs (QCQPs) as a special case. Unlike linear and quadratic programs, these problems give rise to convex-concave minimax reformulations with non-bilinear coupling terms; therefore, the existing primal-dual methods for bilinear couplings are not applicable. To address this challenge, we build on the accelerated primal-dual method with adaptive stepsize search -- as it adapts to the local curvature -- and develop both fixed-frequency and adaptive restart schemes, incorporating both monotone and non-monotone adaptive step-size search strategies. The resulting algorithms require only first-order information and matrix-vector products, making them suitable for large-scale and GPU-accelerated implementation. Under metric subregularity of the KKT mapping, we prove a quadratic growth property for a self-centered smoothed duality gap and establish global linear convergence of the proposed restarted methods. We also establish sufficient conditions under which the metric subregularity holds even for general nonconvex problems over convex polyhedral cones. These results are new and may be of independent interest. Numerical experiments on random QCQPs and kernel matrix learning instances show that the proposed methods, especially with non-monotone adaptive stepsizes and GPU acceleration, achieve strong practical performance.

math.OC↗

A Stochastic GDA Method With Backtracking For Solving Nonconvex Concave Minimax Problems

We propose a stochastic GDA (gradient descent ascent) method with backtracking (SGDA-B) to solve nonconvex-concave (NCC) minimax problems of the form: $\min_{\mathbf{x}} \max_y \sum_{i=1}^N g_i(x_i)+f(\mathbf{x},y)-h(y)$, where $h$ and $g_i$ for $i=1,\cdots,N$ are closed, convex functions, and for some $L,μ\geq 0$, $f$ is $L$-smooth and $f(\mathbf{x},\cdot)$ is $μ$-strongly concave for all $\mathbf{x}$ in the problem domain. We consider the stochastic setting where one only has an access to an unbiased stochastic oracle of $\nabla f$ with a finite variance bound $σ^2$. While most of the existing methods assume knowledge of $L$, $μ$ and/or $σ^2$, SGDA-B is agnostic to all of these problem parameters. Moreover, SGDA-B can support random block-coordinate updates. In the deterministic setting, i.e., $σ^2=0$ and one can compute $\nabla f$ exactly, SGDA-B can compute an $ε$-stationary point within $\mathcal{O}(Lκ^2/ε^2)$ and $\mathcal{O}(L^3/ε^4)$ gradient calls when $μ>0$ and $μ=0$, respectively, where $κ\triangleq L/μ$. In the stochastic setting, i.e., $σ^2>0$, for any $p\in(0,1)$ and $ε>0$, it can compute an $ε$-stationary point with high probability, which requires $\mathcal{O}(L κ^3 ε^{-4} \log^2(1/p))$ and $\tilde{\mathcal{O}}(L^4ε^{-7}\log^2(1/p))$ stochastic oracle calls, with probability at least $1-p$, when $μ>0$ and $μ=0$, respectively. To our knowledge, SGDA-B is the first GDA-type method with backtracking to solve NCC minimax problems and achieves the best complexity among the methods that are agnostic to $L$, $μ$ and $σ^2$. We also provide numerical results for SGDA-B on a distributionally robust learning problem illustrating the potential performance gains that can be achieved by SGDA-B.

math.OC↗

An Accelerated Primal Dual Algorithm with Backtracking for Decentralized Constrained Optimization

We propose a distributed accelerated primal-dual method with backtracking (D-APDB) for cooperative multi-agent constrained consensus optimization problems over an undirected network of agents, where only those agents connected by an edge can directly communicate to exchange large-volume data vectors using a high-speed, short-range communication protocol, e.g., WiFi, and we also assume that the network allows for one-hop simple information exchange beyond immediate neighbors as in LoRaWAN protocol. The objective is to minimize the sum of agent-specific composite convex functions over agent-specific private constraint sets. Unlike existing decentralized primal-dual methods that require knowledge of the Lipschitz constants, D-APDB automatically adapts to unknown smoothness constants by employing a distributed backtracking step-size search. Each agent relies only on first-order oracles associated with its own objective and constraint functions and on local communications with the neighboring agents, without any prior knowledge of Lipschitz constants. We establish $\mathcal{O}(1/K)$ convergence guarantees for sub-optimality, infeasibility and consensus violation, under standard assumptions on smoothness and on the connectivity of the communication graph. To our knowledge, when nodes have private constraints, especially when they are nonlinear convex constraints onto which projections are not cheap to compute, D-APDB is the first distributed method with backtracking that achieves the optimal convergence rate for the class of constrained composite convex optimization problems. We also provide numerical results for D-APDB on a distributed QCQP problem and distributed primal SVM training to illustrate the potential performance gains that can be achieved by D-APDB.

math.OC↗

Adaptive Algorithms for Robust Phase Retrieval

This paper considers the robust phase retrieval, which can be cast as a nonsmooth and nonconvex composite optimization problem. We propose two first-order algorithms with adaptive step sizes: the subgradient algorithm (AdaSubGrad) and the inexact proximal linear algorithm (AdaIPL). Our contribution lies in the novel design of adaptive step sizes based on quantiles of the absolute residuals. Local linear convergences of both algorithms are analyzed under different regimes for the hyper-parameters. Numerical experiments on synthetic datasets and image recovery also demonstrate that our methods are competitive against the existing methods in the literature utilizing predetermined (possibly impractical) step sizes, such as the subgradient methods and the inexact proximal linear method.

math.OC↗

Accelerated Gradient Methods with Biased Gradient Estimates: Risk Sensitivity, High-Probability Guarantees, and Large Deviation Bounds

We study trade-offs between convergence rate and robustness to gradient errors in the context of first-order methods. Our focus is on generalized momentum methods (GMMs)--a broad class that includes Nesterov's accelerated gradient, heavy-ball, and gradient descent methods--for minimizing smooth strongly convex objectives. We allow stochastic gradient errors that may be adversarial and biased, and quantify robustness of these methods to gradient errors via the risk-sensitive index (RSI) from robust control theory. For quadratic objectives with i.i.d. Gaussian noise, we give closed form expressions for RSI in terms of solutions to 2x2 matrix Riccati equations, revealing a Pareto frontier between RSI and convergence rate over the choice of step-size and momentum parameters. We then prove a large-deviation principle for time-averaged suboptimality in the large iteration limit and show that the rate function is, up to a scaling, the convex conjugate of the RSI function. We further show that the rate function and RSI are linked to the $H_\infty$-norm--a measure of robustness to the worst-case deterministic gradient errors--so that stronger worst-case robustness (smaller $H_\infty$-norm) leads to sharper decay of the tail probabilities for the average suboptimality. Beyond quadratics, under potentially biased sub-Gaussian gradient errors, we derive non-asymptotic bounds on a finite-time analogue of the RSI, yielding finite-time high-probability guarantees and non-asymptotic large-deviation bounds for the averaged iterates. In the case of smooth strongly convex functions, we also observe an analogous trade-off between RSI and convergence-rate bounds. To our knowledge, these are the first non-asymptotic guarantees for GMMs with biased gradients and the first risk-sensitive analysis of GMMs. Finally, we provide numerical experiments on a robust regression problem to illustrate our results.

math.OC↗

A Retraction-free Method for Nonsmooth Minimax Optimization over a Compact Manifold

We study the minimax problem $\min_{x\in M} \max_y f_r(x,y):=f(x,y)-h(y)$, where $M$ is a compact submanifold, $f$ is continuously differentiable in $(x, y)$, $h$ is a closed, weakly-convex (possibly non-smooth) function and we assume that the regularized coupling function $-f_r(x,\cdot)$ is either $μ$-PL for some $μ>0$ or concave ($μ= 0$) for any fixed $x$ in the vicinity of $M$. To address the nonconvexity due to the manifold constraint, we use an exact penalty for the constraint $x \in M$, and enforcing a convex constraint $x\in X$ for some $X \supset M$, onto which projections can be computed efficiently. Building upon this new formulation for the manifold minimax problem in question, a single-loop smoothed manifold gradient descent-ascent (sm-MGDA) algorithm is proposed. Theoretically, any limit point of sm-MGDA sequence is a stationary point of the manifold minimax problem and sm-MGDA can generate an $O(ε)$-stationary point of the original problem with $O(1/ε^2)$ and $\tilde{O}(1/ε^4)$ complexity for $μ> 0$ and $μ= 0$ scenarios, respectively. Moreover, for the $μ= 0$ setting, through adopting Tikhonov regularization of the dual, one can improve the complexity to $O(1/ε^3)$ at the expense of asymptotic stationarity. The key component, common in the analysis of all cases, is to connect $ε$-stationary points between the penalized problem and the original problem by showing that the constraint $x \in X$ becomes inactive and the penalty term tends to $0$ along any convergent subsequence. To our knowledge, sm-MGDA is the first retraction-free algorithm for minimax problems over compact submanifolds, and this is a very desirable algorithmic property since through avoiding retractions, one can get away with matrix orthogonalization subroutines required for computing retractions to manifolds arising in practice, which are not GPU friendly.

math.OC↗

A Randomized Block-Coordinate Primal-Dual Method for Large-scale Stochastic Saddle Point Problems

We consider (stochastic) convex-concave saddle point (SP) problems with high-dimensional decision variables, arising in various applications including machine learning problems. To contend with the challenges in computing full gradients, we employ a randomized block-coordinate primal-dual scheme in which randomly selected primal and dual blocks of variables are updated. We consider both deterministic and stochastic settings, where deterministic partial gradients and their randomly sampled estimates are used, respectively, at each iteration. We investigate the convergence of the proposed method under different blocking strategies and provide the corresponding complexity results. While the best-known computational complexity result for computing a saddle point with $\varepsilon$ primal-dual gap for deterministic primal-dual methods using full gradients is $\mathcal O(\max\{m,n\}^2/\varepsilon)$, where $m$ and $n$ denote the dimensions of primal and dual variables, respectively, we show that our proposed randomized block-coordinate method achieves an improved complexity of $\mathcal O(mn/\varepsilon)$ assuming a coordinate-friendly structure on the problem. Moreover, for the stochastic setting where a mini-batch sample gradient is utilized, we show a computational complexity of $\tilde{\mathcal{O}}(m^2n^2/\varepsilon^2)$ through acceleration. Finally, almost sure convergence of the iterate sequence to a saddle point is established.

math.OC↗

AGDA+: Proximal Alternating Gradient Descent Ascent Method with a Nonmonotone Adaptive Step-Size Search for Nonconvex Minimax Problems

We consider double-regularized nonconvex-strongly concave (NCSC) minimax problems of the form $(P):\min_{x\in\mathcal{X}} \max_{y\in\mathcal{Y}}g(x)+f(x,y)-h(y)$, where $g$, $h$ are closed convex, $f$ is $L$-smooth in $(x,y)$ and strongly concave in $y$. We propose a proximal alternating gradient descent ascent method AGDA+ that can adaptively choose nonmonotone primal-dual stepsizes to compute an approximate stationary point for $(P)$ without requiring the knowledge of the global Lipschitz constant $L$ and the concavity modulus $μ$. Using a nonmonotone step-size search (backtracking) scheme, AGDA+ stands out by its ability to exploit the local Lipschitz structure and eliminates the need for precise tuning of hyper-parameters. AGDA+ achieves the optimal iteration complexity of $\mathcal{O}(ε^{-2})$ and it is the first step-size search method for NCSC minimax problems that require only $3$ calls to $\nabla f$ on average per backtracking iteration. The numerical experiments demonstrate its robustness and efficiency.

math.OC↗

High-probability complexity guarantees for nonconvex minimax problems

Stochastic smooth nonconvex minimax problems are prevalent in machine learning, e.g., GAN training, fair classification, and distributionally robust learning. Stochastic gradient descent ascent (GDA)-type methods are popular in practice due to their simplicity and single-loop nature. However, there is a significant gap between the theory and practice regarding high-probability complexity guarantees for these methods on stochastic nonconvex minimax problems. Existing high-probability bounds for GDA-type single-loop methods only apply to convex/concave minimax problems and to particular non-monotone variational inequality problems under some restrictive assumptions. In this work, we address this gap by providing the first high-probability complexity guarantees for nonconvex/PL minimax problems corresponding to a smooth function that satisfies the PL-condition in the dual variable. Specifically, we show that when the stochastic gradients are light-tailed, the smoothed alternating GDA method can compute an $\varepsilon$-stationary point within $O(\frac{\ell κ^2 δ^2}{\varepsilon^4} + \fracκ{\varepsilon^2}(\ell+δ^2\log({1}/{\bar{q}})))$ stochastic gradient calls with probability at least $1-\bar{q}$ for any $\bar{q}\in(0,1)$, where $μ$ is the PL constant, $\ell$ is the Lipschitz constant of the gradient, $κ=\ell/μ$ is the condition number, and $δ^2$ denotes a bound on the variance of stochastic gradients. We also present numerical results on a nonconvex/PL problem with synthetic data and on distributionally robust optimization problems with real data, illustrating our theoretical findings.

math.OC↗

SAPD+: An Accelerated Stochastic Method for Nonconvex-Concave Minimax Problems

We propose a new stochastic method SAPD+ for solving nonconvex-concave minimax problems of the form $\min\max\mathcal{L}(x,y)=f(x)+Φ(x,y)-g(y)$, where $f,g$ are closed convex and $Φ(x,y)$ is a smooth function that is weakly convex in $x$, (strongly) concave in $y$. Let $δ^2$ denote the variance bound for the unbiased stochastic oracle used within SAPD+ to estimate $\nablaΦ$. When $δ>0$, for both strongly concave and merely concave settings, SAPD+ achieves the best known oracle complexities: $\mathcal{O}\Big(κ_y\max\Big\{1,\frac{δ^2}{ε^2}\Big\}\frac{L\mathcal{G}_0}{ε^{2}}\Big)$ for the strongly concave case without assuming compactness of the problem domain, and $\mathcal{O}\Big(\frac{L^3\mathcal{D}_y^2\mathcal{G}_0}{ε^{4}}\Big(1+\frac{δ^2}{ε^2}\Big)\Big)$ for the merely concave case, where $κ_y\geq 1$ is the condition number, $L$ is the Lipschitz constant of $\nabla Φ$, $\mathcal{G}_0$ is the primal-dual gap of the initial point, and $\mathcal{D}_y=\sup\{\|y\|:\ y\in\mathbf{dom} g\}$. We also propose SAPD+ with variance reduction, which enjoys $\mathcal{O}\Big(\max\Big\{κ_y,\sqrt{\fracδε}\Big\}\cdot (1+κ_y\fracδε)\frac{L\mathcal{G}_0}{ε^2}\Big)$ oracle complexity for weakly convex-strongly concave setting --this is the best known upper complexity bound in the literature for this setting and our paper establishes it for the first time. We demonstrate the efficiency of SAPD+ on a distributionally robust learning problem with a nonconvex regularizer and also on a multi-class classification problem in deep learning.

math.OC↗

Robust Accelerated Primal-Dual Methods for Computing Saddle Points

We consider strongly-convex-strongly-concave saddle point problems assuming we have access to unbiased stochastic estimates of the gradients. We propose a stochastic accelerated primal-dual (SAPD) algorithm and show that SAPD sequence, generated using constant primal-dual step sizes, linearly converges to a neighborhood of the unique saddle point. Interpreting the size of the neighborhood as a measure of robustness to gradient noise, we obtain explicit characterizations of robustness in terms of SAPD parameters and problem constants. Based on these characterizations, we develop computationally tractable techniques for optimizing the SAPD parameters, i.e., the primal and dual step sizes, and the momentum parameter, to achieve a desired trade-off between the convergence rate and robustness on the Pareto curve. This allows SAPD to enjoy fast convergence properties while being robust to noise as an accelerated method. SAPD admits convergence guarantees for the distance metric with a variance term optimal up to a logarithmic factor -which can be removed by employing a restarting strategy. We also discuss how convergence and robustness results extend to the convex-concave setting. Finally, we illustrate our framework on distributionally robust logistic regression problem.

math.OC↗

Jointly Improving the Sample and Communication Complexities in Decentralized Stochastic Minimax Optimization

We propose a novel single-loop decentralized algorithm called DGDA-VR for solving the stochastic nonconvex strongly-concave minimax problem over a connected network of $M$ agents. By using stochastic first-order oracles to estimate the local gradients, we prove that our algorithm finds an $ε$-accurate solution with $\mathcal{O}(ε^{-3})$ sample complexity and $\mathcal{O}(ε^{-2})$ communication complexity, both of which are optimal and match the lower bounds for this class of problems. Unlike competitors, our algorithm does not require multiple communications for the convergence results to hold, making it applicable to a broader computational environment setting. To the best of our knowledge, this is the first such algorithm to jointly optimize the sample and communication complexities for the problem considered here.

math.OC↗

A Fast Row-Stochastic Decentralized Method for Distributed Optimization Over Directed Graphs

In this paper, we introduce a fast row-stochastic decentralized algorithm, referred to as FRSD, to solve consensus optimization problems over directed communication graphs. The proposed algorithm only utilizes row-stochastic weights, leading to certain practical advantages in broadcast communication settings over those requiring column-stochastic weights. Under the assumption that each node-specific function is smooth and strongly convex, we show that the FRSD iterate sequence converges with a linear rate to the optimal consensus solution. In contrast to the existing methods for directed networks, FRSD enjoys linear convergence without employing a gradient tracking (GT) technique explicitly, rather it implements GT implicitly with the use of a novel momentum term, which leads to a significant reduction in communication and storage overhead for each node when FRSD is implemented for solving high-dimensional problems over small-to-medium scale networks. In the numerical tests, we compare FRSD with other state-of-the-art methods, which use row-stochastic and/or column-stochastic weights.

math.OC↗

High Probability and Risk-Averse Guarantees for a Stochastic Accelerated Primal-Dual Method

We consider stochastic strongly-convex-strongly-concave (SCSC) saddle point (SP) problems which frequently arise in applications ranging from distributionally robust learning to game theory and fairness in machine learning. We focus on the recently developed stochastic accelerated primal-dual algorithm (SAPD), which admits optimal complexity in several settings as an accelerated algorithm. We provide high probability guarantees for convergence to a neighborhood of the saddle point that reflects accelerated convergence behavior. We also provide an analytical formula for the limiting covariance matrix of the iterates for a class of stochastic SCSC quadratic problems where the gradient noise is additive and Gaussian. This allows us to develop lower bounds for this class of quadratic problems which show that our analysis is tight in terms of the high probability bound dependency to the parameters. We also provide a risk-averse convergence analysis characterizing the ``Conditional Value at Risk'', the ``Entropic Value at Risk'', and the $χ^2$-divergence of the distance to the saddle point, highlighting the trade-offs between the bias and the risk associated with an approximate solution obtained by terminating the algorithm at any iteration.

math.OC↗

A Decentralized Primal-dual Method for Constrained Minimization of a Strongly Convex Function

We propose decentralized primal-dual methods for cooperative multi-agent consensus optimization problems over both static and time-varying communication networks, where only local communications are allowed. The objective is to minimize the sum of agent-specific convex functions over conic constraint sets defined by agent-specific nonlinear functions; hence, the optimal consensus decision should lie in the intersection of these private sets. Under the strong convexity assumption, we provide convergence rates for sub-optimality, infeasibility, and consensus violation in terms of the number of communications required; examine the effect of underlying network topology on the convergence rates.

math.OC↗

A Variance-Reduced Stochastic Accelerated Primal Dual Algorithm

In this work, we consider strongly convex strongly concave (SCSC) saddle point (SP) problems $\min_{x\in\mathbb{R}^{d_x}}\max_{y\in\mathbb{R}^{d_y}}f(x,y)$ where $f$ is $L$-smooth, $f(.,y)$ is $μ$-strongly convex for every $y$, and $f(x,.)$ is $μ$-strongly concave for every $x$. Such problems arise frequently in machine learning in the context of robust empirical risk minimization (ERM), e.g. $\textit{distributionally robust}$ ERM, where partial gradients are estimated using mini-batches of data points. Assuming we have access to an unbiased stochastic first-order oracle we consider the stochastic accelerated primal dual (SAPD) algorithm recently introduced in Zhang et al. [2021] for SCSC SP problems as a robust method against gradient noise. In particular, SAPD recovers the well-known stochastic gradient descent ascent (SGDA) as a special case when the momentum parameter is set to zero and can achieve an accelerated rate when the momentum parameter is properly tuned, i.e., improving the $κ\triangleq L/μ$ dependence from $κ^2$ for SGDA to $κ$. We propose efficient variance-reduction strategies for SAPD based on Richardson-Romberg extrapolation and show that our method improves upon SAPD both in practice and in theory.

math.OC↗

Randomized Gossiping with Effective Resistance Weights: Performance Guarantees and Applications

The effective resistance between a pair of nodes in a weighted undirected graph is defined as the potential difference induced when a unit current is injected at one node and extracted from the other, treating edge weights as the conductance values of edges. The effective resistance is a key quantity of interest in many applications, e.g., solving linear systems, Markov Chains, and continuous-time averaging networks. We consider effective resistances (ER) in the context of designing randomized gossiping methods for the consensus problem, where the aim is to compute the average of node values in a distributed manner through iteratively computing weighted averages among randomly chosen neighbors. We show that employing ER weights improves the averaging time corresponding to the traditional choice of uniform weights -the amount of improvement depends on the network structure. We illustrate these results through numerical experiments. We also present an application of the ER gossiping to distributed optimization: we numerically verified that using ER gossiping within EXTRA and DPGA-W methods improves their practical performance in terms of communication efficiency.

math.OC↗

A Primal-Dual Algorithm with Line Search for General Convex-Concave Saddle Point Problems

In this paper, we propose a primal-dual algorithm with a novel momentum term using the partial gradients of the coupling function that can be viewed as a generalization of the method proposed by Chambolle and Pock in 2016 to solve saddle point problems defined by a convex-concave function $\mathcal L(x,y)=f(x)+Φ(x,y)-h(y)$ with a general coupling term $Φ(x,y)$ that is not assumed to be bilinear. Assuming $\nabla_xΦ(\cdot,y)$ is Lipschitz for any fixed $y$, and $\nabla_yΦ(\cdot,\cdot)$ is Lipschitz, we show that the iterate sequence converges to a saddle point; and for any $(x,y)$, we derive error bounds in terms of $\mathcal L(\bar{x}_k,y)-\mathcal L(x,\bar{y}_k)$ for the ergodic sequence $\{\bar{x}_k,\bar{y}_k\}$. In particular, we show $\mathcal O(1/k)$ rate when the problem is merely convex in $x$. Furthermore, assuming $Φ(x,\cdot)$ is linear for each fixed $x$ and $f$ is strongly convex, we obtain the ergodic convergence rate of $\mathcal O(1/k^2)$ -- we are not aware of another single-loop method in the related literature achieving the same rate when $Φ$ is not bilinear. Finally, we propose a backtracking technique which does not require the knowledge of Lipschitz constants while ensuring the same convergence results. We also consider convex optimization problems with nonlinear functional constraints and we show that using the backtracking scheme, the optimal convergence rate can be achieved even when the dual domain is unbounded. We tested our method against other state-of-the-art first-order algorithms and interior-point methods for solving quadratically constrained quadratic problems with synthetic data, the kernel matrix learning, and regression with fairness constraints arising in machine learning.

math.OC↗