Searcharxiv⌕ Search

arXiv subjects

Etienne de Klerk

Publications and source records attributed to Etienne de Klerk.

At least 19 recordsLinked to original sources

Further analysis and extension of the higher-order Newton method of Ahmadi, Chaudhry, and Zhang

We extend a $d$th-order Newton method for unconstrained optimization by Ahmadi, Chaudhry, and Zhang [Advances in Mathematics, 452:109808] to optimization with SOS-convex polynomial constraints. Consider the problem of minimizing a smooth function $f:\mathbb{R}^{n}\to\mathbb{R}$ subject to SOS-convex polynomial constraints. Given an iterate $x\in\mathbb{R}^n$, Ahmadi et al. define the next iterate $x^{+}$ as the minimizer of the $d$th-order Taylor expansion of $f$ at $x$ with a regularization term of degree $d^{\prime}$, where $d^{\prime}$ is the smallest even number greater than $d$, chosen such that this polynomial is SOS-convex, subject to the constraints. Constructing this polynomial and minimizing it subject to the constraints can both be reduced in time polynomial in $n$ to a semidefinite program (SDP). We prove that, if $f$ is strongly convex and the tensor of the $d$th-order partial derivatives of $f$ is Lipschitz continuous, then our method converges locally to the optimal solution $x^{\ast}$ with order $d$. We further prove that, under certain constraint qualifications, the set of active constraints at $x^{\ast}$ is identified locally in a single iteration. Next, we study the worst-case performance of the third-order Newton method in the unconstrained setting for two classes of univariate $f$ using performance estimation. Finally, we extend a globally convergent modification of the $d$th-order Newton method to the setting of SOS-convex polynomial constraints.

math.OC↗

Squared polynomial approximation kernels for the hypercube: improved error bounds and implications for Lasserre hierarchies

We propose a new family of polynomial approximation kernels for approximating nonnegative polynomials on the hypercube $[-1,1]^n$. Our Kernels produce polynomial sums-of-squares of degree $r$, achieving an $O(\log^3 r/r^2)$ error in the $\ell_1$-norm of the coefficients. This improves on the known error bound $O(1/r)$ from the literature. As a corollary, we obtain an improved convergence rate for the Lasserre hierarchy for polynomial optimization on the hypercube, again improving a known rate by Baldi and Slot from $O(1/r)$ to $O(\log^3 r/r^2)$.

math.OC↗

On the convergence rate of the boosted Difference-of-Convex Algorithm (DCA)

The difference-of-convex algorithm (DCA) is a well-established nonlinear programming technique that solves successive convex optimization problems. These sub-problems are obtained from the difference-of-convex~(DC) decompositions of the objective and constraint functions. We investigate the worst-case performance of the unconstrained DCA, with and without boosting, where boosting simply performs an additional step in the direction generated by the usual DCA method. We show that, for certain classes of DC decompositions, the boosted DCA is provably better in the worst-case than the usual DCA. While several numerical studies have reported that boosted DCA outperforms classical DCA, a theoretical explanation for this behavior has, to the best of our knowledge, not been given until now. Our proof technique relies on semidefinite programming (SDP) performance estimation.

math.OC↗

Revisiting the convergence rate of the Lasserre hierarchy for polynomial optimization over the hypercube

We revisit the problem of minimizing a given polynomial $f$ on the hypercube $[-1,1]^n$. Lasserre's hierarchy (also known as the moment- or sum-of-squares hierarchy) provides a sequence of lower bounds $\{f_{(r)}\}_{r \in \mathbb N}$ on the minimum value $f^*$, where $r$ refers to the allowed degrees in the sum-of-squares hierarchy. A natural question is how fast the hierarchy converges as a function of the parameter $r$. The current state-of-the-art is due to Baldi and Slot [SIAM J. on Applied Algebraic Geometry, 2024] and roughly shows a convergence rate of order $1/r$. Here we obtain closely related results via a different approach: the polynomial kernel method. We also discuss limitations of the polynomial kernel method, suggesting a lower bound of order $1/r^2$ for our approach.

math.OC↗

Least multivariate Chebyshev polynomials on diagonally determined sets

We consider a new multivariate generalization of the classical monic (univariate) Chebyshev polynomial that minimizes the uniform norm on the interval $[-1,1]$. Let $Π^*_n$ be the subset of polynomials of degree at most $n$ in $d$ variables, whose homogeneous part of degree $n$ has coefficients summing up to $1$. The problem is determining a polynomial in $Π^*_n$ with the smallest uniform norm on a domain $Ω$, which we call a least Chebyshev polynomial (associated with $Ω$). Our main result solves the problem for $Ω$ belonging to a non-trivial class of sets that we call diagonally-determined, and establishes the remarkable result that a least Chebyshev polynomial can be given via the classical, univariate, Chebyshev polynomial. In particular, the solution can be independent of the dimension. Diagonally-determined domains include centered balls in $\mathbb{R}^d$ in any norm, but can be non-convex and even non-simply connected. We also introduce a computational procedure, based on semidefinite programming hierarchies, to detect if a given semi-algebraic set is diagonally-determined.

math.OC↗

The link between $1$-norm approximation and effective Positivstellensatze for the hypercube

The Schmüdgen's Positivstellensatz gives a certificate to verify positivity of a strictly positive polynomial $f$ on a compact, basic, semi-algebraic set $\mathbf{K} \subset \mathbb{R}^n$. A Positivstellensatz of this type is called effective if one may bound the degrees of the polynomials appearing in the certificate in terms of properties of $f$. If $\mathbf{K} = [-1,1]^n$ and $0 < f_\min := \min_{x \in \mathbf{K}} f(x)$, then the degrees of the polynomials appearing in the certificate may be bounded by $O\left(\sqrt{\frac{f_\max - f_\min}{f_\min}}\right)$, where $f_\max := \max_{x \in \mathbf{K}} f(x)$, as was recently shown by Laurent and Slot [Optimization Letters 17:515-530, 2023]. The big-O notation suppresses dependence on $n$ and the degree $d$ of $f$. In this paper we show a similar result, but with a better dependence on $n$ and $d$. In particular, our bounds depend on the $1$-norm of the coefficients of $f$, that may readily be calculated.

math.OC↗

Optimization-Aided Construction of Multivariate Chebyshev Polynomials

This article is concerned with an extension of univariate Chebyshev polynomials of the first kind to the multivariate setting, where one chases best approximants to specific monomials by polynomials of lower degree relative to the uniform norm. Exploiting the Moment-SOS hierarchy, we devise a versatile semidefinite-programming-based procedure to compute such best approximants, as well as associated signatures. Applying this procedure in three variables leads to the values of best approximation errors for all monomials up to degree six on the euclidean ball, the simplex, and the cross-polytope. Furthermore, inspired by numerical experiments, we obtain explicit expressions for Chebyshev polynomials in two cases unresolved before, namely for the monomial $x_1^2 x_2^2 x_3$ on the euclidean ball and for the monomial $x_1^2 x_2 x_3$ on the simplex.

math.OC↗

The exact worst-case convergence rate of the alternating direction method of multipliers

Recently, semidefinite programming performance estimation has been employed as a strong tool for the worst-case performance analysis of first order methods. In this paper, we derive new non-ergodic convergence rates for the alternating direction method of multipliers (ADMM) by using performance estimation. We give some examples which show the exactness of the given bounds. We also study the linear and R-linear convergence of ADMM. We establish that ADMM enjoys a global linear convergence rate if and only if the dual objective satisfies the Polyak-Lojasiewicz (PL)inequality in the presence of strong convexity. In addition, we give an explicit formula for the linear convergence rate factor. Moreover, we study the R-linear convergence of ADMM under two new scenarios.

math.OC↗

On the rate of convergence of the Difference-of-Convex Algorithm (DCA)

In this paper, we study the convergence rate of the DCA (Difference-of-Convex Algorithm), also known as the convex-concave procedure, with two different termination criteria that are suitable for smooth and nonsmooth decompositions respectively. The DCA is a popular algorithm for difference-of-convex (DC) problems, and known to converge to a stationary point of the objective under some assumptions. We derive a worst-case convergence rate of $O(1/\sqrt{N})$ after $N$ iterations of the objective gradient norm for certain classes of DC problems, without assuming strong convexity in the DC decomposition, and give an example which shows the convergence rate is exact. We also provide a new convergence rate of $O(1/N)$ for the DCA with the second termination criterion. %In addition, we investigate the DCA with regularization. Moreover, we derive a new linear convergence rate result for the DCA under the assumption of the Polyak-Łojasiewicz inequality. The novel aspect of our analysis is that it employs semidefinite programming performance estimation.

math.OC↗

A predictor-corrector algorithm for semidefinite programming that uses the factor width cone

We propose an interior point method (IPM) for solving semidefinite programming problems (SDPs). The standard interior point algorithms used to solve SDPs work in the space of positive semidefinite matrices. Contrary to that the proposed algorithm works in the cone of matrices of constant \emph{factor width}. This adaptation makes the proposed method more suitable for parallelization than the standard IPM. We prove global convergence and provide a complexity analysis. Our work is inspired by a series of papers by Ahmadi, Dash, Majumdar and Hall, and builds upon a recent preprint by Roig-Solvas and Sznaier [arXiv:2202.12374, 2022].

math.OC↗

Convergence rate analysis of randomized and cyclic coordinate descent for convex optimization through semidefinite programming

In this paper, we study randomized and cyclic coordinate descent for convex unconstrained optimization problems. We improve the known convergence rates in some cases by using the numerical semidefinite programming performance estimation method. As a spin-off we provide a method to analyse the worst-case performance of the Gauss-Seidel iterative method for linear systems where the coefficient matrix is positive semidefinite with a positive diagonal.

math.OC↗

Convergence rate analysis of the gradient descent-ascent method for convex-concave saddle-point problems

In this paper, we study the gradient descent-ascent method for convex-concave saddle-point problems. We derive a new non-asymptotic global convergence rate in terms of distance to the solution set by using the semidefinite programming performance estimation method. The given convergence rate incorporates most parameters of the problem and it is exact for a large class of strongly convex-strongly concave saddle-point problems for one iteration. We also investigate the algorithm without strong convexity and we provide some necessary and sufficient conditions under which the gradient descent-ascent enjoys linear convergence.

math.OC↗

Revisiting semidefinite programming approaches to options pricing: complexity and computational perspectives

In this paper we consider the problem of finding bounds on the prices of options depending on multiple assets without assuming any underlying model on the price dynamics, but only the absence of arbitrage opportunities. We formulate this as a generalized moment problem and utilize the well-known Moment-Sum-of-Squares (SOS) hierarchy of Lasserre to obtain bounds on the range of the possible prices. A complementary approach (also due to Lasserre) is employed for comparison. We present several numerical examples to demonstrate the viability of our approach. The framework we consider makes it possible to incorporate different kinds of observable data, such as moment information, as well as observable prices of options on the assets of interest.

math.OC↗

Conditions for linear convergence of the gradient method for non-convex optimization

In this paper, we derive a new linear convergence rate for the gradient method with fixed step lengths for non-convex smooth optimization problems satisfying the Polyak-Lojasiewicz (PL) inequality. We establish that the PL inequality is a necessary and sufficient condition for linear convergence to the optimal value for this class of problems. We list some related classes of functions for which the gradient method may enjoy linear convergence rate. Moreover, we investigate their relationship with the PL inequality.

math.OC↗

The exact worst-case convergence rate of the gradient method with fixed step lengths for L-smooth functions

In this paper, we study the convergence rate of the gradient (or steepest descent) method with fixed step lengths for finding a stationary point of an $L$-smooth function. We establish a new convergence rate, and show that the bound may be exact in some cases, in particular when all step lengths lie in the interval $(0,1/L]$. In addition, we derive an optimal step length with respect to the new bound.

math.OC↗

Convergence rates of RLT and Lasserre-type hierarchies for the generalized moment problem over the simplex and the sphere

We consider the generalized moment problem (GMP) over the simplex and the sphere. This is a rich setting and it contains NP-hard problems as special cases, like constructing optimal cubature schemes and rational optimization. Using the Reformulation-Linearization Technique (RLT) and Lasserre-type hierarchies, relaxations of the problem are introduced and analyzed. For our analysis we assume throughout the existence of a dual optimal solution as well as strong duality. For the GMP over the simplex we prove a convergence rate of $O(1/r)$ for a linear programming, RLT-type hierarchy, where $r$ is the level of the hierarchy, using a quantitative version of Pólya's Positivstellensatz. As an extension of a recent result by Fang and Fawzi [Math. Program., 2020, https://doi.org/10.1007/s10107-020-01537-7] we prove the Lasserre hierarchy of the GMP [Math. Program., Vol. 112, 65-92, 2008] over the sphere has a convergence rate of $O(1/r^2)$. Moreover, we show the introduced linear RLT-relaxation is a generalization of a hierarchy for minimizing forms of degree $d$ over the simplex, introduced by de Klerk, Laurent and Parrilo [J. Theoretical Computer Science, Vol. 361, 210-225, 2006].

math.OC↗

Minimum energy configurations on a toric lattice as a quadratic assignment problem

We consider three known bounds for the quadratic assignment problem (QAP): an eigenvalue, a convex quadratic programming (CQP), and a semidefinite programming (SDP) bound. Since the last two bounds were not compared directly before, we prove that the SDP bound is stronger than the CQP bound. We then apply these to improve known bounds on a discrete energy minimization problem, reformulated as a QAP, which aims to minimize the potential energy between repulsive particles on a toric grid. Thus we are able to prove optimality for several configurations of particles and grid sizes, complementing earlier results by Bouman, Draisma and Van Leeuwaarden [ SIAM Journal on Discrete Mathematics, 27(3):1295--1312, 2013]. The semidefinite programs in question are too large to solve without pre-processing, and we use a symmetry reduction method by Parrilo and Permenter [Mathematical Programming, 181:51--84, 2020] to make computation of the SDP bounds possible.

math.OC↗