Searcharxiv⌕ Search

arXiv subjects

Ting Kei Pong

Publications and source records attributed to Ting Kei Pong.

At least 19 recordsLinked to original sources

A Gaussian smoothing-based zeroth-order method for Goldstein second-order stationarity

We introduce a new generalized Hessian, called the Goldstein second-order $δ$-subdifferential, and an associated notion of $(ε_1,ε_2,δ)$-second-order stationary point for continuously differentiable functions with locally Lipschitz gradients. We propose a zeroth-order algorithm based on cubic regularization and Gaussian smoothing with homotopy to find such approximate second-order stationary points for Lipschitz differentiable functions, and derive the iteration complexity under a mild coercivity-type assumption on the objective function.

math.OC↗

A smoothing moving balls approximation method for a class of conic-constrained difference-of-convex optimization problems

In this paper, we consider the problem of minimizing a difference-of-convex objective over a nonlinear conic constraint, where the cone is closed, convex, pointed and has a nonempty interior. We assume that the support function of a compact base of the polar cone exhibits a majorizing smoothing approximation, a condition that is satisfied by widely studied cones such as $\mathbb{R}^m_-$ and ${\cal S}^m_-$. Leveraging this condition, we reformulate the conic constraint equivalently as a single constraint involving the aforementioned support function, and adapt the moving balls approximation (MBA) method for its solution. In essence, in each iteration of our algorithm, we approximate the support function by a smooth approximation function and apply one MBA step. The subproblems that arise in our algorithm always involve only one single inequality constraint, and can thus be solved efficiently via one-dimensional root-finding procedures. We design explicit rules to evolve the smooth approximation functions from iteration to iteration and establish the corresponding iteration complexity for obtaining an $(ε_1, ε_2, ε_1 \sqrt{ε_2})$-Karush-Kuhn-Tucker point. In addition, in the convex setting, we establish convergence of the sequence generated, and study its local convergence rate under a standard Hölderian growth condition. Finally, we perform numerical experiments to illustrate the performance of our algorithm.

math.OC↗

Burer-Monteiro factorizability of nuclear norm regularized optimization

This paper studies the relationship between the nuclear norm-regularized minimization problem, which minimizes the sum of a $C^2$ function $h$ and a positive multiple of the nuclear norm, denoted by $f$, and its factorized problem obtained by the Burer-Monteiro technique. We are interested in deriving conditions that ensure every second-order stationary point of the factorized problem corresponds to a global minimizer of $f$, a property we call the $r$-factorizability of $f$ in this paper. Under suitable restricted isometry property (RIP) type assumptions on $h$, we prove the $r$-factorizability of $f$. Moreover, the RIP constant in our paper is tight, in the sense that concrete non-$r$-factorizable $f$ can be constructed when the RIP-type assumption fails to hold. Our technique for constructing such examples is novel and may be of independent interest: specifically, we use a variant of the Von Neumann's trace inequality and relate the existence of such examples to the optimal value of a quadratic program involving the RIP constant, then we explicitly solve this optimization problem to identify the parameter regimes in which such worst-case counterexamples can be constructed.

math.OC↗

Change Point Detection in Precision Matrices with D-trace Loss

We consider the problem of estimating a time-varying sparse precision matrix, which is assumed to evolve in a piecewise constant manner. Building upon the Group Fused LASSO and LASSO penalty functions, we estimate both the precision matrix and the change points. We propose an alternative estimator to the commonly employed Gaussian likelihood loss, namely the D-trace loss. We provide the conditions for the consistency of the estimated change points and of the sparse estimators in each block. We show that the solutions to the corresponding estimation problem exist when some conditions relating to the tuning parameters of the penalty functions are satisfied. Unfortunately, these conditions are not verifiable in general, posing challenges for tuning the parameters in practice. To address this issue, we introduce a modified regularizer and develop a revised problem that always admits solutions: these solutions can be used for detecting possible unsolvability of the original problem or obtaining a solution of the original problem otherwise. An alternating direction method of multipliers (ADMM) is then proposed to solve the revised problem. The relevance of the method is illustrated through numerical experiments.

math.ST↗

A smoothing extended sequential quadratic method for difference-of-convex optimization over a convex composite inequality constraint

We consider the problem of minimizing a difference-of-convex objective over a convex composite inequality constraint and a compact convex set constraint. To solve this problem, we extend the ESQM in [1] via incorporating a variable smoothing scheme. In essence, in each iteration of our algorithm, we apply one proximal gradient step to a smoothed penalty function, constructed based on a smooth approximation of the convex composite constraint function; and we design explicit rules to update the smoothing and penalty parameters. Under suitable constraint qualifications, we establish an iteration complexity of $O(ε^{-3})$ for obtaining an $(ε,ε)$-KKT point. Moreover, in the convex setting, we show that the whole sequence generated by our algorithm is convergent and derive its local convergence rate under a standard Hölderian growth condition.

math.OC↗

A conditional-gradient-based single-loop augmented Lagrangian method for inequality constrained optimization

We consider the problem of minimizing the sum of a Lipschitz differentiable convex function $f$ and a proper closed convex function $h$ that admits efficient linear minimization oracles, subject to multiple smooth convex inequality constraints. We adapt the classical augmented Lagrangian (AL) method for these problems: in each iteration, our algorithm consists of one step of the conditional gradient (CG) method applied to the AL function, followed by an update of the dual variable as in classical AL methods with a diminishing dual stepsize. We study the convergence rate of our algorithm under two standard stepsize rules for the CG method, namely, an open-loop stepsize and the short stepsize, and obtain a convergence rate that matches the best-known complexity for this class of problems. We also establish accelerated rates when $h$ is the indicator function of a uniformly convex set.

math.OC↗

Complexity and convergence analysis of a single-loop SDCAM for Lipschitz composite optimization and beyond

We develop and analyze a single-loop algorithm for minimizing the sum of a Lipschitz differentiable function $f$, a prox-friendly proper closed function $g$ (with a closed domain on which $g$ is continuous) and the composition of another prox-friendly proper closed function $h$ (whose domain is closed on which $h$ is continuous) with a continuously differentiable mapping $c$ (that is Lipschitz continuous and Lipschitz differentiable on the convex closure of the domain of $g$). Such models arise naturally in many contemporary applications, where $f$ is the loss function for data misfit, and $g$ and $h$ are nonsmooth functions for inducing desirable structures in $x$ and $c(x)$. Existing single-loop algorithms mainly focus either on the case where $h$ is Lipschitz continuous or the case where $h$ is an indicator function of a closed convex set. In this paper, we develop a single-loop algorithm for more general possibly non-Lipschitz $h$. Our algorithm is a single-loop variant of the successive difference-of-convex approximation method (SDCAM) proposed in [22]. We show that when $h$ is Lipschitz, our algorithm exhibits an iteration complexity that matches the best known complexity result for obtaining an $(ε_1,ε_2,0)$-stationary point. Moreover, we show that, by assuming additionally that dom $g$ is compact, our algorithm exhibits an iteration complexity of $\tilde{O}(ε^{-4})$ for obtaining an $(ε,ε,ε)$-stationary point when $h$ is merely continuous and real-valued. Furthermore, we consider a scenario where $h$ does not have full domain and establish vanishing bounds on successive changes of iterates. Finally, in all three cases mentioned above, we show that one can construct a subsequence such that any accumulation point $x^*$ satisfies $c(x^*)\in$ dom $h$, and if a standard constraint qualification holds at $x^*$, then $x^*$ is a stationary point.

math.OC↗

Error bounds for perspective cones of a class of nonnegative Legendre functions

Error bounds play a central role in the study of conic optimization problems, including the analysis of convergence rates for numerous algorithms. Curiously, those error bounds are often Hölderian with exponent 1/2. In this paper, we try to explain the prevalence of the 1/2 exponent by investigating generic properties of error bounds for conic feasibility problems where the underlying cone is a perspective cone constructed from a nonnegative Legendre function on $\mathbb{R}$. Our analysis relies on the facial reduction technique and the computation of one-step facial residual functions (1-FRFs). Specifically, under appropriate assumptions on the Legendre function, we show that 1-FRFs can be taken to be Hölderian of exponent 1/2 almost everywhere with respect to the two-dimensional Hausdorff measure. This enables us to further establish that having a uniform Hölderian error bound with exponent 1/2 is a generic property for a class of feasibility problems involving these cones.

math.OC↗

A single-loop proximal-conditional-gradient penalty method

We consider the problem of minimizing a convex separable objective (as a separable sum of two proper closed convex functions $f$ and $g$) over a linear coupling constraint. We assume that $f$ can be decomposed as the sum of a smooth part having Hölder continuous gradient (with exponent $μ\in(0,1]$) and a nonsmooth part that admits efficient proximal mapping computations, while $g$ can be decomposed as the sum of a smooth part having Hölder continuous gradient (with exponent $ν\in(0,1]$) and a nonsmooth part that admits efficient linear oracles. Motivated by the recent works [1,49], we propose a single-loop variant of the standard penalty method, which we call a single-loop proximal-conditional-gradient penalty method (proxCG$^{\rm pen}_{1\ell}$), for this problem. In each iteration of proxCG$^{\rm pen}_{1\ell}$, we successively perform one proximal-gradient step involving $f$ and one conditional-gradient step involving $g$ on the quadratic penalty function, followed by an update of the penalty parameter. We present explicit rules for updating the penalty parameter and the stepsize in the conditional-gradient step in each iteration. Under a standard constraint qualification and domain boundedness assumption, we show that the objective value deviations (from the optimal value) along the sequence generated decay in the order of $t^{-\min\{μ,ν,1/2\}}$ with the associated feasibility violations decaying in the order of $t^{-1/2}$. Moreover, if the nonsmooth parts are indicator functions and the extended objective is a KL function with exponent $α\in[0,1)$, then the distances to the optimal solution set along the sequence generated by proxCG$^{\rm pen}_{1\ell}$ decay asymptotically at a rate of $t^{-(1-α)\min\{μ,ν,1/2\}}$. Finally, we illustrate numerically the behavior of proxCG$^{\rm pen}_{1\ell}$ on solving low rank Hankel matrix completion problems.

math.OC↗

Subdifferentially polynomially bounded functions and Gaussian smoothing-based zeroth-order optimization

We study the class of subdifferentially polynomially bounded (SPB) functions, which is a rich class of locally Lipschitz functions that encompasses all Lipschitz functions, all gradient- or Hessian-Lipschitz functions, and even some non-smooth locally Lipschitz functions. We show that SPB functions are compatible with Gaussian smoothing (GS), in the sense that the GS of any SPB function is well-defined and satisfies a descent lemma akin to gradient-Lipschitz functions, with the Lipschitz constant replaced by a polynomial function. Leveraging this descent lemma, we propose GS-based zeroth-order optimization algorithms with an adaptive stepsize strategy for minimizing SPB functions, and analyze their convergence rates with respect to both relative and absolute stationarity measures. Finally, we also establish the iteration complexity for achieving a $(δ, ε)$-approximate stationary point, based on a novel quantification of Goldstein stationarity via the GS gradient that could be of independent interest.

math.OC↗

Convergence analysis for a variant of manifold proximal point algorithm based on Kurdyka-Łojasiewicz property

We incorporate an iteratively reweighted strategy in the manifold proximal point algorithm (ManPPA) in [12] to solve an enhanced sparsity inducing model for identifying sparse yet nonzero vectors in a given subspace. We establish the global convergence of the whole sequence generated by our algorithm by assuming the Kurdyka-Lojasiewicz (KL) properties of suitable potential functions. We also study how the KL exponents of the different potential functions are related. More importantly, when our enhanced model and algorithm reduce, respectively, to the model and ManPPA with constant stepsize considered in [12], we show that the sequence generated converges linearly as long as the optimal value of the model is positive, and converges finitely when the limit of the sequence lies in a set of weak sharp minima. Our results improve [13, Theorem 2.4], which asserts local quadratic convergence in the presence of weak sharp minima when the constant stepsize is small.

math.OC↗

Kurdyka-Łojasiewicz exponent via Hadamard parametrization

We consider a class of $\ell_1$-regularized optimization problems and the associated smooth "over-parameterized" optimization problems built upon the Hadamard parametrization, or equivalently, the Hadamard difference parametrization (HDP). We characterize the set of second-order stationary points of the HDP-based model and show that they correspond to some stationary points of the corresponding $\ell_1$-regularized model. More importantly, we show that the Kurdyka-Lojasiewicz (KL) exponent of the HDP-based model at a second-order stationary point can be inferred from that of the corresponding $\ell_1$-regularized model under suitable assumptions. Our assumptions are general enough to cover a wide variety of loss functions commonly used in $\ell_1$-regularized models, such as the least squares loss function and the logistic loss function. Since the KL exponents of many $\ell_1$-regularized models are explicitly known in the literature, our results allow us to leverage these known exponents to deduce the KL exponents at second-order stationary points of the corresponding HDP-based models, which were previously unknown. Finally, we demonstrate how these explicit KL exponents at second-order stationary points can be applied to deducing the explicit local convergence rate of a standard gradient descent method for minimizing the HDP-based model.

math.OC↗

Tight error bounds for log-determinant cones without constraint qualifications

In this paper, without requiring any constraint qualifications, we establish tight error bounds for the log-determinant cone, which is the closure of the hypograph of the perspective function of the log-determinant function. This error bound is obtained using the recently developed framework based on one-step facial residual functions.

math.OC↗

Optimal error bounds in the absence of constraint qualifications with applications to the $p$-cones and beyond

We prove tight Hölderian error bounds for all $p$-cones. Surprisingly, the exponents differ in several ways from those that have been previously conjectured; moreover, they illuminate $p$-cones as a curious example of a class of objects that possess properties in 3 dimensions that they do not in 4 or more. Using our error bounds, we analyse least squares problems with $p$-norm regularization, where our results enable us to compute the corresponding KL exponents for previously inaccessible values of $p$. Another application is a (relatively) simple proof that most $p$-cones are neither self-dual nor homogeneous. Our error bounds are obtained under the framework of facial residual functions, and we expand it by establishing for general cones an optimality criterion under which the resulting error bound must be tight.

math.OC↗

Frank-Wolfe-type methods for a class of nonconvex inequality-constrained problems

The Frank-Wolfe (FW) method, which implements efficient linear oracles that minimize linear approximations of the objective function over a fixed compact convex set, has recently received much attention in the optimization and machine learning literature. In this paper, we propose a new FW-type method for minimizing a smooth function over a compact set defined as the level set of a single difference-of-convex function, based on new generalized linear-optimization oracles (LO). We show that these LOs can be computed efficiently with closed-form solutions in some important optimization models that arise in compressed sensing and machine learning. In addition, under a mild strict feasibility condition, we establish the subsequential convergence of our nonconvex FW-type method. Since the feasible region of our generalized LO typically changes from iteration to iteration, our convergence analysis is completely different from those existing works in the literature on FW-type methods that deal with fixed feasible regions among subproblems. Finally, motivated by the away steps for accelerating FW-type methods for convex problems, we further design an away-step oracle to supplement our nonconvex FW-type method, and establish subsequential convergence of this variant. Numerical results on the matrix completion problem with standard datasets are presented to demonstrate the efficiency of the proposed FW-type method and its away-step variant.

math.OC↗

An extended sequential quadratic method with extrapolation

We revisit and adapt the extended sequential quadratic method (ESQM) in [3] for solving a class of difference-of-convex optimization problems whose constraints are defined as the intersection of level sets of Lipschitz differentiable functions and a simple compact convex set. Particularly, for this class of problems, we develop a variant of ESQM, called ESQM with extrapolation (ESQM$_e$), which incorporates Nesterov's extrapolation techniques for empirical acceleration. Under standard constraint qualifications, we show that the sequence generated by ESQM$_e$ clusters at a critical point if the extrapolation parameters are uniformly bounded above by a certain threshold. Convergence of the whole sequence and the convergence rate are established by assuming Kurdyka-Lojasiewicz (KL) property of a suitable potential function and imposing additional differentiability assumptions on the objective and constraint functions. In addition, when the objective and constraint functions are all convex, we show that linear convergence can be established if a certain exact penalty function is known to be a KL function with exponent $\frac12$; we also discuss how the KL exponent of such an exact penalty function can be deduced from that of the original extended objective (i.e., sum of the objective and the indicator function of the constraint set). Finally, we perform numerical experiments to demonstrate the empirical acceleration of ESQM$_{e}$ over a basic version of ESQM, and illustrate its effectiveness by comparing with the natural competing algorithm SCP$_{ls}$ from [35].

math.OC↗

Generalized power cones: optimal error bounds and automorphisms

Error bounds are a requisite for trusting or distrusting solutions in an informed way. Until recently, provable error bounds in the absence of constraint qualifications were unattainable for many classes of cones that do not admit projections with known succinct expressions. We build such error bounds for the generalized power cones, using the recently developed framework of one-step facial residual functions. We also show that our error bounds are tight in the sense of that framework. Besides their utility for understanding solution reliability, the error bounds we discover have additional applications to the algebraic structure of the underlying cone, which we describe. In particular we use the error bounds to compute the dimension of the automorphism group for the generalized power cones, and to identify a set of generalized power cones that are self-dual, irreducible, nonhomogeneous, and perfect

math.OC↗

Convergence rate analysis of a Dykstra-type projection algorithm

Given closed convex sets $C_i$, $i=1,\ldots,\ell$, and some nonzero linear maps $A_i$, $i = 1,\ldots,\ell$, of suitable dimensions, the multi-set split feasibility problem aims at finding a point in $\bigcap_{i=1}^\ell A_i^{-1}C_i$ based on computing projections onto $C_i$ and multiplications by $A_i$ and $A_i^T$. In this paper, we consider the associated best approximation problem, i.e., the problem of computing projections onto $\bigcap_{i=1}^\ell A_i^{-1}C_i$; we refer to this problem as the best approximation problem in multi-set split feasibility settings (BA-MSF). We adapt the Dykstra's projection algorithm, which is classical for solving the BA-MSF in the special case when all $A_i = I$, to solve the general BA-MSF. Our Dykstra-type projection algorithm is derived by applying (proximal) coordinate gradient descent to the Lagrange dual problem, and it only requires computing projections onto $C_i$ and multiplications by $A_i$ and $A_i^T$ in each iteration. Under a standard relative interior condition and a genericity assumption on the point we need to project, we show that the dual objective satisfies the Kurdyka-Lojasiewicz property with an explicitly computable exponent on a neighborhood of the (typically unbounded) dual solution set when each $C_i$ is $C^{1,α}$-cone reducible for some $α\in (0,1]$: this class of sets covers the class of $C^2$-cone reducible sets, which include all polyhedrons, second-order cone, and the cone of positive semidefinite matrices as special cases. Using this, explicit convergence rate (linear or sublinear) of the sequence generated by the Dykstra-type projection algorithm is derived. Concrete examples are constructed to illustrate the necessity of some of our assumptions.

math.OC↗