SearcharxivSearch

arXiv subjects

Shujun Bi

Publications and source records attributed to Shujun Bi.

17 recordsLinked to original sources

Group zero-norm regularized robust loss minimization: proximal MM method and statistical error bound

This study focuses on solving group zero-norm regularized robust loss minimization problems. We propose a proximal Majorization-Minimization (PMM) algorithm to address a class of equivalent Difference-of-Convex (DC) surrogate optimization problems. First, we present the core principles and iterative framework of the PMM method. Under the Kurdyka-{\L}ojasiewicz (KL) property assumption of the potential function, we establish the global convergence of the algorithm and characterize its local (sub)linear convergence rate. Furthermore, for linear observation models with design matrices satisfying restricted eigenvalue conditions, we derive statistical estimation error bounds between the PMM-generated iterates (including their limit points) and the ground truth solution. These bounds not only rigorously quantify the approximation accuracy of the algorithm but also extend previous results on element-wise sparse composite optimization from reference [57]. To efficiently implement the PMM framework, we develop a proximal dual semismooth Newton method for solving critical subproblems. Extensive numerical experiments on both synthetic data and the UCI benchmark demonstrate the superior computational efficiency of our PMM method compared to the proximal Alternating Direction Method of Multipliers (pADMM).

math.OC

Convergence Analysis of an Inexact MBA Method for Constrained DC Problems

This paper concerns a class of constrained difference-of-convex (DC) optimization problems in which, the constraint functions are continuously differentiable and their gradients are strictly continuous. For such nonconvex and nonsmooth optimization problems, we develop an inexact moving balls approximation (MBA) method by a workable inexactness criterion for the solution of subproblems. This criterion is proposed by leveraging a global error bound for the strongly convex program associated with parametric optimization problems. We establish the full convergence of the iterate sequence under the Kurdyka-{\L}ojasiewicz (KL) property of the constructed potential function, achieve the local convergence rate of the iterate and objective value sequences under the KL property of the potential function with exponent $q\in[1/2,1)$, and provide the iteration complexity of $O(1/\epsilon^2)$ to seek an $\epsilon$-KKT point. A verifiable condition is also presented to check whether the potential function has the KL property of exponent $q\in[1/2,1)$. To our knowledge, this is the first implementable inexact MBA method with a complete convergence certificate. Numerical comparison with DCA-MOSEK, a DC algorithm with subproblems solved by MOSEK, is conducted on $\ell_1\!-\!\ell_2$ regularized quadratically constrained optimization problems, which demonstrates the advantage of the inexact MBA in the quality of solutions and running time.

math.OC

Tilt stability of a class of nonlinear semidefinite programs

This paper concerns the tilt stability of local optimal solutions to a class of nonlinear semidefinite programs, which involves a twice continuously differentiable objective function and a convex feasible set. By leveraging the second subderivative of the extended-valued objective function and imposing a suitable restriction on the multiplier, we derive two point-based sufficient characterizations for tilt stability of local optimal solutions around which the objective function has positive semidefinite Hessians, and for a class of linear positive semidefinite cone constraint set, establish a point-based necessary characterization with a certain gap from the sufficient one. For this class of linear positive semidefinite cone constraint case, under a suitable restriction on the set of multipliers, we also establish a point-based sufficient and necessary characterization, which is weaker than the dual constraint nondegeneracy when the set of multipliers is singleton. As far as we know, this is the first work to study point-based sufficient and/or necessary characterizations for nonlinear semidefinite programs with convex constraint sets without constraint nondegeneracy conditions.

math.OC

Sparse Linear Regression: Sequential Convex Relaxation, Robust Restricted Null Space Property, and Variable Selection

For high dimensional sparse linear regression problems, we propose a sequential convex relaxation algorithm (iSCRA-TL1) by solving inexactly a sequence of truncated $\ell_1$-norm regularized minimization problems, in which the working index sets are constructed iteratively with an adaptive strategy. We employ the robust restricted null space property and sequential restricted null space property (rRNSP and rSRNSP) to provide the theoretical certificates of iSCRA-TL1. Specifically, under a mild rRNSP or rSRNSP, iSCRA-TL1 is shown to identify the support of the true $r$-sparse vector by solving at most $r$ truncated $\ell_1$-norm regularized problems, and the $\ell_1$-norm error bound of its iterates from the oracle solution is also established. As a consequence, an oracle estimator of high-dimensional linear regression problems can be achieved by solving at most $r\!+\!1$ truncated $\ell_1$-norm regularized problems. To the best of our knowledge, this is the first sequential convex relaxation algorithm to produce an oracle estimator under a weaker NSP condition within a specific number of steps, provided that the Lasso estimator lacks high quality, say, the supports of its first $r$ largest (in modulus) entries do not coincide with those of the true vector.

math.ST

Error Bounds for Rank-one Double Nonnegative Reformulations of QAP and Exact Penalties

This paper focuses on the error bounds for several equivalent rank-one doubly nonnegative (DNN) conic reformulations of the quadratic assignment problem (QAP), a class of challenging combinatorial optimization problems. We provide three equivalent rank-one DNN reformulations of the QAP, including the one proposed in \cite{Jiang21}, and establish the locally and globally Lipschitzian error bounds for their feasible sets. Then, these error bounds are employed to prove that the penalty problems induced by the difference-of-convexity (DC) reformulation of the rank-one constraint are global exact penalties, and so are the penalty problems for their Burer-Monteiro (BM) factorizations. As a byproduct, the penalty problem for the rank-one DNN reformulation in \cite{Jiang21} is shown to be a global exact penalty without the calmness assumption. Finally, we illustrate the application of these exact penalties by proposing a relaxation approach with one of them to seek a rank-one approximate feasible solution. This relaxation approach is validated to be superior to the commercial solver Gurobi for \textbf{132} benchmark instances in terms of the relative gap between the generated objective value and the known best one and the number of instances with better objective values.

math.OC

A proximal MM method for the zero-norm regularized PLQ composite optimization problem

This paper is concerned with a class of zero-norm regularized piecewise linear-quadratic (PLQ) composite minimization problems, which covers the zero-norm regularized $\ell_1$-loss minimization problem as a special case. For this class of nonconvex nonsmooth problems, we show that its equivalent MPEC reformulation is partially calm on the set of global optima and make use of this property to derive a family of equivalent DC surrogates. Then, we propose a proximal majorization-minimization (MM) method, a convex relaxation approach not in the DC algorithm framework, for solving one of the DC surrogates which is a semiconvex PLQ minimization problem involving three nonsmooth terms. For this method, we establish its global convergence and linear rate of convergence, and under suitable conditions show that the limit of the generated sequence is not only a local optimum but also a good critical point in a statistical sense. Numerical experiments are conducted with synthetic and real data for the proximal MM method with the subproblems solved by a dual semismooth Newton method to confirm our theoretical findings, and numerical comparisons with a convergent indefinite-proximal ADMM for the partially smoothed DC surrogate verify its superiority in the quality of solutions and computing time.

math.OC

Error bound of critical points and KL property of exponent $1/2$ for squared F-norm regularized factorization

This paper is concerned with the squared F(robenius)-norm regularized factorization form for noisy low-rank matrix recovery problems. Under a suitable assumption on the restricted condition number of the Hessian for the loss function, we derive an error bound to the true matrix for the non-strict critical points with rank not more than that of the true matrix. Then, for the squared F-norm regularized factorized least squares loss function, under the noisy and full sample setting we establish its KL property of exponent $1/2$ on its global minimizer set, and under the noisy and partial sample setting achieve this property for a class of critical points. These theoretical findings are also confirmed by solving the squared F-norm regularized factorization problem with an accelerated alternating minimization method.

math.OC

KL property of exponent $1/2$ of $\ell_{2,0}$-norm and DC regularized factorizations for low-rank matrix recovery

This paper is concerned with the factorization form of the rank regularized loss minimization problem. To cater for the scenario in which only a coarse estimation is available for the rank of the true matrix, an $\ell_{2,0}$-norm regularized term is added to the factored loss function to reduce the rank adaptively; and account for the ambiguities in the factorization, a balanced term is then introduced. For the least squares loss, under a restricted condition number assumption on the sampling operator, we establish the KL property of exponent $1/2$ of the nonsmooth factored composite function and its equivalent DC reformulations in the set of their global minimizers. We also confirm the theoretical findings by applying a proximal linearized alternating minimization method to the regularized factorizations.

math.OC

A proximal dual semismooth Newton method for computing zero-norm penalized QR estimator

This paper is concerned with the computation of the high-dimensional zero-norm penalized quantile regression estimator, defined as a global minimizer of the zero-norm penalized check loss function. To seek a desirable approximation to the estimator, we reformulate this NP-hard problem as an equivalent augmented Lipschitz optimization problem, and exploit its coupled structure to propose a multi-stage convex relaxation approach (MSCRA\_PPA), each step of which solves inexactly a weighted $\ell_1$-regularized check loss minimization problem with a proximal dual semismooth Newton method. Under a restricted strong convexity condition, we provide the theoretical guarantee for the MSCRA\_PPA by establishing the error bound of each iterate to the true estimator and the rate of linear convergence in a statistical sense. Numerical comparisons on some synthetic and real data show that MSCRA\_PPA not only has comparable even better estimation performance, but also requires much less CPU time.

math.OC

KL property of exponent $1/2$ of quadratic functions under nonnegative zero-norm constraints and applications

This paper focuses on the quadratic optimization over two classes of nonnegative zero-norm constraints: nonnegative zero-norm sphere constraint and zero-norm simplex constraint, which have important applications in nonnegative sparse eigenvalue problems and sparse portfolio problems, respectively. We establish the KL property of exponent 1/2 for the extended-valued objective function of these nonconvex and nonsmooth optimization problems, and use this crucial property to develop a globally and linearly convergent projection gradient descent (PGD) method. Numerical results are included for nonegative sparse principal component analysis and sparse portfolio problems with synthetic and real data to confirm the theoretical results.

math.OC

Kurdyka-Lojasiewicz Property of Zero-Norm Composite Functions

This paper focuses on a class of zero-norm composite optimization problems. For this class of nonconvex nonsmooth problems, we establish the Kurdyka-Lojasiewicz property of exponent being a half for its objective function under a suitable assumption, and provide some examples to illustrate that such an assumption is not very restricted which, in particular, involve the zero-norm regularized or constrained piecewise linear-quadratic function, the zero-norm regularized or constrained logistic regression function, the zero-norm regularized or constrained quadratic function over a sphere.

math.OC

Equivalent Lipschitz surrogates for zero-norm and rank optimization problems

This paper proposes a mechanism to produce equivalent Lipschitz surrogates for zero-norm and rank optimization problems by means of the global exact penalty for their equivalent mathematical programs with an equilibrium constraint (MPECs). Specifically, we reformulate these combinatorial problems as equivalent MPECs by the variational characterization of the zero-norm and rank function, show that their penalized problems, yielded by moving the equilibrium constraint into the objective, are the global exact penalization, and obtain the equivalent Lipschitz surrogates by eliminating the dual variable in the global exact penalty. These surrogates, including the popular SCAD function in statistics, are also difference of two convex functions (D.C.) if the function and constraint set involved in zero-norm and rank optimization problems are convex. We illustrate an application by designing a multi-stage convex relaxation approach to the rank plus zero-norm regularized minimization problem.

stat.ML

GEP-MSCRA for computing the group zero-norm regularized least squares estimator

This paper concerns with the group zero-norm regularized least squares estimator which, in terms of the variational characterization of the zero-norm, can be obtained from a mathematical program with equilibrium constraints (MPEC). By developing the global exact penalty for the MPEC, this estimator is shown to arise from an exact penalization problem that not only has a favorable bilinear structure but also implies a recipe to deliver equivalent DC estimators such as the SCAD and MCP estimators. We propose a multi-stage convex relaxation approach (GEP-MSCRA) for computing this estimator, and under a restricted strong convexity assumption on the design matrix, establish its theoretical guarantees which include the decreasing of the error bounds for the iterates to the true coefficient vector and the coincidence of the iterates after finite steps with the oracle estimator. Finally, we implement the GEP-MSCRA with the subproblems solved by a semismooth Newton augmented Lagrangian method (ALM) and compare its performance with that of SLEP and MALSAR, the solvers for the weighted $\ell_{2,1}$-norm regularized estimator, on synthetic group sparse regression problems and real multi-task learning problems. Numerical comparison indicates that the GEP-MSCRA has significant advantage in reducing error and achieving better sparsity than the SLEP and the MALSAR do.

math.ST

Calibrated zero-norm regularized LS estimator for high-dimensional error-in-variables regression

This paper is concerned with high-dimensional error-in-variables regression that aims at identifying a small number of important interpretable factors for corrupted data from many applications where measurement errors or missing data can not be ignored. Motivated by CoCoLasso due to Datta and Zou \cite{Datta16} and the advantage of the zero-norm regularized LS estimator over Lasso for clean data, we propose a calibrated zero-norm regularized LS (CaZnRLS) estimator by constructing a calibrated least squares loss with a positive definite projection of an unbiased surrogate for the covariance matrix of covariates, and use the multi-stage convex relaxation approach to compute the CaZnRLS estimator. Under a restricted eigenvalue condition on the true matrix of covariates, we derive the $\ell_2$-error bound of every iterate and establish the decreasing of the error bound sequence, and the sign consistency of the iterates after finite steps. The statistical guarantees are also provided for the CaZnRLS estimator under two types of measurement errors. Numerical comparisons with CoCoLasso and NCL (the nonconvex Lasso proposed by Poh and Wainwright \cite{Loh11}) demonstrate that CaZnRLS not only has the comparable or even better relative RSME but also has the least number of incorrect predictors identified.

math.OC

A multi-stage convex relaxation approach to noisy structured low-rank matrix recovery

This paper concerns with a noisy structured low-rank matrix recovery problem which can be modeled as a structured rank minimization problem. We reformulate this problem as a mathematical program with a generalized complementarity constraint (MPGCC), and show that its penalty version, yielded by moving the generalized complementarity constraint to the objective, has the same global optimal solution set as the MPGCC does whenever the penalty parameter is over a threshold. Then, by solving the exact penalty problem in an alternating way, we obtain a multi-stage convex relaxation approach. We provide theoretical guarantees for our approach under a mild restricted eigenvalue condition, by quantifying the reduction of the error and approximate rank bounds of the first stage convex relaxation (which is exactly the nuclear norm relaxation) in the subsequent stages and establishing the geometric convergence of the error sequence in a statistical sense. Numerical experiments are conducted for some structured low-rank matrix recovery examples to confirm our theoretical findings.

math.OC

Error bounds for rank constrained optimization problems and applications

This paper is concerned with the rank constrained optimization problem whose feasible set is the intersection of the rank constraint set $\mathcal{R}=\!\big\{X\in\mathbb{X}\ |\ {\rm rank}(X)\le \kappa\big\}$ and a closed convex set $\Omega$. We establish the local (global) Lipschitzian type error bounds for estimating the distance from any $X\in \Omega$ ($X\in\mathbb{X}$) to the feasible set and the solution set, respectively, under the calmness of a multifunction associated to the feasible set at the origin, which is specially satisfied by three classes of common rank constrained optimization problems. As an application of the local Lipschitzian type error bounds, we show that the penalty problem yielded by moving the rank constraint into the objective is exact in the sense that its global optimal solution set coincides with that of the original problem when the penalty parameter is over a certain threshold. This particularly offers an affirmative answer to the open question whether the penalty problem (32) in (Gao and Sun, 2010) is exact or not. As another application, we derive the error bounds of the iterates generated by a multi-stage convex relaxation approach to those three classes of rank constrained problems and show that the bounds are nonincreasing as the number of stages increases.

math.OC

Exact penalty decomposition method for zero-norm minimization based on MPEC formulation

We reformulate the zero-norm minimization problem as an equivalent mathematical program with equilibrium constraints and establish that its penalty problem, induced by adding the complementarity constraint to the objective, is exact. Then, by the special structure of the exact penalty problem, we propose a decomposition method that can seek a global optimal solution of the zero-norm minimization problem under the null space condition in [M. A. Khajehnejad et al. IEEE Trans. Signal. Process., 59(2011), pp. 1985-2001] by solving a finite number of weighted $l_1$-norm minimization problems. To handle the weighted $l_1$-norm subproblems, we develop a partial proximal point algorithm where the subproblems may be solved approximately with the limited memory BFGS (L-BFGS) or the semismooth Newton-CG. Finally, we apply the exact penalty decomposition method with the weighted $l_1$-norm subproblems solved by combining the L-BFGS with the semismooth Newton-CG to several types of sparse optimization problems, and compare its performance with that of the penalty decomposition method [Z. Lu and Y. Zhang, SIAM J. Optim., 23(2013), pp. 2448- 2478], the iterative support detection method [Y. L. Wang and W. T. Yin, SIAM J. Sci. Comput., 3(2010), pp. 462-491] and the state-of-the-art code FPC_AS [Z. W. Wen et al. SIAM J. Sci. Comput., 32(2010), pp. 1832-1857]. Numerical comparisons indicate that the proposed method is very efficient in terms of the recoverability and the required computing time.

math.OC