SearcharxivSearch

arXiv subjects

Xiaoqi Yang

Publications and source records attributed to Xiaoqi Yang.

At least 19 recordsLinked to original sources

$\ell_{1\text{-}2}$ Regularization for Sparse Optimization: Consistency and Global Convergence

The $\ell_{1\text{-}2}$ regularization method has a strong sparsity promoting capability in approaching sparse solutions of linear inverse problems and gained successful applications in various mathematics and applied science fields. This paper aims to investigate the consistency theory and global convergent algorithms for the $\ell_{1\text{-}2}$ regularization problem. In the theoretical aspect, we introduce a notion of restricted eigenvalue condition relative to the $\ell_{1\text{-}2}$ penalty, and employ it to establish an oracle property and a recovery bound for the global solution of the $\ell_{1\text{-}2}$ regularization problem. In the algorithmic aspect, we propose two types of iterative thresholding algorithms with the truncation technique and the continuation technique, respectively, to solve the $\ell_{1\text{-}2}$ regularization problem. Moreover, under the assumption of the well-known restricted isometry property, we establish the convergence of the proposed algorithms to the ground true sparse solution within a tolerance relevant to the noise level and the recovery bound. Preliminary numerical results show that our proposed algorithms can approach the ground true sparse solution and significantly enhance the sparsity recovery capability, compared with the popular sparse optimization algorithms in the literature.

math.OC

Dynamic Proximal Gradient Algorithms for Schatten-$p$ Quasi-Norm Regularized Problems

This paper investigates numerical solution methods for the Schatten-$p$ quasi-norm regularized problem with $p \in [0,1]$, which has been widely studied for finding low-rank solutions of linear inverse problems and gained successful applications in various mathematics and applied science fields. We propose a dynamic proximal gradient algorithm that, through the use of the Cayley transformation, avoids computationally expensive singular value decompositions at each iteration, thereby significantly reducing the computational complexity. The algorithm incorporates two step size selection strategies: an adaptive backtracking search and an explicit step size rule. We establish the sublinear convergence of the proposed algorithm for all $p \in [0,1]$ within the framework of the Kurdyka-Lojasiewicz property. Notably, under mild assumptions, we show that the generated sequence converges to a stationary point of the objective function of the problem. For the special case when $p=1$, the linear convergence is further proved under the strict complementarity-type regularity condition commonly used in the linear convergence analysis of the forward-backward splitting algorithms. Preliminary numerical results validate the superior computational efficiency of the proposed algorithm.

math.OC

A Globalized Semismooth Newton Method for Prox-regular Optimization Problems

We are concerned with a class of nonconvex and nonsmooth composite optimization problems, comprising a twice differentiable function and a prox-regular function. We establish a sufficient condition for the proximal mapping of a prox-regular function to be single-valued and locally Lipschitz continuous. By virtue of this property, we propose a hybrid of proximal gradient and semismooth Newton methods for solving these composite optimization problems, which is a globalized semismooth Newton method. The whole sequence is shown to converge to an $L$-stationary point under a Kurdyka-Łojasiewicz exponent assumption. Under an additional error bound condition and some other mild conditions, we prove that the sequence converges to a nonisolated $L$-stationary point at a superlinear convergence rate. Numerical comparison with several existing second order methods reveal that our approach performs comparably well in solving both the $\ell_q(0<q<1)$ quasi-norm regularized problems and the fused zero-norm regularization problems.

math.OC

Lipschitz continuity of solution multifunctions of extended $\ell_1$ regularization problems

The Lasso and the basis pursuit in compressed sensing and machine learning are convex optimization problems with three parameters: the regularization scalar, the observation vector and the data matrix. Relative to the first two parameters, we obtain the Lipschitz continuity of the solution multifunction on its convex domain. When the data matrix of the Lasso also perturbs, where non-polyhedral structure may display, we obtain full characterizations for the Lipschitz continuity of the solution multifunction on the product of a compact and convex set in the space of first two parameters and a neighborhood of the fixed data matrix. Moreover for the solution multifunction of the Lasso, we show that the Lipschitz continuity implies its single-valuedness. Our analysis is based on polyhedron theory, a sufficient condition that ensures the Lipschitz continuity of a polyhedral multifunction with a convex domain, and an explicit representation of the solution multifunction, where the latter is a consequence of the Lipschitz continuity of the solution multifunction relative to the first two parameters.

math.OC

Relative Well-Posedness of Truncated Constrained Systems Accompanied by Variational Calculus

The paper concerns foundations of sensitivity and stability analysis in optimization and related areas, being primarily addressed truncated constrained systems. We consider general models, which are described by multifunctions between Banach spaces and concentrate on characterizing their well-posedness properties that revolve around Lipschitz stability and metric regularity relative to sets. Invoking tools of variational analysis and generalized differentiation, we introduce new robust notions of relative contingent coderivatives. The novel machinery of variational analysis leads us to establishing complete characterizations of such properties and developing basic rules of variational calculus interrelated with the obtained characterizations of well-posedness. Most of the our results valid in general infinite-dimensional settings are also new in finite dimensions.

math.OC

Avoiding strict saddle points of nonconvex regularized problems

In this paper, we consider a class of non-convex and non-smooth sparse optimization problems, which encompass most existing nonconvex sparsity-inducing terms. We show the second-order optimality conditions only depend on the nonzeros of the stationary points. We propose two damped iterative reweighted algorithms including the iteratively reweighted $\ell_1$ algorithm (DIRL$_1$) and the iteratively reweighted $\ell_2$ (DIRL$_2$) algorithm, to solve these problems. For DIRL$_1$, we show the reweighted $\ell_1$ subproblem has support identification property so that DIRL$_1$ locally reverts to a gradient descent algorithm around a stationary point. For DIRL$_2$, we show the solution map of the reweighted $\ell_2$ subproblem is differentiable and Lipschitz continuous everywhere. Therefore, the map of DIRL$_1$ and DIRL$_2$ and their inverse are Lipschitz continuous, and the strict saddle points are their unstable fixed points. By applying the stable manifold theorem, these algorithms are shown to converge only to local minimizers with randomly initialization when the strictly saddle point property is assumed.

math.OC

Out-of-distribution Detection in Medical Image Analysis: A survey

Computer-aided diagnostics has benefited from the development of deep learning-based computer vision techniques in these years. Traditional supervised deep learning methods assume that the test sample is drawn from the identical distribution as the training data. However, it is possible to encounter out-of-distribution samples in real-world clinical scenarios, which may cause silent failure in deep learning-based medical image analysis tasks. Recently, research has explored various out-of-distribution (OOD) detection situations and techniques to enable a trustworthy medical AI system. In this survey, we systematically review the recent advances in OOD detection in medical image analysis. We first explore several factors that may cause a distributional shift when using a deep-learning-based model in clinic scenarios, with three different types of distributional shift well defined on top of these factors. Then a framework is suggested to categorize and feature existing solutions, while the previous studies are reviewed based on the methodology taxonomy. Our discussion also includes evaluation protocols and metrics, as well as the challenge and a research direction lack of exploration.

cs.CV

Stability Criteria and Calculus Rules via Conic Contingent Coderivatives in Banach Spaces

This paper addresses the study of novel constructions of variational analysis and generalized differentiation that are appropriate for characterizing robust stability properties of constrained set-valued mappings/multifunctions between Banach spaces important in optimization theory and its applications. Our tools of generalized differentiation revolves around the newly introduced concept of $\varepsilon$-regular normal cone to sets and associated coderivative notions for set-valued mappings. Based on these constructions, we establish several characterizations of the central stability notion known as the relative Lipschitz-like property of set-valued mappings in infinite dimensions. Applying a new version of the constrained extremal principle of variational analysis, we develop comprehensive sum and chain rules for our major constructions of conic contingent coderivatives for multifunctions between appropriate classes of Banach spaces.

math.OC

An Inexact Projected Regularized Newton Method for Fused Zero-norms Regularization Problems

We are concerned with structured $\ell_0$-norms regularization problems, with a twice continuously differentiable loss function and a box constraint. This class of problems have a wide range of applications in statistics, machine learning and image processing. To the best of our knowledge, there is no effective algorithm in the literature for solving them. In this paper, we first obtain a polynomial-time algorithm to find a point in the proximal mapping of the fused $\ell_0$-norms with a box constraint based on dynamic programming principle. We then propose a hybrid algorithm of proximal gradient method and inexact projected regularized Newton method to solve structured $\ell_0$-norms regularization problems. The whole sequence generated by the algorithm is shown to be convergent by virtue of a non-degeneracy condition, a curvature condition and a Kurdyka-Łojasiewicz property. A superlinear convergence rate of the iterates is established under a locally Hölderian error bound condition on a second-order stationary point set, without requiring the local optimality of the limit point. Finally, numerical experiments are conducted to highlight the features of our considered model, and the superiority of our proposed algorithm.

math.OC

An inexact regularized proximal Newton method for nonconvex and nonsmooth optimization

This paper focuses on the minimization of a sum of a twice continuously differentiable function $f$ and a nonsmooth convex function. An inexact regularized proximal Newton method is proposed by an approximation to the Hessian of $f$ involving the $\varrho$th power of the KKT residual. For $\varrho=0$, we justify the global convergence of the iterate sequence for the KL objective function and its R-linear convergence rate for the KL objective function of exponent $1/2$. For $\varrho\in(0,1)$, by assuming that cluster points satisfy a locally Hölderian error bound of order $q$ on a second-order stationary point set and a local error bound of order $q>1\!+\!\varrho$ on the common stationary point set, respectively, we establish the global convergence of the iterate sequence and its superlinear convergence rate with order depending on $q$ and $\varrho$. A dual semismooth Newton augmented Lagrangian method is also developed for seeking an inexact minimizer of subproblems. Numerical comparisons with two state-of-the-art methods on $\ell_1$-regularized Student's $t$-regressions, group penalized Student's $t$-regressions, and nonconvex image restoration confirm the efficiency of the proposed method.

math.OC

Variational Analysis of Kurdyka-Łojasiewicz Property, Exponent and Modulus

The Kurdyka-Łojasiewicz (KŁ) property, exponent and modulus have played a very important role in the study of global convergence and rate of convergence for optimal algorithms. In this paper, at a stationary point of a locally lower semicontinuous function, we obtain complete characterizations of the KŁ property and the KŁ modulus via the outer limiting subdifferential of an auxilliary function and a newly-introduced subderivative function respectively. In particular, for a class of prox-regular, twice epi-differentiable and subdifferentially continuous functions, we show that the KŁ property and the KŁ modulus can be described by its Moreau envelopes and a quadratic growth condition. We apply the obtained results to establish the KŁ property with exponent $\frac12$ and to provide calculation of the modulus for a smooth function, the pointwise maximum of finitely many smooth functions and regularized functions respectively. These functions often appear in the modelling of structured optimization problems.

math.OC

A regularized Newton method for $\ell_q$-norm composite optimization problems

This paper is concerned with $\ell_q\,(0<q<1)$-norm regularized minimization problems with a twice continuously differentiable loss function. For this class of nonconvex and nonsmooth composite problems, many algorithms have been proposed to solve them and most of which are of the first-order type. In this work, we propose a hybrid of proximal gradient method and subspace regularized Newton method, named HpgSRN. The whole iterate sequence produced by HpgSRN is proved to have a finite length and converge to an $L$-type stationary point under a mild curve-ratio condition and the Kurdyka-Łojasiewicz property of the cost function, which does linearly if further a Kurdyka-Łojasiewicz property of exponent $1/2$ holds. Moreover, a superlinear convergence rate for the iterate sequence is also achieved under an additional local error bound condition. Our convergence results do not require the isolatedness and strict local minimality properties of the $L$-stationary point. Numerical comparisons with ZeroFPR, a hybrid of proximal gradient method and quasi-Newton method for the forward-backward envelope of the cost function, proposed in [A. Themelis, L. Stella, and P. Patrinos, {\em SIAM J. Optim., } 28(2018), pp. 2274-2303] for the $\ell_q$-norm regularized linear and logistic regressions on real data indicate that HpgSRN not only requires much less computing time but also yields comparable even better sparsities and objective function values.

math.OC

Projectional Coderivatives and Calculus Rules

This paper is devoted to the study of a newly introduced tool, projectional coderivatives and the corresponding calculus rules in finite dimensions. We show that when the restricted set has some nice properties, more specifically, is a smooth manifold, the projectional coderivative can be refined as a fixed-point expression. We will also improve the generalized Mordukhovich criterion to give a complete characterization of the relative Lipschitz-like property under such a setting. Chain rules and sum rules are obtained to facilitate the application of the tool to a wider range of problems.

math.OC

Relative Lipschitz-like property of parametric systems via projectional coderivative

This paper concerns upper estimates of the projectional coderivative of implicit mappings and corresponding applications on analyzing the relative Lipschitz-like property. Under different constraint qualifications, we provide upper estimates of the projectional coderivative for solution mappings of parametric systems. For the solution mapping of affine variational inequalities, a generalized critical face condition is obtained for sufficiency of its Lipschitz-like property relative to a polyhedral set within its domain under a constraint qualification. The equivalence between the relative Lipschitz-like property and the local inner-semicontinuity for polyhedral multifunctions is also demonstrated. For the solution mapping of linear complementarity problems with a $Q_0$-matrix, we establish a sufficient and necessary condition for the Lipschitz-like property relative to its convex domain via the generalized critical face condition and its combinatorial nature.

math.OC

Isolated calmness and sharp minima via Hölder graphical derivatives

The paper utilizes Hölder graphical derivatives for characterizing Hölder strong subregularity, isolated calmness and sharp minimum. As applications, we characterize Hölder isolated calmness in linear semi-infinite optimization and Hölder sharp minimizers of some penalty functions for constrained optimization.

math.OC

Lipschitz-like property relative to a set and the generalized Mordukhovich criterion

In this paper we will establish some necessary condition and sufficient condition respectively for a set-valued mapping to have the Lipschitz-like property relative to a closed set by employing regular normal cone and limiting normal cone of a restricted graph of the set-valued mapping. We will obtain a complete characterization for a set-valued mapping to have the Lipschitz-property relative to a closed and convex set by virtue of the projection of the coderivative onto a tangent cone. Furthermore, by introducing a projectional coderivative of set-valued mappings, we establish a verifiable generalized Mordukhovich criterion for the Lipschitz-like property relative to a closed and convex set. We will study the representation of the graphical modulus of a set-valued mapping relative to a closed and convex set by using the outer norm of the corresponding projectional coderivative value. For an extended real-valued function, we will apply the obtained results to investigate its Lipschitz continuity relative to a closed and convex set and the Lipschitz-like property of a level-set mapping relative to a half line.

math.OC

Fully piecewise linear vector optimization problem

We distinguish two kinds of piecewise linear functions and provide an interesting representation for a piecewise linear function between two normed spaces. Based on such a representation, we study a fully piecewise linear vector optimization (PLP) with the objective and constraint functions being piecewise linear. We divide (PLP) into some linear subproblems and structure a finite dimensional reduction method to solve (PLP). Under some mild assumptions, we prove that the Pareto (resp. weak Pareto) solution set of (PLP) is the union of finitely many generalized polyhedra (resp. polyhedra), each of which is contained in a Pareto (resp. weak Pareto) face of some linear subproblem. Our main results are even new in the linear case and further generalize Arrow, Barankin and Blackwell's classical results on linear vector optimization problems in the framework of finite dimensional spaces.

math.OC

Sparse estimation via $\ell_q$ optimization method in high-dimensional linear regression

In this paper, we discuss the statistical properties of the $\ell_q$ optimization methods $(0<q\leq 1)$, including the $\ell_q$ minimization method and the $\ell_q$ regularization method, for estimating a sparse parameter from noisy observations in high-dimensional linear regression with either a deterministic or random design. For this purpose, we introduce a general $q$-restricted eigenvalue condition (REC) and provide its sufficient conditions in terms of several widely-used regularity conditions such as sparse eigenvalue condition, restricted isometry property, and mutual incoherence property. By virtue of the $q$-REC, we exhibit the stable recovery property of the $\ell_q$ optimization methods for either deterministic or random designs by showing that the $\ell_2$ recovery bound $O(ε^2)$ for the $\ell_q$ minimization method and the oracle inequality and $\ell_2$ recovery bound $O(λ^{\frac{2}{2-q}}s)$ for the $\ell_q$ regularization method hold respectively with high probability. The results in this paper are nonasymptotic and only assume the weak $q$-REC. The preliminary numerical results verify the established statistical property and demonstrate the advantages of the $\ell_q$ regularization method over some existing sparse optimization methods.

stat.ML