SearcharxivSearch

arXiv subjects

Rongrong Lin

Publications and source records attributed to Rongrong Lin.

10 recordsLinked to original sources

Smallest Singular Value Estimates for Nonuniform Fourier Matrices via Periodic Nonuniform Sampling

We study the smallest singular value of nonuniform Fourier matrices in two settings: clustered nodes and perturbations of an equispaced grid. By reducing the problem to spectral norm estimates for periodic nonuniform interpolation matrices, we obtain nearly optimal bounds in both cases. For clustered nodes, we derive the first local separation condition in which each required gap depends only on the sizes of the two neighboring clusters. For perturbations with the bound \(1/4\leq L<1/2\), our result confirms the conjecture of Austin and Trefethen on the \(2\)-norm Lebesgue constant up to a logarithmic factor.

math.NA

Computing the Proximal Operator of the $q$-th Power of the $\ell_{1,q}$-norm for Group Sparsity

In this note, we comprehensively characterize the proximal operator of the $q$-th power of the $\ell_{1,q}$-norm (denoted by $\ell_{1,q}^{q}$) with $0\!<\!q\!<\!1$ by exploiting the well-known proximal operator of $|\cdot|^q$ on the real line. In particular, much more explicit characterizations can be obtained whenever $q\!=\!1/2$ and $q\!=\!2/3$ due to the existence of closed-form expressions for the proximal operators of $|\cdot|^{1/2}$ and $|\cdot|^{2/3}$. Numerical experiments demonstrate potential advantages of the $\ell_{1,q}^{q}$ regularization in the }inter-group and intra-group sparse vector recovery.

math.NA

On Choosing Initial Values of Iteratively Reweighted $\ell_1$ Algorithms for the Piece-wise Exponential Penalty

Computing the proximal operator of the sparsity-promoting piece-wise exponential (PiE) penalty $1-e^{-|x|/σ}$ with a given shape parameter $σ>0$, which is treated as a popular nonconvex surrogate of $\ell_0$-norm, is fundamental in feature selection via support vector machines, image reconstruction, zero-one programming problems, compressed sensing, etc. Due to the nonconvexity of PiE, for a long time, its proximal operator is frequently evaluated via an iteratively reweighted $\ell_1$ algorithm, which substitutes PiE with its first-order approximation, however, the obtained solutions only are the critical point. Based on the exact characterization of the proximal operator of PiE, we explore how the iteratively reweighted $\ell_1$ solution deviates from the true proximal operator in certain regions, which can be explicitly identified in terms of $σ$, the initial value and the regularization parameter in the definition of the proximal operator. Moreover, the initial value can be adaptively and simply chosen to ensure that the iteratively reweighted $\ell_1$ solution belongs to the proximal operator of PiE.

math.NA

Kernel Support Vector Machine Classifiers with the $\ell_0$-Norm Hinge Loss

Support Vector Machine (SVM) has been one of the most successful machine learning techniques for binary classification problems. The key idea is to maximize the margin from the data to the hyperplane subject to correct classification on training samples. The commonly used hinge loss and its variations are sensitive to label noise, and unstable for resampling due to its unboundedness. This paper is concentrated on the kernel SVM with the $\ell_0$-norm hinge loss (referred as $\ell_0$-KSVM), which is a composite function of hinge loss and $\ell_0$-norm and then could overcome the difficulties mentioned above. In consideration of the nonconvexity and nonsmoothness of $\ell_0$-norm hinge loss, we first characterize the limiting subdifferential of the $\ell_0$-norm hinge loss and then derive the equivalent relationship among the proximal stationary point, the Karush-Kuhn-Tucker point, and the local optimal solution of $\ell_0$-KSVM. Secondly, we develop an ADMM algorithm for $\ell_0$-KSVM, and obtain that any limit point of the sequence generated by the proposed algorithm is a locally optimal solution. Lastly, some experiments on the synthetic and real datasets are illuminated to show that $\ell_0$-KSVM can achieve comparable accuracy compared with the standard KSVM while the former generally enjoys fewer support vectors.

cs.LG

The Proximal Operator of the Piece-wise Exponential Function and Its Application in Compressed Sensing

This paper characterizes the proximal operator of the piece-wise exponential function $1\!-\!e^{-|x|/σ}$ with a given shape parameter $σ\!>\!0$, which is a popular nonconvex surrogate of $\ell_0$-norm in support vector machines, zero-one programming problems, and compressed sensing, etc. Although Malek-Mohammadi et al. [IEEE Transactions on Signal Processing, 64(21):5657--5671, 2016] once worked on this problem, the expressions they derived were regrettably inaccurate. In a sense, it was lacking a case. Using the Lambert W function and an extensive study of the piece-wise exponential function, we have rectified the formulation of the proximal operator of the piece-wise exponential function in light of their work. We have also undertaken a thorough analysis of this operator. Finally, as an application in compressed sensing, an iterative shrinkage and thresholding algorithm (ISTA) for the piece-wise exponential function regularization problem is developed and fully investigated. A comparative study of ISTA with nine popular non-convex penalties in compressed sensing demonstrates the advantage of the piece-wise exponential penalty.

math.NA

On Reproducing Kernel Banach Spaces: Generic Definitions and Unified Framework of Constructions

Recently, there has been emerging interest in constructing reproducing kernel Banach spaces (RKBS) for applied and theoretical purposes such as machine learning, sampling reconstruction, sparse approximation and functional analysis. Existing constructions include the reflexive RKBS via a bilinear form, the semi-inner-product RKBS, the RKBS with $\ell^1$ norm, the $p$-norm RKBS via generalized Mercer kernels, etc. The definitions of RKBS and the associated reproducing kernel in those references are dependent on the construction. Moreover, relations among those constructions are unclear. We explore a generic definition of RKBS and the reproducing kernel for RKBS that is independent of construction. Furthermore, we propose a framework of constructing RKBSs that unifies existing constructions mentioned above via a continuous bilinear form and a pair of feature maps. A new class of Orlicz RKBSs is proposed. Finally, we develop representer theorems for machine learning in RKBSs constructed in our framework, which also unifies representer theorems in existing RKBSs.

cs.LG

Multi-task Learning in Vector-valued Reproducing Kernel Banach Spaces with the $\ell^1$ Norm

Targeting at sparse multi-task learning, we consider regularization models with an $\ell^1$ penalty on the coefficients of kernel functions. In order to provide a kernel method for this model, we construct a class of vector-valued reproducing kernel Banach spaces with the $\ell^1$ norm. The notion of multi-task admissible kernels is proposed so that the constructed spaces could have desirable properties including the crucial linear representer theorem. Such kernels are related to bounded Lebesgue constants of a kernel interpolation question. We study the Lebesgue constant of multi-task kernels and provide examples of admissible kernels. Furthermore, we present numerical experiments for both synthetic data and real-world benchmark data to demonstrate the advantages of the proposed construction and regularization models.

math.FA

An Optimal Convergence Rate for the Gaussian Regularized Shannon Sampling Series

We consider the reconstruction of a bandlimited function from its finite localized sample data. Truncating the classical Shannon sampling series results in an unsatisfactory convergence rate due to the slow decay of the sinc function. To overcome this drawback, a simple and highly effective method, called the Gaussian regularization of the Shannon series, was proposed in the engineering and has received remarkable attention. It works by multiplying the sinc function in the Shannon series with a regularized Gaussian function. Recently, it was proved that the upper error bound of this method can achieve a convergence rate of the order $O(\frac{1}{\sqrt{n}}\exp(-\frac{\pi-\delta}{2}n))$, where $0<\delta<\pi$ is the bandwidth and $n$ the number of sample data. The convergence rate is by far the best convergence rate among all regularized methods for the Shannon sampling series. The main objective of this article is to present the theoretical justification and numerical verification that the convergence rate is optimal when $0<\delta<\pi/2$ by estimating the lower error bound of the truncated Gaussian regularized Shannon sampling series.

math.NA

Convergence Analysis of the Gaussian Regularized Shannon Sampling Formula

We consider the reconstruction of a bandlimited function from its finite localized sample data. Truncating the classical Shannon sampling series results in an unsatisfactory convergence rate due to the slow decayness of the sinc function. To overcome this drawback, a simple and highly effective method, called the Gaussian regularization of the Shannon series, was proposed in the engineering and has received remarkable attention. It works by multiplying the sinc function in the Shannon series with a regularized Gaussian function. L. Qian (Proc. Amer. Math. Soc., 2003) established the convergence rate of $O(\sqrt{n}\exp(-\frac{π-δ}2n))$ for this method, where $δ<π$ is the bandwidth and $n$ is the number of sample data. C. Micchelli {\it et al.} (J. Complexity, 2009) proposed a different regularized method and obtained the corresponding convergence rate of $O(\frac1{\sqrt{n}}\exp(-\frac{π-δ}2n))$. This latter rate is by far the best among all regularized methods for the Shannon series. However, their regularized method involves the solving of a linear system and is implicit and more complicated. The main objective of this note is to show that the Gaussian regularization of the Shannon series can also achieve the same best convergence rate as that by C. Micchelli {\it et al}. We also show that the Gaussian regularization method can improve the convergence rate for the useful average sampling. Finally, the outstanding performance of numerical experiments justifies our results.

cs.IT

Existence of the Bedrosian Identity for Singular Integral Operators

The Hilbert transform $H$ satisfies the Bedrosian identity $H(fg)=fHg$ whenever the supports of the Fourier transforms of $f,g\in L^2(R)$ are respectively contained in $A=[-a,b]$ and $B=R\setminus(-b,a)$, $0\le a,b\le+\infty$. Attracted by this interesting result arising from the time-frequency analysis, we investigate the existence of such an identity for a general bounded singular integral operator on $L^2(R^d)$ and for general support sets $A$ and $B$. A geometric characterization of the support sets for the existence of the Bedrosian identity is established. Moreover, the support sets for the partial Hilbert transforms are all found. In particular, for the Hilbert transform to satisfy the Bedrosian identity, the support sets must be given as above.

math.CA