Searcharxiv⌕ Search

arXiv subjects

Long Feng

Publications and source records attributed to Long Feng.

84 records · Page 5Linked to original sources

Projected Robust PCA with Application to Smooth Image Recovery

Most high-dimensional matrix recovery problems are studied under the assumption that the target matrix has certain intrinsic structures. For image data related matrix recovery problems, approximate low-rankness and smoothness are the two most commonly imposed structures. For approximately low-rank matrix recovery, the robust principal component analysis (PCA) is well-studied and proved to be effective. For smooth matrix problem, 2d fused Lasso and other total variation based approaches have played a fundamental role. Although both low-rankness and smoothness are key assumptions for image data analysis, the two lines of research, however, have very limited interaction. Motivated by taking advantage of both features, we in this paper develop a framework named projected robust PCA (PRPCA), under which the low-rank matrices are projected onto a space of smooth matrices. Consequently, a large class of image matrices can be decomposed as a low-rank and smooth component plus a sparse component. A key advantage of this decomposition is that the dimension of the core low-rank component can be significantly reduced. Consequently, our framework is able to address a problematic bottleneck of many low-rank matrix problems: singular value decomposition (SVD) on large matrices. Theoretically, we provide explicit statistical recovery guarantees of PRPCA and include classical robust PCA as a special case.

stat.ML↗

Max-sum tests for cross-sectional dependence of high-demensional panel data

We consider a testing problem for cross-sectional dependence for high-dimensional panel data, where the number of cross-sectional units is potentially much larger than the number of observations. The cross-sectional dependence is described through a linear regression model. We study three tests named the sum test, the max test and the max-sum test, where the latter two are new. The sum test is initially proposed by Breusch and Pagan (1980). We design the max and sum tests for sparse and non-sparse residuals in the linear regressions, respectively.And the max-sum test is devised to compromise both situations on the residuals. Indeed, our simulation shows that the max-sum test outperforms the previous two tests. This makes the max-sum test very useful in practice where sparsity or not for a set of data is usually vague. Towards the theoretical analysis of the three tests, we have settled two conjectures regarding the sum of squares of sample correlation coefficients asked by Pesaran (2004 and 2008). In addition, we establish the asymptotic theory for maxima of sample correlations coefficients appeared in the linear regression model for panel data, which is also the first successful attempt to our knowledge. To study the max-sum test, we create a novel method to show asymptotic independence between maxima and sums of dependent random variables. We expect the method itself is useful for other problems of this nature. Finally, an extensive simulation study as well as a case study are carried out. They demonstrate advantages of our proposed methods in terms of both empirical powers and robustness for residuals regardless of sparsity or not.

math.ST↗

Approximate nonparametric maximum likelihood inference for mixture models via convex optimization

Nonparametric maximum likelihood (NPML) for mixture models is a technique for estimating mixing distributions that has a long and rich history in statistics going back to the 1950s, and is closely related to empirical Bayes methods. Historically, NPML-based methods have been considered to be relatively impractical because of computational and theoretical obstacles. However, recent work focusing on approximate NPML methods suggests that these methods may have great promise for a variety of modern applications. Building on this recent work, a class of flexible, scalable, and easy to implement approximate NPML methods is studied for problems with multivariate mixing distributions. Concrete guidance on implementing these methods is provided, with theoretical and empirical support; topics covered include identifying the support set of the mixing distribution, and comparing algorithms (across a variety of metrics) for solving the simple convex optimization problem at the core of the approximate NPML problem. Additionally, three diverse real data applications are studied to illustrate the methods' performance: (i) A baseball data analysis (a classical example for empirical Bayes methods), (ii) high-dimensional microarray classification, and (iii) online prediction of blood-glucose density for diabetes patients. Among other things, the empirical results demonstrate the relative effectiveness of using multivariate (as opposed to univariate) mixing distributions for NPML-based approaches.

stat.ME↗

Sorted Concave Penalized Regression

The Lasso is biased. Concave penalized least squares estimation (PLSE) takes advantage of signal strength to reduce this bias, leading to sharper error bounds in prediction, coefficient estimation and variable selection. For prediction and estimation, the bias of the Lasso can be also reduced by taking a smaller penalty level than what selection consistency requires, but such smaller penalty level depends on the sparsity of the true coefficient vector. The sorted L1 penalized estimation (Slope) was proposed for adaptation to such smaller penalty levels. However, the advantages of concave PLSE and Slope do not subsume each other. We propose sorted concave penalized estimation to combine the advantages of concave and sorted penalizations. We prove that sorted concave penalties adaptively choose the smaller penalty level and at the same time benefits from signal strength, especially when a significant proportion of signals are stronger than the corresponding adaptively selected penalty levels. A local convex approximation, which extends the local linear and quadratic approximations to sorted concave penalties, is developed to facilitate the computation of sorted concave PLSE and proven to possess desired prediction and estimation error bounds. We carry out a unified treatment of penalty functions in a general optimization setting, including the penalty levels and concavity of the above mentioned sorted penalties and mixed penalties motivated by Bayesian considerations. Our analysis of prediction and estimation errors requires the restricted eigenvalue condition on the design, not beyond, and provides selection consistency under a required minimum signal strength condition in addition. Thus, our results also sharpens existing results on concave PLSE by removing the upper sparse eigenvalue component of the sparse Riesz condition.

math.ST↗

Optimal Sign Test for High Dimensional Location Parameters

This article concerns tests for location parameters in cases where the data dimension is larger than the sample size. We propose a family of tests based on the optimality arguments in Le Cam (1986) under elliptical symmetric. The asymptotic normality of these tests are established. By maximizing the asymptotic power function, we propose an uniformly optimal test for all elliptical symmetric distributions. The optimality is also confirmed by a Monte Carlo investigation.

stat.ME↗

Spatial-Sign based High-Dimensional Location Test

In this paper, we consider the problem of testing the mean vector in the high dimensional settings. We proposed a new robust scalar transform invariant test based on spatial sign. The proposed test statistic is asymptotically normal under elliptical distributions. Simulation studies show that our test is very robust and efficient in a wide range of distributions.

stat.ME↗

High Dimensional Spatial Rank Test for Two-Sample Location Problem

This article concerns tests for the two-sample location problem when the dimension is larger than the sample size. The traditional multivariate-rank-based procedures cannot be used in high dimensional settings because the sample scatter matrix is not available. We propose a novel high-dimensional spatial rank test in this article. The asymptotic normality is established. We can allow the dimension being almost the exponential rate of the sample sizes. Simulations demonstrate that it is very robust and efficient in a wide range of distributions.

stat.ME↗

A Note on High Dimensional Two Sample Mean Test

In this paper, we propose a new scalar and shift transform invariant test statistic for the high-dimensional two-sample location test. The expectation of our test is exactly zero under the null hypothesis. And we allow the dimension could be arbitrary large. Theoretical results and simulation comparison show the good performance of our test.

stat.ME↗

Scalar-Invariant Test for High-Dimensional Regression Coefficients

This article is concerned with simultaneous tests on linear regression coefficients in high-dimensional settings. When the dimensionality is larger than the sample size, the classic $F$-test is not applicable since the sample covariance matrix is not invertible. In order to overcome this issue, both Goeman, Finos and van Houwelingen (2011) and Zhong and Chen (2011) proposed their test procedures after excluding the $(\X^{'}\X)^{-1}$ term in $F$-statistics. However, both these two test are not invariant under the group of scalar transformations. In order to treat those variables in a `fair' way, we proposed a new test statistic and establish its asymptotically normal under certain mild conditions. Simulation studies showed that our test procedure performs very well in many cases.

stat.ME↗

High Dimensional Rank Tests for Sphericity

Sphericity test plays a key role in many statistical problems. We propose Spearman's rho-type rank test and Kendall's tau-type rank test for sphericity in the high dimensional settings. We show that these two tests are equivalent. Thanks to the "blessing of dimension", we do not need to estimate any nuisance parameters. Without estimating the location parameter, we can allow the dimension to be arbitrary large. Asymptotic normality of these two tests are also established under elliptical distributions. Simulations demonstrate that they are very robust and efficient in a wide range of settings.

stat.ME↗

Nonparametric maximum likelihood approach to multiple change-point problems

In multiple change-point problems, different data segments often follow different distributions, for which the changes may occur in the mean, scale or the entire distribution from one segment to another. Without the need to know the number of change-points in advance, we propose a nonparametric maximum likelihood approach to detecting multiple change-points. Our method does not impose any parametric assumption on the underlying distributions of the data sequence, which is thus suitable for detection of any changes in the distributions. The number of change-points is determined by the Bayesian information criterion and the locations of the change-points can be estimated via the dynamic programming algorithm and the use of the intrinsic order structure of the likelihood function. Under some mild conditions, we show that the new method provides consistent estimation with an optimal rate. We also suggest a prescreening procedure to exclude most of the irrelevant points prior to the implementation of the nonparametric likelihood method. Simulation studies show that the proposed method has satisfactory performance of identifying multiple change-points in terms of estimation accuracy and computation time.

math.ST↗