SearcharxivSearch

arXiv subjects

Mengxi Yi

Publications and source records attributed to Mengxi Yi.

8 recordsLinked to original sources

Robust regularized covariance matrix estimation: well-posedness and convergent algorithm

In this paper, we study properties of penalized and structured M-estimators of multivariate scatter, based on geodesically convex but not necessarily smooth penalty functions. Existence and uniqueness conditions for these penalized and structured estimators are given. However, we show that the standard fixed-point algorithm which is usually applied to an M-estimation problem does not necessarily converge for penalized M-estimation problems. Hence, we develop a new but simple re-weighting algorithm and prove that it has monotone convergence for a broad class of penalized and structured M-estimators of multivariate scatter.

stat.ME

Stab-GKnock: Controlled variable selection for partially linear models using generalized knockoffs

The recently proposed fixed-X knockoff is a powerful variable selection procedure that controls the false discovery rate (FDR) in any finite-sample setting, yet its theoretical insights are difficult to show beyond Gaussian linear models. In this paper, we make the first attempt to extend the fixed-X knockoff to partially linear models by using generalized knockoff features, and propose a new stability generalized knockoff (Stab-GKnock) procedure by incorporating selection probability as feature importance score. We provide FDR control and power guarantee under some regularity conditions. In addition, we propose a two-stage method under high dimensionality by introducing a new joint feature screening procedure, with guaranteed sure screening property. Extensive simulation studies are conducted to evaluate the finite-sample performance of the proposed method. A real data example is also provided for illustration.

stat.ME

Robust and Resistant Regularized Covariance Matrices

We introduce a class of regularized M-estimators of multivariate scatter and show, analogous to the popular spatial sign covariance matrix (SSCM), that they possess high breakdown points. We also show that the SSCM can be viewed as an extreme member of this class. Unlike the SSCM, this class of estimators takes into account the shape of the contours of the data cloud when down-weighing observations. We also propose a median based cross validation criterion for selecting the tuning parameter for this class of regularized M-estimators. This cross validation criterion helps assure the resulting tuned scatter estimator is a good fit to the data as well as having a high breakdown point. A motivation for this new median based criterion is that when it is optimized over all possible scatter parameters, rather than only over the tuned candidates, it results in a new high breakdown point affine equivariant multivariate scatter statistic.

stat.ME

Test of the Latent Dimension of a Spatial Blind Source Separation Model

We assume a spatial blind source separation model in which the observed multivariate spatial data is a linear mixture of latent spatially uncorrelated Gaussian random fields containing a number of pure white noise components. We propose a test on the number of white noise components and obtain the asymptotic distribution of its statistic for a general domain. We also demonstrate how computations can be facilitated in the case of gridded observation locations. Based on this test, we obtain a consistent estimator of the true dimension. Simulation studies and an environmental application demonstrate that our test is at least comparable to and often outperforms bootstrap-based techniques, which are also introduced in this paper.

math.ST

On Cokriging, Neural Networks, and Spatial Blind Source Separation for Multivariate Spatial Prediction

Multivariate measurements taken at irregularly sampled locations are a common form of data, for example in geochemical analysis of soil. In practical considerations predictions of these measurements at unobserved locations are of great interest. For standard multivariate spatial prediction methods it is mandatory to not only model spatial dependencies but also cross-dependencies which makes it a demanding task. Recently, a blind source separation approach for spatial data was suggested. When using this spatial blind source separation method prior the actual spatial prediction, modelling of spatial cross-dependencies is avoided, which in turn simplifies the spatial prediction task significantly. In this paper we investigate the use of spatial blind source separation as a pre-processing tool for spatial prediction and compare it with predictions from Cokriging and neural networks in an extensive simulation study as well as a geochemical dataset.

eess.SP

Breakdown points of penalized and hybrid M-estimators of covariance

We introduce a class of hybrid M-estimators of multivariate scatter which, analogous to the popular spatial sign covariance matrix (SSCM), possess high breakdown points. We also show that the SSCM can be viewed as an extreme member of this class. Unlike the SSCM, but like the regular M-estimators of scatter, this new class of estimators takes into account the shape of the contours of the data cloud for downweighting observations.

stat.ME

Shrinking the Sample Covariance Matrix using Convex Penalties on the Matrix-Log Transformation

For $q$-dimensional data, penalized versions of the sample covariance matrix are important when the sample size is small or modest relative to $q$. Since the negative log-likelihood under multivariate normal sampling is convex in $Σ^{-1}$, the inverse of its covariance matrix, it is common to add to it a penalty which is also convex in $Σ^{-1}$. More recently, Deng-Tsui (2013) and Yu et al.(2017) have proposed penalties which are functions of the eigenvalues of $Σ$, and are convex in $\log Σ$, but not in $Σ^{-1}$. The resulting penalized optimization problem is not convex in either $\log Σ$ or $Σ^{-1}$. In this paper, we note that this optimization problem is geodesically convex in $Σ$, which allows us to establish the existence and uniqueness of the corresponding penalized covariance matrices. More generally, we show the equivalence of convexity in $\log Σ$ and geodesic convexity for penalties on $Σ$ which are strictly functions of their eigenvalues. In addition, when using such penalties, we show that the resulting optimization problem reduces to to a $q$-dimensional convex optimization problem on the eigenvalues of $Σ$, which can then be readily solved via Newton-Raphson. Finally, we argue that it is better to apply these penalties to the shape matrix $Σ/(\det Σ)^{1/q}$ rather than to $Σ$ itself. A simulation study and an example illustrate the advantages of applying the penalty to the shape matrix.

math.ST

Lassoing Eigenvalues

The properties of penalized sample covariance matrices depend on the choice of the penalty function. In this paper, we introduce a class of non-smooth penalty functions for the sample covariance matrix, and demonstrate how this method results in a grouping of the estimated eigenvalues. We refer to this method as "lassoing eigenvalues" or as the "elasso".

stat.ME