SearcharxivSearch

arXiv subjects

Yuehan Yang

Publications and source records attributed to Yuehan Yang.

15 recordsLinked to original sources

Non-zero block selector: A linear correlation coefficient measure for blocking-selection models

Multiple-group data is widely used in genomic studies, finance, and social science. This study investigates a block structure that consists of covariate and response groups. It examines the block-selection problem of high-dimensional models with group structures for both responses and covariates, where both the number of blocks and the dimension within each block are allowed to grow larger than the sample size. We propose a novel strategy for detecting the block structure, which includes the block-selection model and a non-zero block selector (NBS). We establish the uniform consistency of the NBS and propose three estimators based on the NBS to enhance modeling efficiency. We prove that the estimators achieve the oracle solution and show that they are consistent, jointly asymptotically normal, and efficient in modeling extremely high-dimensional data. Simulations generate complex data settings and demonstrate the superiority of the proposed method. A gene-data analysis also demonstrates its effectiveness.

stat.ME

An iterative algorithm for high-dimensional linear models with both sparse and non-sparse structures

Numerous practical medical problems often involve data that possess a combination of both sparse and non-sparse structures. Traditional penalized regularizations techniques, primarily designed for promoting sparsity, are inadequate to capture the optimal solutions in such scenarios. To address these challenges, this paper introduces a novel algorithm named Non-sparse Iteration (NSI). The NSI algorithm allows for the existence of both sparse and non-sparse structures and estimates them simultaneously and accurately. We provide theoretical guarantees that the proposed algorithm converges to the oracle solution and achieves the optimal rate for the upper bound of the $l_2$-norm error. Through simulations and practical applications, NSI consistently exhibits superior statistical performance in terms of estimation accuracy, prediction efficacy, and variable selection compared to several existing methods. The proposed method is also applied to breast cancer data, revealing repeated selection of specific genes for in-depth analysis.

stat.ME

A model-free feature selection technique of feature screening and random forest based recursive feature elimination

In this paper, we propose a model-free feature selection method for ultra-high dimensional data with mass features. This is a two phases procedure that we propose to use the fused Kolmogorov filter with the random forest based RFE to remove model limitations and reduce the computational complexity. The method is fully nonparametric and can work with various types of datasets. It has several appealing characteristics, i.e., accuracy, model-free, and computational efficiency, and can be widely used in practical problems, such as multiclass classification, nonparametric regression, and Poisson regression, among others. We show that the proposed method is selection consistent and $L_2$ consistent under weak regularity conditions. We further demonstrate the superior performance of the proposed method over other existing methods by simulations and real data examples.

stat.ME

Dimension reduction of high-dimension categorical data with two or multiple responses considering interactions between responses

This paper models categorical data with two or multiple responses, focusing on the interactions between responses. We propose an efficient iterative procedure based on sufficient dimension reduction. We study the theoretical guarantees of the proposed method under the two- and multiple-response models, demonstrating the uniqueness of the proposed estimator and with the high probability that the proposed method recovers the oracle least squares estimators. For data analysis, we demonstrate that the proposed method is efficient in the multiple-response model and performs better than some existing methods built in the multiple-response models. We apply this modeling and the proposed method to an adult dataset and right heart catheterization dataset and obtain meaningful results.

stat.ME

A sequential stepwise screening procedure for sparse recovery in high-dimensional multiresponse models with complex group structures

Multiresponse data with complex group structures in both responses and predictors arises in many fields, yet, due to the difficulty in identifying complex group structures, only a few methods have been studied on this problem. We propose a novel algorithm called sequential stepwise screening procedure (SeSS) for feature selection in high-dimensional multiresponse models with complex group structures. This algorithm encourages the grouping effect, where responses and predictors come from different groups, further, each response group is allowed to relate to multiple predictor groups. To obtain a correct model under the complex group structures, the proposed procedure first chooses the nonzero block and the nonzero row by the canonical correlation measure (CC) and then selects the nonzero entries by the extended Bayesian Information Criterion (EBIC). We show that this method is accurate in extremely sparse models and computationally attractive. The theoretical property of SeSS is established. We conduct simulation studies and consider a real example to compare its performances with existing methods.

stat.ME

Randomization-based joint central limit theorem and efficient covariate adjustment in stratified $2^K$ factorial experiments

Randomized block factorial experiments are widely used in industrial engineering, clinical trials, and social science. Researchers often use a linear model and analysis of covariance to analyze experimental results; however, limited studies have addressed the validity and robustness of the resulting inferences because assumptions for a linear model might not be justified by randomization in randomized block factorial experiments. In this paper, we establish a new finite population joint central limit theorem for usual (unadjusted) factorial effect estimators in randomized block $2^K$ factorial experiments. Our theorem is obtained under a randomization-based inference framework, making use of an extension of the vector form of the Wald--Wolfowitz--Hoeffding theorem for a linear rank statistic. It is robust to model misspecification, numbers of blocks, block sizes, and propensity scores across blocks. To improve the estimation and inference efficiency, we propose four covariate adjustment methods. We show that under mild conditions, the resulting covariate-adjusted factorial effect estimators are consistent, jointly asymptotically normal, and generally more efficient than the unadjusted estimator. In addition, we propose Neyman-type conservative estimators for the asymptotic covariances to facilitate valid inferences. Simulation studies and a clinical trial data analysis demonstrate the benefits of the covariate adjustment methods.

stat.ME

Design-based theory for Lasso adjustment in randomized block experiments and rerandomized experiments

Blocking, a special case of rerandomization, is routinely implemented in the design stage of randomized experiments to balance the baseline covariates. This study proposes a regression adjustment method based on the least absolute shrinkage and selection operator (Lasso) to efficiently estimate the average treatment effect in randomized block experiments with high-dimensional covariates. We derive the asymptotic properties of the proposed estimator and outline the conditions under which this estimator is more efficient than the unadjusted one. We provide a conservative variance estimator to facilitate valid inferences. Our framework allows one treated or control unit in some blocks and heterogeneous propensity scores across blocks, thus including paired experiments and finely stratified experiments as special cases. We further accommodate rerandomized experiments and a combination of blocking and rerandomization. Moreover, our analysis allows both the number of blocks and block sizes to tend to infinity, as well as heterogeneous treatment effects across blocks without assuming a true outcome data-generating model. Simulation studies and two real-data analyses demonstrate the advantages of the proposed method.

stat.ME

Interaction Pursuit Biconvex Optimization

Multivariate regression models are widely used in various fields such as biology and finance. In this paper, we focus on two key challenges: (a) When should we favor a multivariate model over a series of univariate models; (b) If the numbers of responses and predictors are allowed to greatly exceed the sample size, how to reduce the computational cost and provide precise estimation. The proposed method, Interaction Pursuit Biconvex Optimization (IPBO), explores the regression relationship allowing the predictors and responses derived from different multivariate normal distributions with general covariance matrices. In practice, the correlation structures within are complex and interact on each other based on the regression function. The proposed method solves this problem by building a structured sparsity penalty to encourages the shared structure between the network and the regression coefficients. We prove theoretical results under interpretable conditions, and provide an efficient algorithm to compute the estimator. Simulation studies and real data examples compare the proposed method with several existing methods, indicating that IPBO works well.

stat.ME

Regression-adjusted average treatment effect estimates in stratified randomized experiments

Researchers often use linear regression to analyse randomized experiments to improve treatment effect estimation by adjusting for imbalances of covariates in the treatment and control groups. Our work offers a randomization-based inference framework for regression adjustment in stratified randomized experiments. Under mild conditions, we re-establish the finite population central limit theorem for a stratified experiment. We prove that both the stratified difference-in-means and the regression-adjusted average treatment effect estimators are consistent and asymptotically normal. The asymptotic variance of the latter is no greater and is typically lesser than that of the former. We also provide conservative variance estimators to construct large-sample confidence intervals for the average treatment effect.

math.ST

MSP: A Multi-step Screening Procedure for Sparse Recovery

We propose a Multi-step Screening Procedure (MSP) for the recovery of sparse linear models in high-dimensional data. This method is based on a repeated small penalty strategy that quickly converges to an estimate within a few iterations. Specifically, in each iteration, an adaptive lasso regression with a small penalty is fit within the reduced feature space obtained from the previous step, rendering its computational complexity roughly comparable with the Lasso. MSP is shown to select the true model under complex correlation structures among the predictors and response, even when the irrepresentable condition fails. Further, under suitable regularity conditions, MSP achieves the optimal minimax rate $(q \log n /n)^{1/2}$ for the upper bound of $l_2$-norm error. Numerical comparisons show that the method works effectively both in model selection and estimation, and the MSP fitted model is stable over a range of small tuning parameter values, eliminating the need to choose the tuning parameter by cross-validation. We also apply MSP to financial data and show that MSP is successful in asset allocation selection.

stat.ME

Sparse Laplacian Shrinkage with the Graphical Lasso Estimator for Regression Problems

This paper considers a high-dimensional linear regression problem where there are complex correlation structures among predictors. We propose a graph-constrained regularization procedure, named Sparse Laplacian Shrinkage with the Graphical Lasso Estimator (SLS-GLE). The procedure uses the estimated precision matrix to describe the specific information on the conditional dependence pattern among predictors, and encourages both sparsity on the regression model and the graphical model. We introduce the Laplacian quadratic penalty adopting the graph information, and give detailed discussions on the advantages of using the precision matrix to construct the Laplacian matrix. Theoretical properties and numerical comparisons are presented to show that the proposed method improves both model interpretability and accuracy of estimation. We also apply this method to a financial problem and prove that the proposed procedure is successful in assets selection.

stat.ME

Smooth Adjustment for Correlated Effects

This paper considers a high dimensional linear regression model with corrected variables. A variety of methods have been developed in recent years, yet it is still challenging to keep accurate estimation when there are complex correlation structures among predictors and the response. We propose an adaptive and "reversed" penalty for regularization to solve this problem. This penalty doesn't shrink variables but focuses on removing the shrinkage bias and encouraging grouping effect. Combining the l_1 penalty and the Minimax Concave Penalty (MCP), we propose two methods called Smooth Adjustment for Correlated Effects (SACE) and Generalized Smooth Adjustment for Correlated Effects (GSACE). Compared with the traditional adaptive estimator, the proposed methods have less influence from the initial estimator and can reduce the false negatives of the initial estimation. The proposed methods can be seen as linear functions of the new penalty's tuning parameter, and are shown to estimate the coefficients accurately in both extremely highly correlated variables situation and weakly correlated variables situation. Under mild regularity conditions we prove that the methods satisfy certain oracle property. We show by simulations and applications that the proposed methods often outperforms other methods.

stat.ME

Penalized regression adjusted causal effect estimates in high dimensional randomized experiments

Regression adjustments are often considered by investigators to improve the estimation efficiency of causal effect in randomized experiments when there exists many pre-experiment covariates. In this paper, we provide conditions that guarantee the penalized regression including the Ridge, Elastic Net and Adapive Lasso adjusted causal effect estimators are asymptotic normal and we show that their asymptotic variances are no greater than that of the simple difference-in-means estimator, as long as the penalized estimators are risk consistent. We also provide conservative estimators for the asymptotic variance which can be used to construct asymptotically conservative confidence intervals for the average causal effect (ACE). Our results are obtained under the Neyman-Rubin potential outcomes model of randomized experiment when the number of covariates is large. Simulation study shows the advantages of the penalized regression adjusted ACE estimators over the difference-in-means estimator.

math.ST

Model Selection Consistency of Lasso for Empirical Data

Large-scale empirical data, the sample size and the dimension are high, often exhibit various characteristics. For example, the noise term follows unknown distributions or the model is very sparse that the number of critical variables is fixed while dimensionality grows with $n$. We consider the model selection problem of lasso for this kind of data. We investigate both theoretical guarantees and simulations, and show that the lasso is robust for various kinds of data.

math.ST

Adaptive elastic net and Separate Selection from Least Squares for ultra-high dimensional regression models

This paper studies the asymptotic properties of the adaptive elastic net in ultra-high dimensional sparse linear regression models and proposes a new method called SSLS (Separate Selection from Least Squares) to improve prediction accuracy. Besides, we prove that SSLS has the superior performance both in the theoretical part and empirical part. In this paper, we prove that the probability of adaptive elastic net selecting wrong variables can decays at an exponential rate with very few conditions. Irrepresentable Condition or similar constraint isn't necessary in our proof. We derive accurate bounds of bias and mean squared error (MSE) which both depend on the choice of parameters, and also show that there exists a bias of asymptotic normality of the adaptive elastic net. Furthermore, simulations and empirical part both show that the prediction accuracy of the penalized least squares requires more improvement. Therefore, we propose SSLS to improve the prediction. It selects variable first, reducing high dimension to low dimension by using the adaptive elastic net in this paper. In the second step, the coefficients are constructed based on the OLS estimation. We show that the bias of SSLS can decays at an exponential rate. Also, MSE decays to zero. Finally, we prove that the variable selection consistency of SSLS implies the asymptotic normality of SSLS. Simulations given in this paper illustrate the performance of the SSLS, adaptive elastic net and other penalized least squares. The index tracking problem in stock market is studied in the empirical part with other methods.

stat.ME