Searcharxiv⌕ Search

arXiv subjects

Hu Yang

Publications and source records attributed to Hu Yang.

26 records · Page 2Linked to original sources

Interaction Pursuit Biconvex Optimization

Multivariate regression models are widely used in various fields such as biology and finance. In this paper, we focus on two key challenges: (a) When should we favor a multivariate model over a series of univariate models; (b) If the numbers of responses and predictors are allowed to greatly exceed the sample size, how to reduce the computational cost and provide precise estimation. The proposed method, Interaction Pursuit Biconvex Optimization (IPBO), explores the regression relationship allowing the predictors and responses derived from different multivariate normal distributions with general covariance matrices. In practice, the correlation structures within are complex and interact on each other based on the regression function. The proposed method solves this problem by building a structured sparsity penalty to encourages the shared structure between the network and the regression coefficients. We prove theoretical results under interpretable conditions, and provide an efficient algorithm to compute the estimator. Simulation studies and real data examples compare the proposed method with several existing methods, indicating that IPBO works well.

stat.ME↗

Sparse Laplacian Shrinkage with the Graphical Lasso Estimator for Regression Problems

This paper considers a high-dimensional linear regression problem where there are complex correlation structures among predictors. We propose a graph-constrained regularization procedure, named Sparse Laplacian Shrinkage with the Graphical Lasso Estimator (SLS-GLE). The procedure uses the estimated precision matrix to describe the specific information on the conditional dependence pattern among predictors, and encourages both sparsity on the regression model and the graphical model. We introduce the Laplacian quadratic penalty adopting the graph information, and give detailed discussions on the advantages of using the precision matrix to construct the Laplacian matrix. Theoretical properties and numerical comparisons are presented to show that the proposed method improves both model interpretability and accuracy of estimation. We also apply this method to a financial problem and prove that the proposed procedure is successful in assets selection.

stat.ME↗

On the penalized maximum likelihood estimation of high-dimensional approximate factor model

In this paper, we mainly focus on the penalized maximum likelihood estimation (MLE) of the high-dimensional approximate factor model. Since the current estimation procedure can not guarantee the positive definiteness of the error covariance matrix, by reformulating the estimation of error covariance matrix and based on the lagrangian duality, we propose an accelerated proximal gradient (APG) algorithm to give a positive definite estimate of the error covariance matrix. Combined the APG algorithm with EM method, a new estimation procedure is proposed to estimate the high-dimensional approximate factor model. The new method not only gives positive definite estimate of error covariance matrix but also improves the efficiency of estimation for the high-dimensional approximate factor model. Although the proposed algorithm can not guarantee a global unique solution, it enjoys a desirable non-increasing property. The efficiency of the new algorithm on estimation and forecasting is also investigated via simulation and real data analysis.

stat.CO↗

Smooth Adjustment for Correlated Effects

This paper considers a high dimensional linear regression model with corrected variables. A variety of methods have been developed in recent years, yet it is still challenging to keep accurate estimation when there are complex correlation structures among predictors and the response. We propose an adaptive and "reversed" penalty for regularization to solve this problem. This penalty doesn't shrink variables but focuses on removing the shrinkage bias and encouraging grouping effect. Combining the l_1 penalty and the Minimax Concave Penalty (MCP), we propose two methods called Smooth Adjustment for Correlated Effects (SACE) and Generalized Smooth Adjustment for Correlated Effects (GSACE). Compared with the traditional adaptive estimator, the proposed methods have less influence from the initial estimator and can reduce the false negatives of the initial estimation. The proposed methods can be seen as linear functions of the new penalty's tuning parameter, and are shown to estimate the coefficients accurately in both extremely highly correlated variables situation and weakly correlated variables situation. Under mild regularity conditions we prove that the methods satisfy certain oracle property. We show by simulations and applications that the proposed methods often outperforms other methods.

stat.ME↗

Semiparametric model averaging for high dimensional conditional quantile prediction

In this article, we propose a penalized high dimensional semiparametric model average quantile prediction approach that is robust for forecasting the conditional quantile of the response. We consider a two-step estimation procedure. In the first step, we use a local linear regression approach to estimate the individual marginal quantile functions, and approximate the conditional quantile of the response by an affine combination of one-dimensional marginal quantile regression functions. In the second step, based on the nonparametric kernel estimates of the marginal quantile regression functions, we utilize a penalized method to estimate the suitable model weights vector involved in the approximation. The objective of the second step is to select significant variables whose marginal quantile functions make a significant contribution to estimating the joint multivariate conditional quantile function. Under some mild conditions, we have established the asymptotic properties of the proposed robust estimator. Finally, simulations and a real data analysis have been used to illustrate the proposed method.

math.ST↗

Model Selection Consistency of Lasso for Empirical Data

Large-scale empirical data, the sample size and the dimension are high, often exhibit various characteristics. For example, the noise term follows unknown distributions or the model is very sparse that the number of critical variables is fixed while dimensionality grows with $n$. We consider the model selection problem of lasso for this kind of data. We investigate both theoretical guarantees and simulations, and show that the lasso is robust for various kinds of data.

math.ST↗

A note on the condition number of the scaled total least squares problem

In this paper, we consider the explicit expressions of the normwise condition number for the scaled total least squares problem. Some techniques are introduced to simplify the expression of the condition number, and some new results are derived. Based on these new results, new expressions of the condition number for the total least squares problem can be deduced as a special case. New forms of the condition number enjoy some storage and computational advantages. We also proposed three different methods to estimate the condition number. Some numerical experiments are carried out to illustrate the effectiveness of our results.

math.NA↗

Adaptive elastic net and Separate Selection from Least Squares for ultra-high dimensional regression models

This paper studies the asymptotic properties of the adaptive elastic net in ultra-high dimensional sparse linear regression models and proposes a new method called SSLS (Separate Selection from Least Squares) to improve prediction accuracy. Besides, we prove that SSLS has the superior performance both in the theoretical part and empirical part. In this paper, we prove that the probability of adaptive elastic net selecting wrong variables can decays at an exponential rate with very few conditions. Irrepresentable Condition or similar constraint isn't necessary in our proof. We derive accurate bounds of bias and mean squared error (MSE) which both depend on the choice of parameters, and also show that there exists a bias of asymptotic normality of the adaptive elastic net. Furthermore, simulations and empirical part both show that the prediction accuracy of the penalized least squares requires more improvement. Therefore, we propose SSLS to improve the prediction. It selects variable first, reducing high dimension to low dimension by using the adaptive elastic net in this paper. In the second step, the coefficients are constructed based on the OLS estimation. We show that the bias of SSLS can decays at an exponential rate. Also, MSE decays to zero. Finally, we prove that the variable selection consistency of SSLS implies the asymptotic normality of SSLS. Simulations given in this paper illustrate the performance of the SSLS, adaptive elastic net and other penalized least squares. The index tracking problem in stock market is studied in the empirical part with other methods.

stat.ME↗