SearcharxivSearch

arXiv subjects

Linh Nghiem

Publications and source records attributed to Linh Nghiem.

9 recordsLinked to original sources

Screening methods for linear errors-in-variables models in high dimensions

Microarray studies, in order to identify genes associated with an outcome of interest, usually produce noisy measurements for a large number of gene expression features from a small number of subjects. One common approach to analyzing such high-dimensional data is to use linear errors-in-variables models; however, current methods for fitting such models are computationally expensive. In this paper, we present two efficient screening procedures, namely corrected penalized marginal screening and corrected sure independence screening, to reduce the number of variables for final model building. Both screening procedures are based on fitting corrected marginal regression models relating the outcome to each contaminated covariate separately, which can be computed efficiently even with a large number of features. Under mild conditions, we show that these procedures achieve screening consistency and reduce the number of features considerably, even when the number of covariates grows exponentially with the sample size. Additionally, if the true covariates are weakly correlated, corrected penalized marginal screening can achieve full variable selection consistency. Through simulation studies and an analysis of gene expression data for bone mineral density of Norwegian women, we demonstrate that the two new screening procedures make estimation of linear errors-in-variables models computationally scalable in high dimensional settings, and improve finite sample estimation and selection performance compared with estimators that do not employ a screening stage.

stat.ME

Sparse Sliced Inverse Regression via Cholesky Matrix Penalization

We introduce a new sparse sliced inverse regression estimator called Cholesky matrix penalization and its adaptive version for achieving sparsity in estimating the dimensions of the central subspace. The new estimators use the Cholesky decomposition of the covariance matrix of the covariates and include a regularization term in the objective function to achieve sparsity in a computationally efficient manner. We establish the theoretical values of the tuning parameters that achieve estimation and variable selection consistency for the central subspace. Furthermore, we propose a new projection information criterion to select the tuning parameter for our proposed estimators and prove that the new criterion facilitates selection consistency. The Cholesky matrix penalization estimator inherits the strength of the Matrix Lasso and the Lasso sliced inverse regression estimator; it has superior performance in numerical studies and can be adapted to other sufficient dimension methods in the literature.

stat.ME

Bayesian Regularization of Gaussian Graphical Models with Measurement Error

We consider a framework for determining and estimating the conditional pairwise relationships of variables when the observed samples are contaminated with measurement error in high dimensional settings. Assuming the true underlying variables follow a multivariate Gaussian distribution, if no measurement error is present, this problem is often solved by estimating the precision matrix under sparsity constraints. However, when measurement error is present, not correcting for it leads to inconsistent estimates of the precision matrix and poor identification of relationships. We propose a new Bayesian methodology to correct for the measurement error from the observed samples. This Bayesian procedure utilizes a recent variant of the spike-and-slab Lasso to obtain a point estimate of the precision matrix, and corrects for the contamination via the recently proposed Imputation-Regularization Optimization procedure designed for missing data. Our method is shown to perform better than the naive method that ignores measurement error in both identification and estimation accuracy. To show the utility of the method, we apply the new method to establish a conditional gene network from a microarray dataset.

stat.ME

Estimation in linear errors-in-variables models with unknown error distribution

Parameter estimation in linear errors-in-variables models typically requires that the measurement error distribution be known (or estimable from replicate data). A generalized method of moments approach can be used to estimate model parameters in the absence of knowledge of the error distributions, but requires the existence of a large number of model moments. In this paper, parameter estimation based on the phase function, a normalized version of the characteristic function, is considered. This approach requires the model covariates to have asymmetric distributions, while the error distributions are symmetric. Parameter estimation is then based on minimizing a distance function between the empirical phase functions of the noisy covariates and the outcome variable. No knowledge of the measurement error distribution is required to calculate this estimator. Both the asymptotic and finite sample properties of the estimator are considered. The connection between the phase function approach and method of moments is also discussed. The estimation of standard errors is also considered and a modified bootstrap algorithm is proposed for fast computation. The newly proposed estimator is competitive when compared to generalized method of moments, even while making fewer model assumptions on the measurement error. Finally, the proposed method is applied to a real dataset concerning the measurement of air pollution.

stat.ME

Lagged Exact Bayesian Online Changepoint Detection with Parameter Estimation

Identifying changes in the generative process of sequential data, known as changepoint detection, has become an increasingly important topic for a wide variety of fields. A recently developed approach, which we call EXact Online Bayesian Changepoint Detection (EXO), has shown reasonable results with efficient computation for real time updates. The method is based on a \textit{forward} recursive message-passing algorithm. However, the detected changepoints from these methods are unstable. We propose a new algorithm called Lagged EXact Online Bayesian Changepoint Detection (LEXO) that improves the accuracy and stability of the detection by incorporating $\ell$-time lags to the inference. The new algorithm adds a recursive \textit{backward} step to the forward EXO and has computational complexity linear in the number of added lags. Estimation of parameters associated with regimes is also developed. Simulation studies with three common changepoint models show that the detected changepoints from LEXO are much more stable and parameter estimates from LEXO have considerably lower MSE than EXO. We illustrate applicability of the methods with two real world data examples comparing the EXO and LEXO.

stat.ML

Simulation-Selection-Extrapolation: Estimation in High-Dimensional Errors-in-Variables Models

This paper considers errors-in-variables models in a high-dimensional setting where the number of covariates can be much larger than the sample size, and there are only a small number of non-zero covariates. The presence of measurement error in the covariates can result in severely biased parameter estimates, and also affects the ability of penalized methods such as the lasso to recover the true sparsity pattern. A new estimation procedure called SIMSELEX (SIMulation-SELection-EXtrapolation) is proposed. This procedure augments the traditional SIMEX approach with a variable selection step based on the group lasso. The SIMSELEX estimator is shown to perform well in variable selection, and has significantly lower estimation error than naive estimators that ignore measurement error. SIMSELEX can be applied in a variety of errors-in-variables settings, including linear models, generalized linear models, and Cox survival models. It is furthermore shown how SIMSELEX can be applied to spline-based regression models. SIMSELEX estimators are compared to the corrected lasso and the conic programming estimator for a linear model, and to the conditional scores lasso for a logistic regression model. Finally, the method is used to analyze a microarray dataset that contains gene expression measurements of favorable histology Wilms tumors.

stat.ME

Phase Function Density Deconvolution with Heteroscedastic Measurement Error of Unknown Type

It is important to properly correct for measurement error when estimating density functions associated with biomedical variables. These estimators that adjust for measurement error are broadly referred to as density deconvolution estimators. While most methods in the literature assume the distribution of the measurement error to be fully known, a recently proposed method based on the empirical phase function (EPF) can deal with the situation when the measurement error distribution is unknown. The EPF density estimator has only been considered in the context of additive and homoscedastic measurement error; however, the measurement error of many biomedical variables is heteroscedastic in nature. In this paper, we developed a phase function approach for density deconvolution when the measurement error has unknown distribution and is heteroscedastic. A weighted empirical phase function (WEPF) is proposed where the weights are used to adjust for heteroscedasticity of measurement error. The asymptotic properties of the WEPF estimator are evaluated. Simulation results show that the weighting can result in large decreases in mean integrated squared error (MISE) when estimating the phase function. The estimation of the weights from replicate observations is also discussed. Finally, the construction of a deconvolution density estimator using the WEPF is compared to an existing deconvolution estimator that adjusts for heteroscedasticity, but assumes the measurement error distribution to be fully known. The WEPF estimator proves to be competitive, especially when considering that it relies on the minimal assumption of the distribution of measurement error.

stat.ME

A Heuristic Method for Scheduling Band Concert Tours

Scheduling band concert tours is an important and challenging task faced by many band management companies and producers. A band has to perform in various cities over a period of time, and the specific route they follow is subject to numerous constraints, such as: venue availability, travel limits, and required rest periods. A good tour must consider several objectives regarding the desirability of certain days of the week, as well as travel cost. We developed and implemented a heuristic algorithm in Java, which was based on simulated annealing, to automatically generate good tours that both satisfied the above constraints and improved objectives significantly when compared to the best manual tour created by the client. Our program also enabled the client to see and explore trade-offs among objectives while choosing the best tour that meets the requirements of the business.

cs.CY

Risk-return relationship: An empirical study of different statistical methods for estimating the Capital Asset Pricing Models (CAPM) and the Fama-French model for large cap stocks

The Capital Asset Pricing Model (CAPM) is one of the original models in explaining risk-return relationship in the financial market. However, when applying the CAPM into reality, it demonstrates a lot of shortcomings. While improving the performance of the model, many studies, on one hand, have attempted to apply different statistical methods to estimate the model, on the other hand, have added more predictors to the model. First, the thesis focuses on reviewing the CAPM and comparing popular statistical methods used to estimate it, and then, the thesis compares predictive power of the CAPM and the Fama-French model, which is an important extension of the CAPM. Through an empirical study on the data set of large cap stocks, we have demonstrated that there is no statistical method that would recover the expected relationship between systematic risk (represented by beta) and return from the CAPM, and that the Fama-French model does not have a better predictive performance than the CAPM on individual stocks. Therefore, the thesis provides more evidence to support the incorrectness of the CAPM and the limitation of the Fama-French model in explaining risk-return relationship.

q-fin.ST