SearcharxivSearch

arXiv subjects

Xuewen Lu

Publications and source records attributed to Xuewen Lu.

7 recordsLinked to original sources

Bayesian Model Averaging under Predictor Redundancy via Density-Ratio Posterior Compression

Bayesian model averaging in support-indexed regression induces a posterior distribution over active predictor supports. Under predictor redundancy, posterior mass can spread across many nearly interchangeable supports, making exact-support summaries unstable or hard to interpret even when prediction is stable. We study how to report an already fitted Bayesian model averaging posterior without changing the Bayesian target. A report uses hard or soft regions of support space, and its compressed reporting law is compared with the reference posterior through an explicit density ratio. This ratio gives computable total-variation and Kullback--Leibler distortion, bounds for bounded predictive summaries, retained-mass diagnostics, and fallback-weight diagnostics. The framework covers fixed hard regions, metric-ball regions, posterior-cluster regions, and pooled-pruned region dictionaries. We prove exact error formulas and validation bounds for these region reports, and give conditions under which a few regions can replace a long list of individual supports. In simulations, our region reports often give shorter and clearer summaries while preserving the main posterior information, and the density-ratio diagnostics show when too much information has been lost.

stat.ML

Deep Neural Networks for Doubly Robust Estimation with Nonprobability Survey Samples

Integrating probability and nonprobability survey samples is an important problem in modern survey sampling. Nonprobability samples often contain rich outcome information but may lack population representativeness, whereas probability samples provide design-based auxiliary information but may not contain the study variable. We propose a deep neural network (DNN)-assisted doubly robust framework for estimating the finite population mean from these two data sources. The proposed method models the logit sampling score for the nonprobability sample as an unknown nonparametric function and estimates it by maximizing a pseudo-likelihood that combines information from the nonprobability sample and a reference probability sample. The DNN parameters are optimized using the ADAM algorithm. The resulting DNN-estimated sampling scores are incorporated into a DNN-assisted inverse-probability weighted estimator and a deep doubly robust estimator. We establish consistency and convergence rates under regularity conditions and evaluate the finite-sample performance of the proposed estimators through simulation studies and an empirical application using Pew Research Center and Behavioral Risk Factor Surveillance System data. The results suggest that the proposed estimators can improve robustness to parametric propensity-score misspecification, especially when the true selection mechanism is nonlinear.

math.ST

Bernstein-von Mises Theorem for Sparse Generalized Linear Model

We study spike-and-slab priors for generalized linear models with possible grouped sparsity. The main result is an oracle Bernstein--von Mises theorem for the fractional posterior under supportwise likelihood assumptions. The proof develops sparse local asymptotic normality and Laplace approximation around support-specific pseudo-true centers, and combines them with fixed-prior mass, support penalization, recovery geometry, and beta-min separation to obtain contraction, support recovery, Gaussian mixture approximation, and collapse to the oracle Gaussian law. Model-entry verifications are given for Gaussian regression and for logistic, Poisson, probit, Gamma log-link, and negative-binomial log-link regression under stated sufficient conditions. The ordinary posterior is treated only through restricted Gaussian and canonical-link extensions, with coverage under additional active-dimension and moment conditions.

math.ST

Variable Selection with Broken Adaptive Ridge Regression for Interval-Censored Competing Risks Data

Competing risks data refer to situations where the occurrence of one event pre- cludes the possibility of other events happening, resulting in multiple mutually exclusive events. This data type is commonly encountered in medical research and clinical trials, exploring the interplay between different events and informing decision-making in fields such as healthcare and epidemiology. We develop a penal- ized variable selection procedure to handle such complex data in an interval-censored setting. We consider a broad class of semiparametric transformation regression mod- els, including popular models such as proportional and non-proportional hazards models. To promote sparsity and select variables specific to each event, we employ the broken adaptive ridge (BAR) penalty. This approach allows us to simultane- ously select important risk factors and estimate their effects for each event under investigation. We establish the oracle property of the BAR procedure and evaluate its performance through simulation studies. The proposed method is applied to a real-life HIV cohort dataset, further validating its applicability in practice.

stat.ME

Variable selection in the joint frailty model of recurrent and terminal events using Broken Adaptive Ridge regression

We introduce a novel method to simultaneously perform variable selection and estimation in the joint frailty model of recurrent and terminal events using the Broken Adaptive Ridge Regression penalty. The BAR penalty can be summarized as an iteratively reweighted squared $L_2$-penalized regression, which approximates the $L_0$-regularization method. Our method allows for the number of covariates to diverge with the sample size. Under certain regularity conditions, we prove that the BAR estimator implemented under the model framework is consistent and asymptotically normally distributed, which are known as the oracle properties in the variable selection literature. In our simulation studies, we compare our proposed method to the Minimum Information Criterion (MIC) method. We apply our method on the Medical Information Mart for Intensive Care (MIMIC-III) database, with the aim of investigating which variables affect the risks of repeated ICU admissions and death during ICU stay.

stat.ME

Broken Adaptive Ridge Method for Variable Selection in Generalized Partly Linear Models with Application to the Coronary Artery Disease Data

Motivated by the CATHGEN data, we develop a new statistical learning method for simultaneous variable selection and parameter estimation under the context of generalized partly linear models for data with high-dimensional covariates. The method is referred to as the broken adaptive ridge (BAR) estimator, which is an approximation of the $L_0$-penalized regression by iteratively performing reweighted squared $L_2$-penalized regression. The generalized partly linear model extends the generalized linear model by including a non-parametric component to construct a flexible model for modeling various types of covariate effects. We employ the Bernstein polynomials as the sieve space to approximate the non-parametric functions so that our method can be implemented easily using the existing R packages. Extensive simulation studies suggest that the proposed method performs better than other commonly used penalty-based variable selection methods. We apply the method to the CATHGEN data with a binary response from a coronary artery disease study, which motivated our research, and obtained new findings in both high-dimensional genetic and low-dimensional non-genetic covariates.

stat.ME

Penalized Variable Selection with Broken Adaptive Ridge Regression for Semi-competing Risks Data

Semi-competing risks data arise when both non-terminal and terminal events are considered in a model. Such data with multiple events of interest are frequently encountered in medical research and clinical trials. In this framework, terminal event can censor the non-terminal event but not vice versa. It is known that variable selection is practical in identifying significant risk factors in high-dimensional data. While some recent works on penalized variable selection deal with these competing risks separately without incorporating possible correlation between them, we perform variable selection in an illness-death model using shared frailty where semiparametric hazard regression models are used to model the effect of covariates. We propose a broken adaptive ridge (BAR) penalty to encourage sparsity and conduct extensive simulation studies to compare its performance with other popular methods. We perform variable selection in an event specific manner so that the potential risk factors and covariates effects can be estimated and selected, simultaneously corresponding to each event in the study. The grouping effect, as well as the oracle property of the proposed BAR procedure are investigated using simulation studies. The proposed method is then applied to real-life data arising from a Colon Cancer study.

stat.ME