SearcharxivSearch

arXiv subjects

Shangyuan Ye

Publications and source records attributed to Shangyuan Ye.

5 recordsLinked to original sources

Cluster-induced target shift and synthetic approximation for high-dimensional clustered data

In high-dimensional clustered data, covariates whose distributions differ across clusters can act as proxies for unobserved cluster effects. We show that the marginal-model LASSO implicitly uses sparse combinations of such covariates to shift the estimation target from structural coefficients to a contaminated vector, which inflates false selections. We therefore propose the Synthetic Heterogeneous-Effects LASSO (SHEL), which augments the regression design with separately penalized, outcome-independent, cluster-constant synthetic covariates to correct the cluster-induced target shift. Under a fixed nuisance-effects formulation, we derive oracle properties and establish consistency of the structural coefficients when the residual heterogeneity is weakly aligned with the augmented design. We also demonstrate the distinction between the structural parameter under small residual heterogeneity and the population working parameter when residual heterogeneity persists. Further, we develop a debiased estimator with a cluster-robust sandwich variance, and establish asymptotic results. Simulations show far fewer false selections than the marginal LASSO, and an analysis of longitudinal neutrophil transcriptomic data from hospitalized COVID-19 patients yields a more parsimonious gene set that retains the severity markers of the original study.

stat.ME

Design-Assisted Regression

We consider regression problems in which the marginal distribution of the covariates is informative for estimation and variable selection, rather than merely auxiliary. Motivated by random-design, high-dimensional, and latent-effect settings, we propose a general design-assisted regression framework in which the estimating criterion depends on both the conditional model for $Y \mid \bfX$ and structured features of the covariate distribution. The framework identifies two roles of design information: stabilizing weak design directions through quadratic regularization and correcting latent-effect distortion through nuisance augmentation. We establish oracle properties for the resulting estimator, separate the effects of stochastic error, shrinkage, and approximation, and compare it with a benchmark sparse procedure that ignores design information. These results show that the proposed framework improves estimation while preserving first-order prediction performance. Numerical studies and two real-data applications illustrate the practical impact of incorporating design information.

stat.ME

High-dimensional Statistical Inference and Variable Selection Using Sufficient Dimension Association

Simultaneous variable selection and statistical inference is challenging in high-dimensional data analysis. Most existing post-selection inference methods require explicitly specified regression models, which are often linear, as well as sparsity in the regression model. The performance of such procedures can be poor under either misspecified nonlinear models or a violation of the sparsity assumption. In this paper, we propose a sufficient dimension association (SDA) technique that measures the association between each predictor and the response variable conditioning on other predictors in the high-dimensional setting. Our proposed SDA method requires neither a specific form of regression model nor sparsity in the regression. Alternatively, our method assumes normalized or Gaussian-distributed predictors with a Markov blanket property. We propose an estimator for the SDA and prove asymptotic properties for the estimator. We construct three types of test statistics for the SDA and propose a multiple testing procedure to control the false discovery rate. Extensive simulation studies have been conducted to show the validity and superiority of our SDA method. Gene expression data from the Alzheimer Disease Neuroimaging Initiative are used to demonstrate a real application.

stat.ME

A marginalized three-part interrupted time series regression model for proportional data

Interrupted time series (ITS) is often used to evaluate the effectiveness of a health policy intervention that accounts for the temporal dependence of outcomes. When the outcome of interest is a percentage or percentile, the data can be highly skewed, bounded in $[0, 1]$, and have many zeros or ones. A three-part Beta regression model is commonly used to separate zeros, ones, and positive values explicitly by three submodels. However, incorporating temporal dependence into the three-part Beta regression model is challenging. In this article, we propose a marginalized zero-one-inflated Beta time series model that captures the temporal dependence of outcomes through copula and allows investigators to examine covariate effects on the marginal mean. We investigate its practical performance using simulation studies and apply the model to a real ITS study.

stat.ME

Orthogonal Series Density Estimation for Complex Surveys

We propose an orthogonal series density estimator for complex surveys, where samples are neither independent nor identically distributed. The proposed estimator is proved to be design-unbiased and asymptotically design-consistent. The asymptotic normality is proved under both design and combined spaces. Two data driven estimators are proposed based on the proposed oracle estimator. We show the efficiency of the proposed estimators in simulation studies. A real survey data example is provided for an illustration.

stat.ME