Cluster-induced target shift and synthetic approximation for high-dimensional clustered data
In high-dimensional clustered data, covariates whose distributions differ across clusters can act as proxies for unobserved cluster effects. We show that the marginal-model LASSO implicitly uses sparse combinations of such covariates to shift the estimation target from structural coefficients to a contaminated vector, which inflates false selections. We therefore propose the Synthetic Heterogeneous-Effects LASSO (SHEL), which augments the regression design with separately penalized, outcome-independent, cluster-constant synthetic covariates to correct the cluster-induced target shift. Under a fixed nuisance-effects formulation, we derive oracle properties and establish consistency of the structural coefficients when the residual heterogeneity is weakly aligned with the augmented design. We also demonstrate the distinction between the structural parameter under small residual heterogeneity and the population working parameter when residual heterogeneity persists. Further, we develop a debiased estimator with a cluster-robust sandwich variance, and establish asymptotic results. Simulations show far fewer false selections than the marginal LASSO, and an analysis of longitudinal neutrophil transcriptomic data from hospitalized COVID-19 patients yields a more parsimonious gene set that retains the severity markers of the original study.