SearcharxivSearch

arXiv subjects

Huali Zhao

Publications and source records attributed to Huali Zhao.

3 recordsLinked to original sources

Simulation-free extrapolation for misspecified models induced by categorizing an error-prone continuous covariate

Epidemiological studies often categorize continuous exposures for interpretation even when the underlying outcome-exposure association is continuous. The fitted categorical regression is a misspecified model because it replaces the continuous exposure with categories. With measurement error, categorization also misclassifies latent exposure categories, so the observed regression generally targets different means and contrasts. Existing estimating-equation and simulation-based extrapolation approaches require, respectively, outcome-model-specific derivations and pseudo-data generation with repeated fitting. We introduce simulation-free extrapolation (SIMFEX), which estimates misclassification probabilities and latent category proportions from replicates, computes the mean-scale trajectory without pseudo-data or repeated outcome-model fitting, and extrapolates it to the no-misclassification endpoint. Without additional error-free covariates, this construction applies across known one-to-one links, and the resulting estimator is consistent under stated conditions. With error-free covariates, the contrast relation remains exact for identity-link additive models when misclassification probabilities and category proportions are covariate-invariant; the nonidentity-link version provides a practical approximation. Simulations show substantial bias reduction relative to the naive analysis and coverage generally close to nominal. In the UK Biobank analysis, SIMFEX produced larger estimated high-versus-low fat-intake contrasts than the naive analysis for body mass index and obesity, illustrating how category misclassification can change the magnitude and uncertainty of prespecified contrasts.

stat.ME

Augmented transfer regression learning for completely missing covariates

Large-scale population-level datasets, such as the UK Biobank and the All of Us Research Program, often lack covariates needed for a specific analysis, such as genetic or lifestyle measures, while related studies measure them. This creates a cross-population missing data problem in which covariates are completely unobserved in the target population, rather than partially missing within one dataset. We propose an augmented transfer regression learning method for this setting. The key identifying condition is a sub-population shift assumption: the joint distribution of the outcome and observed covariates may differ across source and target populations, but the conditional distribution of the missing covariates given observed variables is invariant. We combine importance-weighted estimating equations with imputation terms for first- and second-order moments of the missing covariates. The resulting estimator is doubly robust and remains consistent if either the density ratio model or both imputation models are correctly specified. It is n^{1/2}-consistent and asymptotically normal, and attains the semiparametric efficiency bound when both

stat.ME

Debiased high-dimensional regression calibration for errors-in-variables log-contrast models

Motivated by the challenges in analyzing gut microbiome and metagenomic data, this work aims to tackle the issue of measurement errors in high-dimensional regression models that involve compositional covariates. This paper marks a pioneering effort in conducting statistical inference on high-dimensional compositional data affected by mismeasured or contaminated data. We introduce a calibration approach tailored for the linear log-contrast model. Under relatively lenient conditions regarding the sparsity level of the parameter, we have established the asymptotic normality of the estimator for inference. Numerical experiments and an application in microbiome study have demonstrated the efficacy of our high-dimensional calibration strategy in minimizing bias and achieving the expected coverage rates for confidence intervals. Moreover, the potential application of our proposed methodology extends well beyond compositional data, suggesting its adaptability for a wide range of research contexts.

stat.ME