SearcharxivSearch

arXiv subjects

Ruohui Chen

Publications and source records attributed to Ruohui Chen.

4 recordsLinked to original sources

Clustering Informed Inverse Probability Weighting Strategies for Causal Effect Estimation in Observational Studies

Inverse probability weighting (IPW) is widely used to estimate causal effects in observational studies but depends on adequate propensity-score specification. We compare three strategies for addressing treatment assignment heterogeneity: standard IPW, clustering augmented IPW with cluster specific propensity score models, and a global propensity score model including estimated cluster membership as a covariate. Through simulations with and without latent cluster structure and under correctly specified and omitted covariate propensity score models, we evaluate bias, mean squared error (MSE), and confidence interval coverage across sample sizes of 100 to 500. Both cluster informed strategies reduced bias and MSE from omitted covariate misspecification relative to standard IPW, but neither uniformly dominated: clustering augmented IPW achieved lower MSE when latent cluster structure was present, whereas the global model generally provided lower bias and better coverage at smaller sample sizes. We also apply the methods to 966 breast cancer patients treated with carboplatin, using generalized propensity scores to estimate the dose response relationship between treatment cycles and hypersensitivity reaction risk. Standard and clustered analyses produced similar pooled estimates, while clustering additionally provided subgroup specific estimates and diagnostic profiles. Overall, cluster informed strategies may improve robustness to propensity score misspecification, with relative performance depending on subgroup structure, sample size, and inferential priorities.

stat.ME

Semiparametric Regression Models for Explanatory Variables with Missing Data due to Detection Limit

Detection limit (DL) has become an increasingly ubiquitous issue in statistical analyses of biomedical studies, such as cytokine, metabolite and protein analysis. In regression analysis, if an explanatory variable is left-censored due to concentrations below the DL, one may limit analyses to observed data. In many studies, additional, or surrogate, variables are available to model, and incorporating such auxiliary modeling information into the regression model can improve statistical power. Although methods have been developed along this line, almost all are limited to parametric models for both the regression and left-censored explanatory variable. While some recent work has considered semiparametric regression for the censored DL-effected explanatory variable, the regression of primary interest is still left parametric, which not only makes it prone to biased estimates, but also suffers from high computational cost and inefficiency due to maximizing an extremely complex likelihood function and bootstrap inference. In this paper, we propose a new approach by considering semiparametric generalized linear models (SPGLM) for the primary regression and parametric or semiparametric models for DL-effected explanatory variable. The semiparametric and semiparametric combination provides the most robust inference, while the semiparametric and parametric case enables more efficient inference. The proposed approach is also much easier to implement and allows for leveraging sample splitting and cross fitting (SSCF) to improve computational efficiency in variance estimation. In particular, our approach improves computational efficiency over bootstrap by 450 times. We use simulated and real study data to illustrate the approach.

stat.ME

A doubly robust estimator for the Mann Whitney Wilcoxon Rank Sum Test when applied for causal inference in observational studies

The Mann-Whitney-Wilcoxon rank sum test (MWWRST) is a widely used method for comparing two treatment groups in randomized control trials, particularly when dealing with highly skewed data. However, when applied to observational study data, the MWWRST often yields invalid results for causal inference. To address this limitation, Wu et al. (2014) introduced an approach that incorporates inverse probability weighting (IPW) into this rank-based statistics to mitigate confounding effects. Subsequently, Mao (2018), Zhang et al. (2019), and Ai et al. (2020) extended this IPW estimator to develop doubly robust estimators. Nevertheless, each of these approaches has notable limitations. Mao's method imposes stringent assumptions that may not align with real-world study data. Zhang et al.'s (2019) estimators rely on bootstrap inference, which suffers from computational inefficiency and lacks known asymptotic properties. Meanwhile, Ai et al. (2020) primarily focus on testing the null hypothesis of equal distributions between two groups, which is a more stringent assumption that may not be well-suited to the primary practical application of MWWRST. In this paper, we aim to address these limitations by leveraging functional response models (FRM) to develop doubly robust estimators. We demonstrate the performance of our proposed approach using both simulated and real study data.

stat.ME

A linear mixed model approach for measurement error adjustment: applications to sedentary behavior assessment from wearable devices

In recent years, wearable devices have become more common to capture a wide range of health behaviors, especially for physical activity and sedentary behavior. These sensor-based measures are deemed to be objective and thus less prone to self-reported biases, inherent in questionnaire assessments. While this is undoubtedly a major advantage, there can still be measurement errors from the device recordings, which pose serious challenges for conducting statistical analysis and obtaining unbiased risk estimates. There is a vast literature proposing statistical methods for adjusting for measurement errors in self-reported behaviors, such as in dietary intake. However, there is much less research on error correction for sensor-based device measures, especially sedentary behavior. In this paper, we address this gap. Exploiting the excessive multiple-day assessments typically collected when sensor devices are deployed, we propose a two-stage linear mixed effect model (LME) based approach to correct bias caused by measurement errors. We provide theoretical proof of the debiasing process using the Best Linear Unbiased Predictors (BLUP), and use both simulation and real data from a cohort study to demonstrate the performance of the proposed approach while comparing to the na\"ive plug-in approach that directly uses device measures without appropriately adjusting measurement errors. Our results indicate that employing our easy-to-implement BLUP correction method can greatly reduce biases in disease risk estimates and thus enhance the validity of study findings.

stat.ME