SearcharxivSearch

arXiv subjects

Lucy Shao

Publications and source records attributed to Lucy Shao.

3 recordsLinked to original sources

Win-Ratio Regression for Prioritized Composite Outcomes in Observational Studies: Doubly Robust and Efficient Estimation with Future-Score Correction

Prioritized pairwise outcomes are useful when clinical events follow a natural hierarchy, but censoring before pair resolution complicates estimation. We develop a win-ratio regression framework for this setting by defining a complete-data target over follow-up and deriving an estimating equation for the observed data. The central idea is future-score correction (FC): when censoring prevents later pairwise comparisons from being observed, the method replaces the remaining score with its conditional expectation given the observed history. This correction recovers pairwise information beyond that provided by inverse censoring weights alone. Additionally, we incorporate treatment weighting and baseline outcome augmentation to address baseline confounding. Together, these components yield double robustness for treatment assignment and censoring. Inference is obtained from U-statistic theory. Under standard regularity conditions, the AIPW-FC estimator is asymptotically normal and efficient when all nuisance functions are correctly specified. Simulations with 30%, 50%, and 65% censoring show that efficiency gains from future-score correction increase with the censoring rate, with relative efficiency reaching 1.50 under 65% censoring and near-nominal coverage for AIPW-FC. An application to OneFlorida electronic health record data illustrates the method for a composite outcome that prioritizes death over hospitalization.

stat.ME

Why Is the Double-Robust Estimator for Causal Inference Not Doubly Robust for Variance Estimation?

Doubly robust estimators (DRE) are widely used in causal inference because they yield consistent estimators of average causal effect when at least one of the nuisance models, the propensity for treatment (exposure) or the outcome regression, is correct. However, double robustness does not extend to variance estimation; the influence-function (IF)-based variance estimator is consistent only when both nuisance parameters are correct. This raises concerns about applying DRE in practice, where model misspecification is inevitable. The recent paper by Shook-Sa et al. (2025, Biometrics, 81(2), ujaf054) demonstrated through Monte Carlo simulations that the IF-based variance estimator is biased. However, the paper's findings are empirical. The key question remains: why does the variance estimator fail in double robustness, and under what conditions do alternatives succeed, such as the ones demonstrated in Shook-Sa et al. 2025. In this paper, we develop a formal theory to clarify the efficiency properties of DRE that underlie these empirical findings. We also introduce alternative strategies, including a mixture-based framework underlying the sample-splitting and crossfitting approaches, to achieve valid inference with misspecified nuisance parameters. Our considerations are illustrated with simulation and real study data.

stat.ME

Semiparametric Regression Models for Explanatory Variables with Missing Data due to Detection Limit

Detection limit (DL) has become an increasingly ubiquitous issue in statistical analyses of biomedical studies, such as cytokine, metabolite and protein analysis. In regression analysis, if an explanatory variable is left-censored due to concentrations below the DL, one may limit analyses to observed data. In many studies, additional, or surrogate, variables are available to model, and incorporating such auxiliary modeling information into the regression model can improve statistical power. Although methods have been developed along this line, almost all are limited to parametric models for both the regression and left-censored explanatory variable. While some recent work has considered semiparametric regression for the censored DL-effected explanatory variable, the regression of primary interest is still left parametric, which not only makes it prone to biased estimates, but also suffers from high computational cost and inefficiency due to maximizing an extremely complex likelihood function and bootstrap inference. In this paper, we propose a new approach by considering semiparametric generalized linear models (SPGLM) for the primary regression and parametric or semiparametric models for DL-effected explanatory variable. The semiparametric and semiparametric combination provides the most robust inference, while the semiparametric and parametric case enables more efficient inference. The proposed approach is also much easier to implement and allows for leveraging sample splitting and cross fitting (SSCF) to improve computational efficiency in variance estimation. In particular, our approach improves computational efficiency over bootstrap by 450 times. We use simulated and real study data to illustrate the approach.

stat.ME