SearcharxivSearch

arXiv subjects

Xin M. Tu

Publications and source records attributed to Xin M. Tu.

5 recordsLinked to original sources

Semiparametric Estimation of Delayed-Outcome Treatment Effects Using Short-Term Surrogates under Administrative Censoring

The multi-site registry studies, such as Stepped-wedge cluster-randomized trials (SW-CRT), staggered-enrollment RCTs, etc., share a structural feature: the primary long-term outcome is administratively censored for a non-negligible fraction of units, with censoring driven by calendar design rather than by the outcome itself. Standard inverse-probability-of-censoring weighting becomes unstable when observation probabilities $g_Δ$ concentrate near zero for late-crossing units, while parametric mixed-model analyses discard the information in any short-term intermediate measurement and rely on correct specification of the secular time trend. We study semiparametric estimation of the average treatment effect when a short-term surrogate, which is observed for all units and conditionally independent of the censoring mechanism given baseline covariates, is available. Identification takes a nested-integral form in which the outcome regression is marginalized over the conditional surrogate distribution, so the observation mechanism does not enter the target functional as an inverse weight. We show that a density-plug-in one-step debiased machine-learning construction for this functional leaves a second-order cross-product remainder $R_{SY}$ that has no doubly-robust complement in the efficient influence function and is not eliminated by cross-fitting . We propose a surrogate-assisted AIPW estimator (SA-AIPW) that integrates over the empirical surrogate distribution through treatment weighting rather than estimating the conditional surrogate density, and so structurally avoids $R_{SY}$. For clustered data, the estimator is shown to be $\sqrt{J}$-consistent and asymptotically linear under a product-rate double-robustness condition.

stat.ME

Win-Ratio Regression for Prioritized Composite Outcomes in Observational Studies: Doubly Robust and Efficient Estimation with Future-Score Correction

Prioritized pairwise outcomes are useful when clinical events follow a natural hierarchy, but censoring before pair resolution complicates estimation. We develop a win-ratio regression framework for this setting by defining a complete-data target over follow-up and deriving an estimating equation for the observed data. The central idea is future-score correction (FC): when censoring prevents later pairwise comparisons from being observed, the method replaces the remaining score with its conditional expectation given the observed history. This correction recovers pairwise information beyond that provided by inverse censoring weights alone. Additionally, we incorporate treatment weighting and baseline outcome augmentation to address baseline confounding. Together, these components yield double robustness for treatment assignment and censoring. Inference is obtained from U-statistic theory. Under standard regularity conditions, the AIPW-FC estimator is asymptotically normal and efficient when all nuisance functions are correctly specified. Simulations with 30%, 50%, and 65% censoring show that efficiency gains from future-score correction increase with the censoring rate, with relative efficiency reaching 1.50 under 65% censoring and near-nominal coverage for AIPW-FC. An application to OneFlorida electronic health record data illustrates the method for a composite outcome that prioritizes death over hospitalization.

stat.ME

Semiparametric Regression Models for Explanatory Variables with Missing Data due to Detection Limit

Detection limit (DL) has become an increasingly ubiquitous issue in statistical analyses of biomedical studies, such as cytokine, metabolite and protein analysis. In regression analysis, if an explanatory variable is left-censored due to concentrations below the DL, one may limit analyses to observed data. In many studies, additional, or surrogate, variables are available to model, and incorporating such auxiliary modeling information into the regression model can improve statistical power. Although methods have been developed along this line, almost all are limited to parametric models for both the regression and left-censored explanatory variable. While some recent work has considered semiparametric regression for the censored DL-effected explanatory variable, the regression of primary interest is still left parametric, which not only makes it prone to biased estimates, but also suffers from high computational cost and inefficiency due to maximizing an extremely complex likelihood function and bootstrap inference. In this paper, we propose a new approach by considering semiparametric generalized linear models (SPGLM) for the primary regression and parametric or semiparametric models for DL-effected explanatory variable. The semiparametric and semiparametric combination provides the most robust inference, while the semiparametric and parametric case enables more efficient inference. The proposed approach is also much easier to implement and allows for leveraging sample splitting and cross fitting (SSCF) to improve computational efficiency in variance estimation. In particular, our approach improves computational efficiency over bootstrap by 450 times. We use simulated and real study data to illustrate the approach.

stat.ME

On Semiparametric Efficiency of an Emerging Class of Regression Models for Between-subject Attributes

The semiparametric regression models have attracted increasing attention owing to their robustness compared to their parametric counterparts. This paper discusses the efficiency bound for functional response models (FRM), an emerging class of semiparametric regression that serves as a timely solution for research questions involving pairwise observations. This new paradigm is especially appealing to reduce astronomical data dimensions for those arising from wearable devices and high-throughput technology, such as microbiome Beta-diversity, viral genetic linkage, single-cell RNA sequencing, etc. Despite the growing applications, the efficiency of their estimators has not been investigated carefully due to the extreme difficulty to address the inherent correlations among pairs. Leveraging the Hilbert-space-based semiparametric efficiency theory for classical within-subject attributes, this manuscript extends such asymptotic efficiency into the broader regression involving between-subject attributes and pinpoints the most efficient estimator, which leads to a sensitive signal-detection in practice. With pairwise outcomes burgeoning immensely as effective dimension-reduction summaries, the established theory will not only fill the critical gap in identifying the most efficient semiparametric estimator but also propel wide-ranging implementations of this new paradigm for between-subject attributes.

stat.ME

Tests for comparing time-invariant and time-varying spectra based on the Anderson-Darling statistic

Based on periodogram-ratios of two univariate time series at different frequency points, two tests are proposed for comparing their spectra. One is an Anderson-Darling-like statistic for testing the equality of two time-invariant spectra. The other is the maximum of Anderson-Darling-like statistics for testing the equality of two spectra no matter that they are time-invariant and time-varying. Both of two tests are applicable for independent or dependent time series. Several simulation examples show that the proposed statistics outperform those that are also based on periodogram-ratios but constructed by the Pearson-like statistics.

stat.ME