SearcharxivSearch

arXiv subjects

Guangyu Tong

Publications and source records attributed to Guangyu Tong.

15 recordsLinked to original sources

Cluster randomized crossover trials with very few clusters but multiple periods: which analyses for continuous outcomes should be used?

Cluster randomized crossover (CRXO) trials are often used when individual randomization is impractical and the number of available clusters is limited. However, statistical analysis of CRXO trials is complex because of the need to account for complex correlation structures over time. It becomes especially challenging when very few clusters are used because standard modeling assumptions may lead to unstable variance estimates, poor confidence interval coverage, and inflated type I error. This study evaluates individual-level mixed-effects and fixed-effects models with and without a cluster-period random effect, cluster-period summary analysis using normal- or \(t\)-based inference, and two-period crossover-difference estimators. Using extensive simulation studies under both nested exchangeable and discrete time decay correlation structures, we compare model performance in terms of bias, root mean squared error, coverage probability, type I error, and convergence. Across scenarios, all models produced approximately unbiased treatment effect estimates, but their inferential performance differed substantially. Models that explicitly accounted for cluster-period heterogeneity generally provided the most reliable control of coverage and type I error, whereas simpler exchangeable models performed adequately only when the true correlation structure closely matched their assumptions. Cluster-period level analysis performance improved with increasing numbers of periods but was unreliable in the sparsest designs. Overall, the findings suggest that in CRXO trials with very few clusters, accurate modeling of cluster-period correlation is more important than the choice between fixed and random cluster intercepts, and that results from extremely sparse designs should be interpreted with caution.

stat.ME

Statistical inference with win statistics in cluster-randomized trials with composite outcomes

Win statistics have become increasingly popular for analyzing hierarchical composite endpoints in clinical trials, because they summarize treatment benefit through pairwise comparisons that respect the clinical importance order among outcome components. The win ratio, win odds, net benefit, and desirability of outcome ranking (DOOR) are all based on the same underlying pairwise comparison methodology and can complement one another to show the strength of the treatment effect. Despite recent progress on win statistics, statistical inference for win statistics in cluster randomized trials (CRTs) remains underdeveloped. In this paper, we provide a comprehensive survey of testing procedures for the win ratio, win odds, net benefit, and DOOR in parallel-arm CRTs with hierarchical composite outcomes. Then based on each win statistic, we compare different testing procedures, including Wald tests based on cluster rank sum statistics and bivariate clustered U-statistics, tests that use a cluster jackknife variance, a score permutation test, a permutation based procedure with analytical variance estimation, and likelihood ratio test derived from clustered jackknife estimates. Through simulation studies that consider varying scenarios such as different cluster sizes, intracluster correlations, and censoring-induced ties, we characterize the finite-sample type I error and power of each procedure across a range of practical settings with small and large numbers of clusters.We illustrate our methods by reanalyzing the Strategies to Reduce Injuries and Develop Confidence in Elders (STRIDE) pragmatic CRT, and implement all win statistics methods in the WinsCRT R package.

stat.ME

Leveraging machine learning to estimate individualized treatment effects in cluster-randomized trials

Cluster-randomized trials (CRTs) are widely used to evaluate interventions delivered at the clinic, practice, or community level. Although standard analyses typically target average treatment effects, such summaries mask potentially meaningful variation in treatment response across individuals and clusters. This work addresses the estimation of conditional average treatment effects (CATEs) for continuous outcomes in two-arm parallel CRTs by defining causal estimands that incorporate both individual- and cluster-level baseline covariates while marginalizing over unobserved cluster heterogeneity. To estimate these quantities, we develop a unified framework based on mixed-effects machine learning, integrating and extending a range of existing approaches, including Bayesian additive regression trees with random effects, multilevel Bayesian causal forests, mixed-effects random forests, several mixed-effects gradient boosting procedures, and generalized additive mixed models, while incorporating cluster-specific random intercepts to account for within-cluster dependence. We evaluate these methods across diverse simulation scenarios and demonstrate their use in the Task Shifting and Blood Pressure Control in Ghana CRT, which investigates strategies for improving hypertension management. Drawing on these investigations, we provide practical guidance for applying mixed-effects machine learning to quantify treatment-effect heterogeneity in CRTs, together with reproducible code that enables investigators to implement all methods within a coherent workflow.

stat.ME

Optimizing Complex Health Intervention Packages through the Learn-As-you-GO (LAGO) Design

In the face of vast numbers of preventable deaths worldwide and gaping disparities in their distribution, we cannot afford to conduct null and inconclusive effectiveness and implementation trials of evidence-based interventions. The gold standard in biomedical research, the individually randomized clinical trial, is ill-suited as the primary tool for knowledge generation for contextually relevant, scalable, complex public health interventions of multi-component strategies. In this paper, we discuss the new Learn-As-you-GO (LAGO) design. In LAGO trials, the components of a complex intervention package are repeatedly optimized in pre-planned stages, until the package achieves its outcome and power goals with minimized cost and/or other optimization criteria, such as maximizing patient satisfaction. In this paper, the inputs to, and outputs of, LAGO are described, along with its general methodology. The methods are illustrated in the BetterBirth study, a large trial that aimed to reduce maternal and neonatal mortality in Uttar Pradesh, India, using the WHO essential birth practices checklist. Despite its scale, the BetterBirth study failed to demonstrate a significant effect of the intervention package on the primary health endpoint that included maternal mortality. We show how this unfortunate outcome could have been remedied had LAGO been used. LAGO is further illustrated through the discussion of several ongoing LAGO-informed implementation trials of HIV and non-communicable diseases in the United States and Sub-Saharan Africa. The Learn-As-you-GO (LAGO) design optimizes a complex, multi-level intervention for minimum cost, pre-specified power, and a pre-specified effectiveness goal, by adapting the intervention as the study is conducted, reducing risk of trial failure.

stat.ME

Doubly robust estimators of the restricted mean time in favor estimands in individual- and cluster-randomized trials

Progressive multi-state survival outcomes are common in trials with recurrent or sequential events and require treatment effect estimands that remain interpretable without proportional intensity or Markov assumptions. The restricted mean time in favor of treatment (RMT-IF) extends the restricted mean survival time to ordered multi-state processes and provides such an interpretable estimand. However, existing RMT-IF methods are nonparametric, assume covariate-independent censoring for independent observations, and do not accommodate cluster-randomized trials (CRTs), limiting both efficiency and applicability. We develop a class of doubly robust estimators for RMT-IF under right censoring using an augmented inverse-probability weighting framework that combines stage-specific outcome regression with arm-specific censoring models, yielding consistency when either nuisance model is correctly specified. We further extend the framework to CRTs by formalizing both cluster-level and individual-level average RMT-IF estimands to address informative cluster size and by constructing corresponding doubly robust estimators that account for within-cluster correlation. For inference, we employ model-agnostic jackknife variance estimators in both individually randomized and cluster-randomized settings. Extensive simulation studies demonstrate finite-sample performance, and the methods are illustrated using two randomized trial examples.

stat.ME

Can discrete-time analyses be trusted for stepped wedge trials with continuous recruitment?

In stepped wedge cluster randomized trials (SW-CRTs), interventions are sequentially rolled out to clusters over multiple periods. It is common practice to analyze data from SW-CRTs using linear mixed models that treat time as discrete. However, a recent systematic review found that 95.1% of cross-sectional SW-CRTs recruit individuals continuously over time. Despite the high prevalence of such continuous recruitment designs, there has been limited guidance on how to draw model-robust inference when analyzing such SW-CRTs. In this article, we investigate through simulations the implications of using such discrete-time linear mixed models in the case of continuous recruitment designs with a continuous outcome. Specifically, in the data-generating process, we characterize continuous recruitment using a continuous-time exponential decay correlation structure in the presence or absence of a fixed continuous period effect, addressing scenarios both with and without a random or exposure-time-dependent intervention effect. We then analyze the simulated data under three popular discrete-time working correlation structures: simple exchangeable, nested exchangeable, and discrete-time exponential decay, with a robust sandwich variance estimator. Our results demonstrate that discrete-time analysis often yields negligible bias and that the robust variance estimator with the Mancl and DeRouen correction consistently achieves nominal coverage and type I error rate. One important exception occurs when recruitment patterns vary systematically between control and intervention periods, where discrete-time analysis leads to slightly biased estimates. Finally, we illustrate these findings by reanalyzing a completed SW-CRT.

stat.ME

Uncovering Treatment Effect Heterogeneity in Pragmatic Gerontology Trials

Detecting heterogeneity in treatment response enriches the interpretation of gerontologic trials. In aging research, estimating the effect of the intervention on clinically meaningful outcomes faces analytical challenges when it is truncated by death. For example, in the Whole Systems Demonstrator trial, a large cluster-randomized study evaluating telecare among older adults, the overall effect of the intervention on quality of life was found to be null. However, this marginal intervention estimate obscures potential heterogeneity of individuals responding to the intervention, particularly among those who survive to the end of follow-up. To explore this heterogeneity, we adopt a causal framework grounded in principal stratification, targeting the Survivor Average Causal Effect (SACE)-the treatment effect among "always-survivors," or those who would survive regardless of treatment assignment. We extend this framework using Bayesian Additive Regression Trees (BART), a nonparametric machine learning method, to flexibly model both latent principal strata and stratum-specific potential outcomes. This enables the estimation of the Conditional SACE (CSACE), allowing us to uncover variation in treatment effects across subgroups defined by baseline characteristics. Our analysis reveals that despite the null average effect, some subgroups experience distinct quality of life benefits (or lack thereof) from telecare, highlighting opportunities for more personalized intervention strategies. This study demonstrates how embedding machine learning methods, such as BART, within a principled causal inference framework can offer deeper insights into trial data with complex features including truncation by death and clustering-key considerations in analyzing pragmatic gerontology trials.

stat.AP

Doubly robust estimation and sensitivity analysis with outcomes truncated by death in multi-arm clinical trials

In clinical trials, the observation of participant outcomes may frequently be hindered by death, leading to ambiguity in defining a scientifically meaningful final outcome for those who die. Principal stratification methods are valuable tools for addressing the average causal effect among always-survivors, i.e., the average treatment effect among a subpopulation defined as those who would survive regardless of treatment assignment. Although robust methods for the truncation-by-death problem in two-arm clinical trials have been previously studied, its expansion to multi-arm clinical trials remains elusive. In this article, we study the identification of a class of survivor average causal effect estimands with multiple treatments under monotonicity and principal ignorability, and first propose simple weighting and regression approaches for point estimation. As a further improvement, we derive the efficient influence function to motivate doubly robust estimators for the survivor average causal effects in multi-arm clinical trials. We also propose sensitivity methods under violations of key causal assumptions. Extensive simulations are conducted to investigate the finite-sample performance of the proposed methods against the existing methods, and a real data example is used to illustrate how to operationalize the proposed estimators and the sensitivity methods in practice.

stat.ME

Covariate-adjusted win statistics in randomized clinical trials with ordinal outcomes

Ordinal outcomes are common in clinical settings where they often represent increasing levels of disease progression or different levels of functional impairment. In this article, we focus on representing the average treatment effect for ordinal outcomes via intrinsic pairwise outcome comparisons captured through win estimands, such as the win ratio and win difference. Recognizing the value of baseline covariate adjustment toward enhanced precision, we first develop propensity score weighting estimators, including both inverse probability weighting (IPW) and overlap weighting (OW), tailored to estimating win estimands. Furthermore, we develop augmented weighting estimators that leverage an additional ordinal outcome regression to potentially improve efficiency over weighting alone. Leveraging the theory of U-statistics, we establish the asymptotic theory for all estimators, and derive closed-form variance estimators to support statistical inference. We also prove that all of the covariate-adjusted estimators do not compromise consistency for the target estimand even when the associated working models are incorrectly specified; hence these covariate-adjusted estimators are model-robust. Through simulations we demonstrate the enhanced efficiency of the weighted estimators over the unadjusted estimator, with the augmented weighting estimators showing a further improvement in efficiency except for extreme cases. Finally, we illustrate our proposed methods with the ORCHID trial, and implement our covariate adjustment methods in an R package winPSW.

stat.ME

Principal stratification with recurrent events truncated by a terminal event: A nested Bayesian nonparametric approach

Recurrent events often serve as key endpoints in clinical studies but may be prematurely truncated by terminal events such as death, creating selection bias and complicating causal inference. To address this challenge, we develop a Bayesian nonparametric framework to address potential selection bias due to truncation by death within the continuous-time principal stratification framework. We introduce causal estimands for recurrent events in the presence of a terminal event and derive a partial identification result for the estimand under a dual-frailty framework, enabling transparent sensitivity analysis for non-identifiable parameters. We then propose a flexible Bayesian nonparametric prior, the enriched dependent Dirichlet process, specifically designed for joint modeling of recurrent and terminal events, addressing a limitation where standard Dirichlet process priors create random partitions dominated by recurrent events, yielding poor predictive performance for terminal events. Simulations are carried out to show that our method has superior performance compared to existing methods. We apply the proposed new Bayesian nonparametric methods to infer the causal effect of a structured exercise program on rehospitalizations, which are subject to truncation by death.

stat.ME

Bayesian inference for cluster-randomized trials with multivariate outcomes subject to both truncation by death and missingness

Cluster-randomized trials (CRTs) on fragile populations frequently encounter complex attrition problems where the reasons for missing outcomes can be heterogeneous, with participants who are known alive, known to have died, or with unknown survival status, and with complex and distinct missing data mechanisms for each group. Although existing methods have been developed to address death truncation in CRTs, no existing methods can jointly accommodate participants who drop out for reasons unrelated to mortality or serious illnesses, or those with an unknown survival status. This paper proposes a Bayesian framework for estimating survivor average causal effects in CRTs while accounting for different types of missingness. Our approach uses a multivariate outcome that jointly estimates the causal effects, and in the posterior estimates, we distinguish the individual-level and the cluster-level survivor average causal effect. We perform simulation studies to evaluate the performance of our model and found low bias and high coverage on key parameters across several different scenarios. We use data from a geriatric CRT to illustrate the use of our model. Although our illustration focuses on the case of a bivariate continuous outcome, our model is straightforwardly extended to accommodate more than two endpoints as well as other types of endpoints (e.g., binary). Thus, this work provides a general modeling framework for handling complex missingness in CRTs and can be applied to a wide range of settings with aging and palliative care populations.

stat.ME

A tutorial on conducting sample size and power calculations for detecting treatment effect heterogeneity in cluster randomized trials with linear mixed models

Cluster-randomized trials (CRTs) are a well-established class of designs for evaluating community-based interventions. An essential task in planning these trials is determining the number of clusters and cluster sizes needed to achieve sufficient statistical power for detecting a clinically relevant effect size. While methods for evaluating the average treatment effect (ATE) for the entire study population are well-established, sample size methods for testing heterogeneity of treatment effects (HTEs), i.e., treatment-covariate interaction or difference in subpopulation-specific treatment effects, in CRTs have only recently been developed. For pre-specified analyses of HTEs in CRTs, effect-modifying covariates should, ideally, be accompanied by sample size or power calculations to ensure the trial has adequate power for the planned analyses. Power analysis for testing HTEs is more complex than for ATEs due to the additional design parameters that must be specified. Power and sample size formulas for testing HTEs via linear mixed effects (LME) models have been separately derived for different cluster-randomized designs, including single and multi-period parallel designs, crossover designs, and stepped-wedge designs, and for continuous and binary outcomes. This tutorial provides a consolidated reference guide for these methods and enhances their accessibility through an online R Shiny calculator. We further discuss key considerations for conducting sample size and power calculations to test pre-specified HTE hypotheses in CRTs, highlighting the importance of specifying advanced estimates of intracluster correlation coefficients for both outcomes and covariates, and their implications for power. The sample size methodology and calculator functionality are demonstrated through a real CRT example.

stat.ME

Moving toward best practice when using propensity score weighting in survey observational studies

Propensity score weighting is a common method for estimating treatment effects with survey data. The method is applied to minimize confounding using measured covariates that are often different between individuals in treatment and control. However, existing literature does not reach a consensus on the optimal use of survey weights for population-level inference in the propensity score weighting analysis. Under the balancing weights framework, we provided a unified solution for incorporating survey weights in both the propensity score of estimation and the outcome regression model. We derived estimators for different target populations, including the combined, treated, controlled, and overlap populations. We provide a unified expression of the sandwich variance estimator and demonstrate that the survey-weighted estimator is asymptotically normal, as established through the theory of M-estimators. Through an extensive series of simulation studies, we examined the performance of our derived estimators and compared the results to those of alternative methods. We further carried out two case studies to illustrate the application of the different methods of propensity score analysis with complex survey data. We concluded with a discussion of our findings and provided practical guidelines for propensity score weighting analysis of observational data from complex surveys.

stat.ME

A Bayesian Machine Learning Approach for Estimating Heterogeneous Survivor Causal Effects: Applications to a Critical Care Trial

Motivated by the Acute Respiratory Distress Syndrome Network (ARDSNetwork) ARDS respiratory management (ARMA) trial, we developed a flexible Bayesian machine learning approach to estimate the average causal effect and heterogeneous causal effects among the always-survivors stratum when clinical outcomes are subject to truncation. We adopted Bayesian additive regression trees (BART) to flexibly specify separate models for the potential outcomes and latent strata membership. In the analysis of the ARMA trial, we found that the low tidal volume treatment had an overall benefit for participants sustaining acute lung injuries on the outcome of time to returning home, but substantial heterogeneity in treatment effects among the always-survivors, driven most strongly by sex and the alveolar-arterial oxygen gradient at baseline (a physiologic measure of lung function and source of hypoxemia). These findings illustrate how the proposed methodology could guide the prognostic enrichment of future trials in the field. We also demonstrated through a simulation study that our proposed Bayesian machine learning approach outperforms other parametric methods in reducing the estimation bias in both the average causal effect and heterogeneous causal effects for always-survivors.

stat.AP

PSweight: An R Package for Propensity Score Weighting Analysis

Propensity score weighting is an important tool for comparative effectiveness research.Besides the inverse probability of treatment weights (IPW), recent development has introduced a general class of balancing weights, corresponding to alternative target populations and estimands. In particular, the overlap weights (OW) lead to optimal covariate balance and estimation efficiency, and a target population of scientific and policy interest. We develop the R package PSweight to provide a comprehensive design and analysis platform for causal inference based on propensity score weighting. PSweight supports (i) a variety of balancing weights, (ii) binary and multiple treatments,(iii) simple and augmented weighting estimators, (iv) nuisance-adjusted sandwich variances, and(v) ratio estimands. PSweight also provides diagnostic tables and graphs for covariate balance assessment. We demonstrate the functionality of the package using a data example from the NationalChild Development Survey (NCDS), where we evaluate the causal effect of educational attainment on income.

stat.ME