SearcharxivSearch

arXiv subjects

Yanyao Yi

Publications and source records attributed to Yanyao Yi.

16 recordsLinked to original sources

Shape-Preserving Covariate Adjustment via Empirical Likelihood in Randomized Experiment

Covariate adjustment improves estimation efficiency in randomized experiments, but standard calibration and augmentation methods, when applied to distribution or survival functions, do not preserve monotonicity---a fundamental property of the estimand. We propose using empirical likelihood with covariate-balancing constraints to construct a covariate-adjusted empirical measure for each treatment arm. Estimators of a broad class of distributional functionals, including cumulative distribution functions, survival functions, quantiles, and restricted mean survival times, are then derived as plug-in functionals of this measure, automatically inheriting proper shape constraints. We establish asymptotic normality with an explicit, guaranteed efficiency gain over unadjusted estimators. The asymptotic distributions are invariant to the randomization scheme, providing a unified inference procedure under simple randomization and all commonly used covariate-adaptive designs satisfying a mild balancing condition. This unified construction, adjusting the empirical measure once and deriving all estimators from it, offers a principled reconciliation of covariate adjustment with shape preservation. Simulations and an application to the SURPASS-4 trial confirm the theoretical gains.

stat.ME

From Estimands to Robust Inference of Treatment Effects in Master Protocol Trials

Master protocol trials use a single overarching protocol to evaluate multiple interventions, diseases, or disease subtypes, where individuals are often randomized to different subsets of intervention arms based on individual characteristics, enrollment timing, and intervention availability. While offering increased flexibility, this constrained and non-uniform intervention assignment poses two fundamental inferential challenges: the precise definition of treatment effects and robust, efficient inference on these effects. These challenges arise primarily because some commonly used analysis approaches may target estimands defined on populations that inadvertently depend on the intervention allocation ratio, making them impossible to fully pre-specify, thereby undermining interpretability and opening the door to ambiguity, post-hoc decisions, and potential bias. This article, for the first time, presents a formal estimand framework for master protocol trials with precise specification of the population. The proposed entire concurrently eligible (ECE) trial population not only preserves the integrity of randomized comparisons but also remains invariant to the randomization ratio. Then, we develop weighting and post-stratification methods to estimate treatment effects under the same minimal assumptions used in traditional randomized trials. We also consider model-assisted covariate adjustment to fully unlock the efficiency potential of master protocol trials while maintaining robustness against model misspecification. The SIMPLIFY trial, a master protocol assessing continuation versus discontinuation of two common therapies in cystic fibrosis, is utilized to highlight the practical significance of this research. All analyses are conducted using the R package RobinCID.

stat.ME

Robust and Data-Adaptive Integration of Nonconcurrent Data in Platform Trials via Gaussian Processes

A platform trial is an innovative clinical trial design that enables simultaneous and continuous evaluation of multiple treatments within a single master protocol. Existing robust methods restrict analyses to concurrently randomized participants due to concerns that including nonconcurrent data may introduce bias from temporal trends. However, this exclusion represents a missed opportunity to improve efficiency. We propose a Gaussian process framework for incorporating nonconcurrent data that exploits temporal smoothness, a key feature of platform trials. The framework includes single-task and multi-task formulations and provides data-adaptive integration of nonconcurrent data with uncertainty quantification. The connection to kernel ridge regression yields a transparent frequentist interpretation of how nonconcurrent data are integrated. We establish two theoretical guarantees: incorporating nonconcurrent controls reduces the posterior variance of the treatment effect, and the resulting bias is controlled by a non-increasing bound. We extend the framework to discrete outcomes and to covariate adjustment, illustrate it on a hypothetical platform trial constructed from SURMOUNT-1, and provide an implementation in the R package RobinCID.

stat.ME

Evolving Longitudinal Patient Histories and Re-enrollment in Master Protocol Trials

A master protocol trial uses a single overarching protocol to test multiple therapies, often across several diseases or subtypes. Although such trials offer considerable flexibility and efficiency, their constrained and non-uniform treatment assignment raises two core challenges: precisely defining treatment effects and conducting robust, efficient inference. These challenges intensify when participants can re-enroll to receive additional eligible therapies over time. To address these issues, we first define a clinically meaningful estimand with a clear population specification for master protocol trials that allow re-enrollment across multiple episodes. Specifically, we define the episode-specific entire concurrently eligible (ECE) population, which preserves the integrity of randomized comparisons and remains invariant to randomization ratios and operational formats. We then introduce a per-episode added-effect estimand that aggregates episode-specific effects into an interpretable overall measure. For inference, we develop weighting and post-stratification estimators under the same minimal assumptions as conventional randomized trials, with model-assisted covariate adjustment to improve efficiency. We establish asymptotic distributions for all estimators and provide cluster-robust variance estimators that properly account for within-participant correlation induced by re-enrollment. We evaluate our methods through extensive simulations and apply our methods to SIMPLIFY, a master protocol trial comparing continuation versus discontinuation of two common cystic fibrosis therapies. All analyses are conducted using the \textsf{R} package \textsf{RobinCID}.

stat.ME

Estimating treatment effects with competing intercurrent events in randomized controlled trials

The analysis of randomized controlled trials is often complicated by intercurrent events (IEs) -- events that occur after treatment initiation and affect either the interpretation or existence of outcome measurements. Examples include treatment discontinuation or the use of additional medications. In two recent clinical trials for systemic lupus erythematosus with complications of IEs, we classify the IEs into two broad categories: effect-informative (e.g., treatment discontinuation due to adverse events or lack of efficacy) and effect-uninformative (e.g., treatment discontinuation due to external factors such as pandemics or relocation). To define a clinically meaningful estimand, we adopt tailored strategies for each category of IEs. For effect-informative IEs, which are often informative about a patient's outcome, we use the composite variable strategy that assigns an outcome value indicative of treatment failure. For effect-uninformative IEs, we apply the hypothetical strategy, assuming their timing is conditionally independent of the outcome given treatment and baseline covariates, and hypothesizing a scenario in which such events do not occur. A central yet previously overlooked challenge is the presence of competing IEs, where the first IE censors all subsequent ones. Despite its ubiquity in practice, this issue has not been explicitly recognized or addressed in previous data analyses due to the lack of rigorous statistical methodology. In this paper, we propose a principled framework to formulate the estimand, establish its nonparametric identification and semiparametric estimation theory, and introduce weighting, outcome regression, and doubly robust estimators. We apply our methods to analyze the two systemic lupus erythematosus trials, demonstrating the robustness and practical utility of the proposed framework.

stat.ME

Covariate Adjustment for Wilcoxon Two Sample Statistic and Test

We apply covariate adjustment to the Wincoxon two sample statistic and Wincoxon-Mann-Whitney test in comparing two treatments. The covariate adjustment through calibration not only improves efficiency in estimation/inference but also widens the application scope of the Wilcoxon two sample statistic and Wincoxon-Mann-Whitney test to situations where covariate-adaptive randomization is used. We motivate how to adjust covariates to reduce variance, establish the asymptotic distribution of adjusted Wincoxon two sample statistic, and provide explicitly the guaranteed efficiency gain. The asymptotic distribution of adjusted Wincoxon two sample statistic is invariant to all commonly used covariate-adaptive randomization schemes so that a unified formula can be used in inference regardless of which covariate-adaptive randomization is applied.

stat.ME

The RobinCar Family: R Tools for Robust Covariate Adjustment in Randomized Clinical Trials

Purpose: Covariate adjustment is a powerful statistical technique that can increase efficiency in clinical trials. Recent guidance from the U.S. FDA provided recommendations and best practices for using covariate adjustment. However, there has existed a gap between the extensive statistical literature on covariate adjustment and software that is easy to use and abides by these best practices. Methods: We have developed the RobinCar Family, which is comprised of RobinCar and RobinCar2. These two R packages enable covariate-adjusted analyses for continuous, discrete, and time-to-event outcomes that follow best practices. For continuous and discrete outcomes, the functions in the RobinCar Family facilitate traditional forms of covariate adjustment such as ANCOVA as well as more recent approaches like ANHECOVA, G-computation with generalized linear models and machine learning models, and adjustment for a super-covariate (as in PROCOVA(TM)). Functions for time-to-event outcomes implement the covariate-adjusted log-rank test, the stratified covariate-adjusted log-rank test, and the marginal covariate-adjusted hazard ratio. The RobinCar Family is supported by the ASA Biopharmaceutical Section Covariate Adjustment Scientific Working Group. Results: We provide an accessible overview of the covariate-adjusted statistical methods, and describe how they are implemented in RobinCar and RobinCar2. We highlight important usage notes for clinical trial practitioners. Conclusion: We apply RobinCar and RobinCar2 functions by analyzing data from the AIDS Clinical Trials Group Study 175, demonstrating that they are straightforward and user-friendly.

stat.ME

Improve the Precision of Area Under the Curve Estimation for Recurrent Events Through Covariate Adjustment

The area under the curve (AUC) of the mean cumulative function (MCF) has recently been introduced as a novel estimand for evaluating treatment effects in recurrent event settings, offering an alternative to the commonly used Lin-Wei-Yang-Ying (LWYY) model. The AUC of the MCF provides a clinically interpretable summary measure that captures the overall burden of disease progression, regardless of whether the proportionality assumption holds. To improve the precision of the AUC estimation while preserving its unconditional interpretability, we propose a nonparametric covariate adjustment approach. This approach guarantees efficiency gain compared to unadjusted analysis, as demonstrated by theoretical asymptotic distributions, and is universally applicable to various randomization schemes, including both simple and covariate-adaptive designs. Extensive simulations across different scenarios further support its advantage in increasing statistical power. Our findings highlight the importance of covariate adjustment for the analysis of AUC in recurrent event settings, offering practical guidance for its application in randomized clinical trials.

stat.ME

Clarifying the Role of the Mantel-Haenszel Risk Difference Estimator in Randomized Clinical Trials

The Mantel-Haenszel (MH) risk difference estimator, commonly used in randomized clinical trials for binary outcomes, calculates a weighted average of stratum-specific risk difference estimators. Traditionally, this method requires the stringent assumption that risk differences are homogeneous across strata, also known as the common (constant) risk difference assumption. In our article, we relax this assumption and adopt a modern perspective, viewing the MH risk difference estimator as an approach for covariate adjustment in randomized clinical trials, distinguishing its use from that in meta-analysis and observational studies. We demonstrate that, under reasonable restrictions on risk difference variability, the MH risk difference estimator consistently estimates the average treatment effect within a standard super-population framework, which is often the primary interest in randomized clinical trials, in addition to estimating a weighted average of stratum-specific risk differences. We rigorously study its properties under the large-stratum and sparse-stratum asymptotic regimes, as well as under mixed-regime settings. Furthermore, for either estimand, we propose a unified robust variance estimator that improves over the popular variance estimators by Greenland and Robins (1985) and Sato et al. (1989) and has provable consistency across these asymptotic regimes, regardless of assuming common risk differences. Extensions of our theoretical results also provide new insights into the Mantel-Haenszel test, the post-stratification estimator, and settings with multiple treatments. Our findings are thoroughly validated through simulations and a clinical trial example.

stat.ME

A General Form of Covariate Adjustment in Randomized Clinical Trials

In randomized clinical trials, adjusting for baseline covariates can improve credibility and efficiency for demonstrating and quantifying treatment effects. This article studies the augmented inverse propensity weighted (AIPW) estimator, which is a general form of covariate adjustment that uses linear, generalized linear, and non-parametric or machine learning models for the conditional mean of the response given covariates. Under covariate-adaptive randomization, we establish general theorems that show a complete picture of the asymptotic normality, {efficiency gain, and applicability of AIPW estimators}. In particular, we provide for the first time a rigorous theoretical justification of using machine learning methods with cross-fitting for dependent data under covariate-adaptive randomization. Based on the general theorems, we offer insights on the conditions for guaranteed efficiency gain and universal applicability {under different randomization schemes}, which also motivate a joint calibration strategy using some constructed covariates after applying AIPW. Our methods are implemented in the R package RobinCar.

stat.ME

Robust Variance Estimation for Covariate-Adjusted Unconditional Treatment Effect in Randomized Clinical Trials with Binary Outcomes

To improve precision of estimation and power of testing hypothesis for an unconditional treatment effect in randomized clinical trials with binary outcomes, researchers and regulatory agencies recommend using g-computation as a reliable method of covariate adjustment. However, the practical application of g-computation is hindered by the lack of an explicit robust variance formula that can be used for different unconditional treatment effects of interest. To fill this gap, we provide explicit and robust variance estimators for g-computation estimators and demonstrate through simulations that the variance estimators can be reliably applied in practice.

stat.ME

Covariate-Adjusted Log-Rank Test: Guaranteed Efficiency Gain and Universal Applicability

Nonparametric covariate adjustment is considered for log-rank type tests of treatment effect with right-censored time-to-event data from clinical trials applying covariate-adaptive randomization. Our proposed covariate-adjusted log-rank test has a simple explicit formula and a guaranteed efficiency gain over the unadjusted test. We also show that our proposed test achieves universal applicability in the sense that the same formula of test can be universally applied to simple randomization and all commonly used covariate-adaptive randomization schemes such as the stratified permuted block and Pocock and Simon's minimization, which is not a property enjoyed by the unadjusted log-rank test. Our method is supported by novel asymptotic theory and empirical results for type I error and power of tests.

stat.ME

Testing for Treatment Effect Twice Using Internal and External Controls in Clinical Trials

Leveraging external controls -- relevant individual patient data under control from external trials or real-world data -- has the potential to reduce the cost of randomized controlled trials (RCTs) while increasing the proportion of trial patients given access to novel treatments. However, due to lack of randomization, RCT patients and external controls may differ with respect to covariates that may or may not have been measured. Hence, after controlling for measured covariates, for instance by matching, testing for treatment effect using external controls may still be subject to unmeasured biases. In this paper, we propose a sensitivity analysis approach to quantify the magnitude of unmeasured bias that would be needed to alter the study conclusion that presumed no unmeasured biases are introduced by employing external controls. Whether leveraging external controls increases power or not depends on the interplay between sample sizes and the magnitude of treatment effect and unmeasured biases, which may be difficult to anticipate. This motivates a combined testing procedure that performs two highly correlated analyses, one with and one without external controls, with a small correction for multiple testing using the joint distribution of the two test statistics. The combined test provides a new method of sensitivity analysis designed for data fusion problems, which anchors at the unbiased analysis based on RCT only and spends a small proportion of the type I error to also test using the external controls. In this way, if leveraging external controls increases power, the power gain compared to the analysis based on RCT only can be substantial; if not, the power loss is small. The proposed method is evaluated in theory and power calculations, and applied to a real trial.

stat.ME

A matching design for augmenting a randomized clinical trial with external control

The use of information from real world to assess the effectiveness of medical products is becoming increasingly popular and more acceptable by regulatory agencies. According to a strategic real-world evidence framework published by U.S. Food and Drug Administration, a hybrid randomized controlled trial that augments internal control arm with real-world data is a pragmatic approach worth more attention. In this paper, we aim to improve on existing matching designs for such a hybrid randomized controlled trial. In particular, we propose to match the entire concurrent randomized clinical trial (RCT) such that (1) the matched external control subjects used to augment the internal control arm are as comparable as possible to the RCT population, (2) every active treatment arm in an RCT with multiple treatments is compared with the same control group, and (3) matching can be conducted and the matched set locked before treatment unblinding to better maintain the data integrity. Besides a weighted estimator, we also introduce a bootstrap method to obtain its variance estimation. The finite sample performance of the proposed method is evaluated by simulations based on data from a real clinical trial.

stat.ME

Toward Better Practice of Covariate Adjustment in Analyzing Randomized Clinical Trials

In randomized clinical trials, adjustments for baseline covariates at both design and analysis stages are highly encouraged by regulatory agencies. A recent trend is to use a model-assisted approach for covariate adjustment to gain credibility and efficiency while producing asymptotically valid inference even when the model is incorrect. In this article we present three considerations for better practice when model-assisted inference is applied to adjust for covariates under simple or covariate-adaptive randomized trials: (1) guaranteed efficiency gain: a model-assisted method should often gain but never hurt efficiency; (2) wide applicability: a valid procedure should be applicable, and preferably universally applicable, to all commonly used randomization schemes; (3) robust standard error: variance estimation should be robust to model misspecification and heteroscedasticity. To achieve these, we recommend a model-assisted estimator under an analysis of heterogeneous covariance working model including all covariates utilized in randomization. Our conclusions are based on an asymptotic theory that provides a clear picture of how covariate-adaptive randomization and regression adjustment alter statistical efficiency. Our theory is more general than the existing ones in terms of studying arbitrary functions of response means (including linear contrasts, ratios, and odds ratios), multiple arms, guaranteed efficiency gain, optimality, and universal applicability.

stat.ME

Inference on Average Treatment Effect under Minimization and Other Covariate-Adaptive Randomization Methods

Covariate-adaptive randomization schemes such as the minimization and stratified permuted blocks are often applied in clinical trials to balance treatment assignments across prognostic factors. The existing theoretical developments on inference after covariate-adaptive randomization are mostly limited to situations where a correct model between the response and covariates can be specified or the randomization method has well-understood properties. Based on stratification with covariate levels utilized in randomization and a further adjusting for covariates not used in randomization, in this article we propose several estimators for model free inference on average treatment effect defined as the difference between response means under two treatments. We establish asymptotic normality of the proposed estimators under all popular covariate-adaptive randomization schemes including the minimization whose theoretical property is unclear, and we show that the asymptotic distributions are invariant with respect to covariate-adaptive randomization methods. Consistent variance estimators are constructed for asymptotic inference. Asymptotic relative efficiencies and finite sample properties of estimators are also studied. We recommend using one of our proposed estimators for valid and model free inference after covariate-adaptive randomization.

stat.ME