SearcharxivSearch

arXiv subjects

Patrick J. Heagerty

Publications and source records attributed to Patrick J. Heagerty.

14 recordsLinked to original sources

Statistical inference with win statistics in cluster-randomized trials with hierarchical composite outcomes

Win statistics have become increasingly popular for analyzing hierarchical composite endpoints in clinical trials. The win ratio, win odds, net benefit, and desirability of outcome ranking (DOOR) share a pairwise-comparison framework and provide complementary summaries of treatment benefit. Despite recent progress in individually randomized trials, statistical inference for these measures in cluster-randomized trials (CRTs) remains underdeveloped. We provide a unified development and comparison of six testing procedures for all four win measures in parallel-arm CRTs: three Wald tests, two randomization-based tests, and a jackknife empirical likelihood ratio test. Through simulations with hierarchical semi-competing risks outcomes, we evaluate type I error rate and power across varying design parameters. With 20 clusters, the randomization-based procedures provided the most stable type I error rate control, while the clustered rank-sum Wald test with a t-distribution was the most reliable among the Wald tests, occasionally carrying a conservative test size. For the win ratio and win odds, the permutation test also yielded higher power than the clustered rank-sum Wald test. With 100 clusters, differences in type I error rate and power were small. These findings favor permutation inference in small CRTs, with the clustered rank-sum Wald test providing an analytic alternative, particularly for net benefit and DOOR. On the other hand, the analytic Wald procedures offer a computationally convenient choice for CRTs with a large number of clusters. We illustrate the methods by reanalyzing the STRIDE trial and implement all procedures in the WinsCRT R package.

stat.ME

From Estimands to Robust Inference of Treatment Effects in Master Protocol Trials

Master protocol trials use a single overarching protocol to evaluate multiple interventions, diseases, or disease subtypes, where individuals are often randomized to different subsets of intervention arms based on individual characteristics, enrollment timing, and intervention availability. While offering increased flexibility, this constrained and non-uniform intervention assignment poses two fundamental inferential challenges: the precise definition of treatment effects and robust, efficient inference on these effects. These challenges arise primarily because some commonly used analysis approaches may target estimands defined on populations that inadvertently depend on the intervention allocation ratio, making them impossible to fully pre-specify, thereby undermining interpretability and opening the door to ambiguity, post-hoc decisions, and potential bias. This article, for the first time, presents a formal estimand framework for master protocol trials with precise specification of the population. The proposed entire concurrently eligible (ECE) trial population not only preserves the integrity of randomized comparisons but also remains invariant to the randomization ratio. Then, we develop weighting and post-stratification methods to estimate treatment effects under the same minimal assumptions used in traditional randomized trials. We also consider model-assisted covariate adjustment to fully unlock the efficiency potential of master protocol trials while maintaining robustness against model misspecification. The SIMPLIFY trial, a master protocol assessing continuation versus discontinuation of two common therapies in cystic fibrosis, is utilized to highlight the practical significance of this research. All analyses are conducted using the R package RobinCID.

stat.ME

Model-robust standardization in stepped wedge cluster randomized trials

Stepped-wedge cluster-randomized trials (SW-CRTs) are widely used in healthcare and implementation science, enabling all clusters to receive the intervention through a staggered rollout. Traditional model-based methods, including generalized estimating equations and mixed models, yield estimates that depend on implicit weighting schemes and parametric assumptions, and therefore may target ambiguous estimands under model misspecification. In this article, we propose a model-robust standardization framework for SW-CRTs that generalizes existing methods from parallel-arm CRTs to address informative sizes. We define causal estimands including horizontal-individual, horizontal-cluster, vertical-individual, and vertical-cluster average treatment effects under a super population framework and introduce a simple procedure that standardizes parametric and semiparametric working models for estimand-aligned analysis. For any specified working model, the resulting estimators remain consistent for their target estimands even if the working regression model is misspecified; moreover, their efficiency improves as the working model more closely approximates the true data-generating process. We evaluate the finite-sample properties of our proposed estimators through extensive simulations. Finally, we illustrate the application of our methods through reanalyses of two real-world SW-CRTs.

stat.ME

Robust and Data-Adaptive Integration of Nonconcurrent Data in Platform Trials via Gaussian Processes

A platform trial is an innovative clinical trial design that enables simultaneous and continuous evaluation of multiple treatments within a single master protocol. Existing robust methods restrict analyses to concurrently randomized participants due to concerns that including nonconcurrent data may introduce bias from temporal trends. However, this exclusion represents a missed opportunity to improve efficiency. We propose a Gaussian process framework for incorporating nonconcurrent data that exploits temporal smoothness, a key feature of platform trials. The framework includes single-task and multi-task formulations and provides data-adaptive integration of nonconcurrent data with uncertainty quantification. The connection to kernel ridge regression yields a transparent frequentist interpretation of how nonconcurrent data are integrated. We establish two theoretical guarantees: incorporating nonconcurrent controls reduces the posterior variance of the treatment effect, and the resulting bias is controlled by a non-increasing bound. We extend the framework to discrete outcomes and to covariate adjustment, illustrate it on a hypothetical platform trial constructed from SURMOUNT-1, and provide an implementation in the R package RobinCID.

stat.ME

Robust and Efficient Semiparametric Inference for the Stepped Wedge Design

Stepped wedge designs (SWDs) are increasingly used to evaluate longitudinal cluster-level interventions but pose substantial challenges for valid inference. Because crossover times are randomized, intervention effects are intrinsically confounded with secular time trends, while heterogeneity across clusters, complex correlation structures, baseline covariate imbalances, and small numbers of clusters further complicate inference. We propose a unified semiparametric framework for estimating possibly time-varying intervention effects in SWDs. Under a semiparametric model on treatment contrast, we develop a nonstandard semiparametric efficiency theory that accommodates correlated observations within clusters, varying cluster-period sizes, and weakly dependent treatment assignments. The resulting estimator is consistent and asymptotically normal even under misspecified covariance structure and control cluster-period means, and is efficient when both are correctly specified. To enable inference with few clusters, we exploit the permutation structure of treatment assignment to propose a standard error estimator that reflects finite-sample variability, with a leave-one-out correction to reduce plug-in bias. The framework also allows incorporation of effect modification and adjustment for imbalanced precision variables through design-based adjustment or double adjustment that additionally incorporates an outcome-based component. Simulations and application to a public health trial demonstrate the robustness and efficiency of the proposed method relative to standard approaches.

stat.ME

A tutorial on conducting sample size and power calculations for detecting treatment effect heterogeneity in cluster randomized trials with linear mixed models

Cluster-randomized trials (CRTs) are a well-established class of designs for evaluating community-based interventions. An essential task in planning these trials is determining the number of clusters and cluster sizes needed to achieve sufficient statistical power for detecting a clinically relevant effect size. While methods for evaluating the average treatment effect (ATE) for the entire study population are well-established, sample size methods for testing heterogeneity of treatment effects (HTEs), i.e., treatment-covariate interaction or difference in subpopulation-specific treatment effects, in CRTs have only recently been developed. For pre-specified analyses of HTEs in CRTs, effect-modifying covariates should, ideally, be accompanied by sample size or power calculations to ensure the trial has adequate power for the planned analyses. Power analysis for testing HTEs is more complex than for ATEs due to the additional design parameters that must be specified. Power and sample size formulas for testing HTEs via linear mixed effects (LME) models have been separately derived for different cluster-randomized designs, including single and multi-period parallel designs, crossover designs, and stepped-wedge designs, and for continuous and binary outcomes. This tutorial provides a consolidated reference guide for these methods and enhances their accessibility through an online R Shiny calculator. We further discuss key considerations for conducting sample size and power calculations to test pre-specified HTE hypotheses in CRTs, highlighting the importance of specifying advanced estimates of intracluster correlation coefficients for both outcomes and covariates, and their implications for power. The sample size methodology and calculator functionality are demonstrated through a real CRT example.

stat.ME

Evolving Longitudinal Patient Histories and Re-enrollment in Master Protocol Trials

A master protocol trial uses a single overarching protocol to test multiple therapies, often across several diseases or subtypes. Although such trials offer considerable flexibility and efficiency, their constrained and non-uniform treatment assignment raises two core challenges: precisely defining treatment effects and conducting robust, efficient inference. These challenges intensify when participants can re-enroll to receive additional eligible therapies over time. To address these issues, we first define a clinically meaningful estimand with a clear population specification for master protocol trials that allow re-enrollment across multiple episodes. Specifically, we define the episode-specific entire concurrently eligible (ECE) population, which preserves the integrity of randomized comparisons and remains invariant to randomization ratios and operational formats. We then introduce a per-episode added-effect estimand that aggregates episode-specific effects into an interpretable overall measure. For inference, we develop weighting and post-stratification estimators under the same minimal assumptions as conventional randomized trials, with model-assisted covariate adjustment to improve efficiency. We establish asymptotic distributions for all estimators and provide cluster-robust variance estimators that properly account for within-participant correlation induced by re-enrollment. We evaluate our methods through extensive simulations and apply our methods to SIMPLIFY, a master protocol trial comparing continuation versus discontinuation of two common cystic fibrosis therapies. All analyses are conducted using the \textsf{R} package \textsf{RobinCID}.

stat.ME

Factors affecting power in stepped wedge trials when the treatment effect varies with time

Stepped wedge cluster randomized trials (SW-CRTs) have historically been analyzed using immediate treatment (IT) models, which assume the effect of the treatment is immediate after treatment initiation and subsequently remains constant over time. However, recent research has shown that this assumption can lead to severely misleading results if treatment effects vary with exposure time, i.e. time since the intervention started. Models that account for time-varying treatment effects, such as the exposure time indicator (ETI) model, allow researchers to target estimands such as the time-averaged treatment effect (TATE) over an interval of exposure time, or the point treatment effect (PTE) representing a treatment contrast at one time point. However, this increased flexibility results in reduced power. In this paper, we use public power calculation software and simulation to characterize factors affecting SW-CRT power. Key elements include choice of estimand, study design considerations, and analysis model selection. or common SW-CRT designs, the sample size (clusters per sequence or individuals per cluster-period) must be increased substantially, commonly by a factor of 1.5 to 3, but often by much more, to maintain 90\% power when switching from an IT model to an ETI model (targeting the TATE over the study). However, the inflation factor is lower for TATE estimands over shorter periods that exclude longer exposure times. In general, SW-CRT designs (including the "staircase" variant) have much greater power for estimating "short-term effects" relative to "long-term effects". For an ETI model targeting a TATE estimand, substantial power can be gained by adding time points to the start of the study or increasing baseline sample size, but surprisingly little power is gained from adding time points to the end of the study. More restrictive choices for modeling the exposure... [truncated]

stat.ME

Analysis of stepped wedge cluster randomized trials in the presence of a time-varying treatment effect

Stepped wedge cluster randomized controlled trials are typically analyzed using models that assume the full effect of the treatment is achieved instantaneously. We provide an analytical framework for scenarios in which the treatment effect varies as a function of exposure time (time since the start of treatment) and define the "effect curve" as the magnitude of the treatment effect on the linear predictor scale as a function of exposure time. The "time-averaged treatment effect", (TATE) and "long-term treatment effect" (LTE) are summaries of this curve. We analytically derive the expectation of the estimator resulting from a model that assumes an immediate treatment effect and show that it can be expressed as a weighted sum of the time-specific treatment effects corresponding to the observed exposure times. Surprisingly, although the weights sum to one, some of the weights can be negative. This implies that the estimator may be severely misleading and can even converge to a value of the opposite sign of the true TATE or LTE. We describe several models that can be used to simultaneously estimate the entire effect curve, the TATE, and the LTE, some of which make assumptions about the shape of the effect curve. We evaluate these models in a simulation study to examine the operating characteristics of the resulting estimators and apply them to two real datasets.

stat.ME

Surrogate-guided sampling designs for classification of rare outcomes from electronic medical records data

Scalable and accurate identification of specific clinical outcomes has been enabled by machine-learning applied to electronic medical record (EMR) systems. The development of classification models requires the collection of a complete labeled data set, where true clinical outcomes are obtained by human expert manual review. For example, the development of natural language processing algorithms requires the abstraction of clinical text data to obtain outcome information necessary for training models. However, if the outcome is rare then simple random sampling results in very few cases and insufficient information to develop accurate classifiers. Since large scale detailed abstraction is often expensive, time-consuming, and not feasible, more efficient strategies are needed. Under such resource constrained settings, we propose a class of enrichment sampling designs, where selection for abstraction is stratified by auxiliary variables related to the true outcome of interest. Stratified sampling on highly specific variables results in targeted samples that are more enriched with cases, which we show translates to increased model discrimination and better statistical learning performance. We provide mathematical details, and simulation evidence that links sampling designs to their resulting prediction model performance. We discuss the impact of our proposed sampling on both model training and validation. Finally, we illustrate the proposed designs for outcome label collection and subsequent machine-learning, using radiology report text data from the Lumbar Imaging with Reporting of Epidemiology (LIRE) study.

stat.ME

A tutorial on evaluating time-varying discrimination accuracy for survival models used in dynamic decision-making

Many medical decisions involve the use of dynamic information collected on individual patients toward predicting likely transitions in their future health status. If accurate predictions are developed, then a prognostic mode can identify patients at greatest risk for future adverse events, and may be used clinically to define populations appropriate for targeted intervention. In practice, a prognostic model is often used to guide decisions at multiple time points over the course of disease, and classification performance, i.e. sensitivity and specificity, for distinguishing high-risk versus low-risk individuals may vary over time as an individual's disease status and prognostic information change. In this tutorial, we detail contemporary statistical methods that can characterize the time-varying accuracy of prognostic survival models when used for dynamic decision-making. Although statistical methods for evaluating prognostic models with simple binary outcomes are well established, methods appropriate for survival outcomes are less well known and require time-dependent extensions of sensitivity and specificity to fully characterize longitudinal biomarkers or models. The methods we review are particularly important in that they allow for appropriate handling of censored outcomes commonly encountered with event-time data. We highlight the importance of determining whether clinical interest is in predicting cumulative (or prevalent) cases over a fixed future time interval versus predicting incident cases over a range of follow-up time, and whether patient information is static or updated over time. We discuss implementation of time-dependent ROC approaches using relevant R statistical software packages. The statistical summaries are illustrated using a liver prognostic model to guide transplantation in primary biliary cirrhosis.

stat.ME

A Novel Tool to Evaluate the Accuracy of Predicting Survival in Cystic Fibrosis

Background: Effective allocation of limited donor lungs in cystic fibrosis (CF) requires accurate survival predictions, so that high-risk patients may be prioritized for transplantation. In practice, decisions about allocation are made dynamically, using routinely updated assessments. We present a novel tool for evaluating risk prediction models that, unlike traditional methods, captures the dynamic nature of decision-making. Methods: Predicted risk is used as a score to rank incident deaths versus patients who survive, with the goal of ranking the deaths higher. The mean rank across deaths at a given time measures time-specific predictive accuracy; when assessed over time, it reflects time-varying accuracy. Results: Applying this approach to CF Registry data on patients followed from 1993-2011, we show that traditional methods do not capture the performance of models used dynamically in the clinical setting. Previously proposed multivariate risk scores perform no better than forced expiratory volume in 1 second as a percentage of predicted normal (FEV1%) alone. Despite its value for survival prediction, FEV1% has a low sensitivity of 45% over time (for fixed specificity of 95%), leaving room for improvement in prediction. Finally, prediction accuracy with annually-updated FEV1% shows minor differences compared to FEV1% updated every 2 years, which may have clinical implications regarding the optimal frequency of updating clinical information. Conclusions: It is imperative to continue to develop models that accurately predict survival in CF. Our proposed approach can serve as the basis for evaluating the predictive ability of these models by better accounting for their dynamic clinical use.

stat.AP

Biased sampling designs to improve research efficiency: Factors influencing pulmonary function over time in children with asthma

Substudies of the Childhood Asthma Management Program [Control. Clin. Trials 20 (1999) 91-120; N. Engl. J. Med. 343 (2000) 1054-1063] seek to identify patient characteristics associated with asthma symptoms and lung function. To determine if genetic measures are associated with trajectories of lung function as measured by forced vital capacity (FVC), children in the primary cohort study retrospectively had candidate loci evaluated. Given participant burden and constraints on financial resources, it is often desirable to target a subsample for ascertainment of costly measures. Methods that can leverage the longitudinal outcome on the full cohort to selectively measure informative individuals have been promising, but have been restricted in their use to analysis of the targeted subsample. In this paper we detail two multiple imputation analysis strategies that exploit outcome and partially observed covariate data on the nonsampled subjects, and we characterize alternative design and analysis combinations that could be used for future studies of pulmonary function and other outcomes. Candidate predictor (e.g., IL10 cytokine polymorphisms) associations obtained from targeted sampling designs can be estimated with very high efficiency compared to standard designs. Further, even though multiple imputation can dramatically improve estimation efficiency for covariates available on all subjects (e.g., gender and baseline age), relatively modest efficiency gains were observed in parameters associated with predictors that are exclusive to the targeted sample. Our results suggest that future studies of longitudinal trajectories can be efficiently conducted by use of outcome-dependent designs and associated full cohort analysis.

stat.AP

Evaluating epoetin dosing strategies using observational longitudinal data

Epoetin is commonly used to treat anemia in chronic kidney disease and End Stage Renal Disease subjects undergoing dialysis, however, there is considerable uncertainty about what level of hemoglobin or hematocrit should be targeted in these subjects. In order to address this question, we treat epoetin dosing guidelines as a type of dynamic treatment regimen. Specifically, we present a methodology for comparing the effects of alternative treatment regimens on survival using observational data. In randomized trials patients can be assigned to follow a specific management guideline, but in observational studies subjects can have treatment paths that appear to be adherent to multiple regimens at the same time. We present a cloning strategy in which each subject contributes follow-up data to each treatment regimen to which they are continuously adherent and artificially censored at first nonadherence. We detail an inverse probability weighted log-rank test with a valid asymptotic variance estimate that can be used to test survival distributions under two regimens. To compare multiple regimens, we propose several marginal structural Cox proportional hazards models with robust variance estimation to account for the creation of clones. The methods are illustrated through simulations and applied to an analysis comparing epoetin dosing regimens in a cohort of 33,873 adult hemodialysis patients from the United States Renal Data System.

stat.AP