Searcharxiv⌕ Search

arXiv subjects

Michael O. Harhay

Publications and source records attributed to Michael O. Harhay.

At least 19 recordsLinked to original sources

Heterogeneous survivor average causal effects beyond monotonicity: Applications to a clinical trial evaluating mechanical ventilation strategies

Clinical trials in critical care often evaluate outcomes that are truncated by death, such as time to discharge alive, for which treatment effects are not well defined among patients who would not survive under one or both treatment strategies. Principal stratification provides a natural framework for defining survivor causal effects, but existing approaches often rely on monotonicity assumptions that may be implausible in settings where treatment can affect survival in competing directions. This concern is illustrated by the ARDS Network trial of lower versus higher positive end-expiratory pressure (PEEP), in which higher PEEP may benefit some patients by improving alveolar recruitment while harming others through over-distention or hemodynamic compromise. Moreover, the largely null average findings of the original trial do not rule out the possibility of clinically meaningful subgroups that may benefit from or be harmed by higher PEEP. We propose a monotonicity-free framework for identifying conditional survivor average causal effects (CSACE) using an interpretable sensitivity parameter that characterizes latent principal-stratum membership. We then develop a flexible Bayesian estimation strategy based on BART, a posterior-mean-based variable importance measure, and a distribution-aware conditional inference tree that uses the full posterior draws of individualized CSACEs. In simulation studies, the proposed BART-based approach improves individualized CSACE estimation and variable-importance recovery compared with linear modeling, especially under nonlinear treatment-effect heterogeneity. Applied to the ARDS PEEP trial, our method reveals clinically interpretable heterogeneity among estimated always-survivors, identifying subgroups with posterior evidence of benefit or harm that are obscured by the overall null trial results.

stat.ME↗

Saturation in G: simple & robust causal inference in cluster randomized trials with informative cluster sizes

Cluster randomized trials (CRTs) can exhibit informative cluster sizes (ICS) where cluster size is associated with outcomes and/or treatment effects. Under ICS, the individual and cluster-average treatment effects (iATE, cATE) can diverge, and the conventional linear mixed-effects model (LMM) and generalized estimating equation (GEE) with an exchangeable working correlation can produce data-dependent weighted contrasts that are not consistent for either estimand. In these settings with ICS, we propose easy to implement "cluster-size saturated models with g-computation" (CS-g), which employ a simple two-step adjustment to standard practice: (1.) augment the appropriately weighted working LMM or GEE with a saturated continuous cluster-size main effect and treatment x cluster-size interaction, and (2.) apply g-computation to target an interpretable marginal estimand. We prove that the appropriately weighted cluster-size saturated LMM with g-computation and more general cluster-size saturated GEE with g-computation can consistently target the iATE and cATE, among a broad class of interpretable estimands, while allowing for ICS. Crucially, this consistency holds under arbitrary misspecification of other model components, including the functional form of the saturated cluster-size terms. Furthermore, we demonstrate exact finite-sample equivalence between these consistent CS-g estimators and their model-robust standardization counterparts. Across simulations with continuous and binary outcomes, the proposed CS-g estimators were unbiased, more efficient than other consistent estimators, and returned greater power to detect ICS. A re-analysis of the PPACT P-CRT further illustrates the approach. Altogether, CS-g offers a simple, robust, and efficient route to target interpretable marginal effects in P-CRTs with ICS.

stat.ME↗

Randomization inference for stepped-wedge designs with noncompliance with application to a palliative care pragmatic trial

While palliative care is increasingly commonly delivered to hospitalized patients with serious illnesses, few studies have estimated its causal effects. Courtright et al. (2016) adopted a stepped-wedge cluster-randomized design to assess the effect of palliative care on a patient-centered outcome. The randomized intervention was a nudge to administer palliative care but did not guarantee receipt of palliative care, resulting in noncompliance. A subsequent analysis using methods suited for standard trial designs produced statistically anomalous results, as an intention-to-treat analysis found no effect while an instrumental variable analysis did (Courtright et al. 2024). This highlights the need for a more principled approach to address noncompliance in stepped-wedge designs. We provide a formal causal inference framework for the stepped-wedge design with noncompliance by introducing a relevant causal estimand and corresponding estimators and inferential procedures. Through numerical studies, we compare an array of estimators and provide practical guidance in choosing an analysis method. Finally, we apply our recommended methods to reanalyze the palliative care pragmatic trial, producing point estimates suggesting a larger effect than the original analysis, but intervals that did not reach statistical significance.

stat.ME↗

Beyond principal ignorability: Nonparametric sensitivity bounds for principal stratification

Principal stratification is an effective framework addressing intermediate variables in causal inference. However, point identification of the principal causal effects (PCEs) often requires the untestable principal ignorability (PI) assumption. This article develops a nonparametric sensitivity analysis framework for evaluating PI violations. We introduce a margin-free bounding factor parameterized by the selection and outcome relative risks of an unmeasured confounder. Using this bounding factor, we derive sharp nonparametric bounds for each PCE. We prove that these bounds nest within the worst-case nonparametric bounds with and without the monotonicity assumption. We then discuss Cornfield-type conditions and principal E-values that quantify the minimum joint magnitude of unmeasured confounding required to nullify the target PCE. Furthermore, we generalize this methodology to principal generalized causal effects, extending the sensitivity bounds and falsification thresholds to the recent pairwise comparison estimands evaluated over a product space.

stat.ME↗

A tutorial on conducting sample size and power calculations for detecting treatment effect heterogeneity in cluster randomized trials with linear mixed models

Cluster-randomized trials (CRTs) are a well-established class of designs for evaluating community-based interventions. An essential task in planning these trials is determining the number of clusters and cluster sizes needed to achieve sufficient statistical power for detecting a clinically relevant effect size. While methods for evaluating the average treatment effect (ATE) for the entire study population are well-established, sample size methods for testing heterogeneity of treatment effects (HTEs), i.e., treatment-covariate interaction or difference in subpopulation-specific treatment effects, in CRTs have only recently been developed. For pre-specified analyses of HTEs in CRTs, effect-modifying covariates should, ideally, be accompanied by sample size or power calculations to ensure the trial has adequate power for the planned analyses. Power analysis for testing HTEs is more complex than for ATEs due to the additional design parameters that must be specified. Power and sample size formulas for testing HTEs via linear mixed effects (LME) models have been separately derived for different cluster-randomized designs, including single and multi-period parallel designs, crossover designs, and stepped-wedge designs, and for continuous and binary outcomes. This tutorial provides a consolidated reference guide for these methods and enhances their accessibility through an online R Shiny calculator. We further discuss key considerations for conducting sample size and power calculations to test pre-specified HTE hypotheses in CRTs, highlighting the importance of specifying advanced estimates of intracluster correlation coefficients for both outcomes and covariates, and their implications for power. The sample size methodology and calculator functionality are demonstrated through a real CRT example.

stat.ME↗

Design-based nested instrumental variable analysis

Two binary instrumental variables (IVs) are nested if individuals who comply under one binary IV also comply under the other. This situation often arises when the two IVs represent different intensities of encouragement or discouragement to take the treatment, with one stronger than the other. In a nested IV structure, treatment effects can be identified for two latent subgroups: always-compliers and switchers. Always-compliers are individuals who comply even under the weaker IV, while switchers are those who do not comply under the weaker IV but do under the stronger IV. We introduce a novel pair-of-pairs nested IV design, where each matched stratum consists of four units organized in two pairs. We develop design-based inference for the always-complier sample average treatment effect and switcher sample average treatment effect. In a nested IV analysis, IV assignment is randomized within each IV pair; however, whether a study unit receives the weaker or stronger IV may not be randomized. To address this complication, we then propose a novel partly biased randomization scheme and study design-based inference under this new scheme. Using extensive simulation studies, we demonstrate the validity of the proposed method even in challenging scenarios with small sample sizes and a low proportion of switchers. Applying the nested IV framework, we estimated that 52.2% (95% CI: 50.4%-53.9%) of participants enrolled at the Henry Ford Health System in the Prostate, Lung, Colorectal, and Ovarian Cancer Screening Trial were always-compliers, while 26.7% (95% CI: 24.5%-28.9%) were switchers. Among always-compliers, flexible sigmoidoscopy was associated with a trend toward a decreased colorectal cancer rate. No effect was detected among switchers. This offers a richer interpretation of why no increase in the intention-to-treat effect was observed after 1997, even though the compliance rate rose.

stat.ME↗

Leveraging machine learning to estimate individualized treatment effects in cluster-randomized trials

Cluster-randomized trials (CRTs) are widely used to evaluate interventions delivered at the clinic, practice, or community level. Although standard analyses typically target average treatment effects, such summaries mask potentially meaningful variation in treatment response across individuals and clusters. This work addresses the estimation of conditional average treatment effects (CATEs) for continuous outcomes in two-arm parallel CRTs by defining causal estimands that incorporate both individual- and cluster-level baseline covariates while marginalizing over unobserved cluster heterogeneity. To estimate these quantities, we develop a unified framework based on mixed-effects machine learning, integrating and extending a range of existing approaches, including Bayesian additive regression trees with random effects, multilevel Bayesian causal forests, mixed-effects random forests, several mixed-effects gradient boosting procedures, and generalized additive mixed models, while incorporating cluster-specific random intercepts to account for within-cluster dependence. We evaluate these methods across diverse simulation scenarios and demonstrate their use in the Task Shifting and Blood Pressure Control in Ghana CRT, which investigates strategies for improving hypertension management. Drawing on these investigations, we provide practical guidance for applying mixed-effects machine learning to quantify treatment-effect heterogeneity in CRTs, together with reproducible code that enables investigators to implement all methods within a coherent workflow.

stat.ME↗

Principal stratification with recurrent events truncated by a terminal event: A nested Bayesian nonparametric approach

Recurrent events often serve as key endpoints in clinical studies but may be prematurely truncated by terminal events such as death, creating selection bias and complicating causal inference. To address this challenge, we develop a Bayesian nonparametric framework to address potential selection bias due to truncation by death within the continuous-time principal stratification framework. We introduce causal estimands for recurrent events in the presence of a terminal event and derive a partial identification result for the estimand under a dual-frailty framework, enabling transparent sensitivity analysis for non-identifiable parameters. We then propose a flexible Bayesian nonparametric prior, the enriched dependent Dirichlet process, specifically designed for joint modeling of recurrent and terminal events, addressing a limitation where standard Dirichlet process priors create random partitions dominated by recurrent events, yielding poor predictive performance for terminal events. Simulations are carried out to show that our method has superior performance compared to existing methods. We apply the proposed new Bayesian nonparametric methods to infer the causal effect of a structured exercise program on rehospitalizations, which are subject to truncation by death.

stat.ME↗

Who's Winning? Clarifying Estimands Based on Win Statistics in Cluster Randomized Trials

Treatment effect estimands based on win statistics, including the win ratio, win odds, and win difference are increasingly popular targets for summarizing endpoints in clinical trials. Such win estimands offer an intuitive approach for prioritizing outcomes by clinical importance. The implementation and interpretation of win estimands is complicated in cluster randomized trials (CRTs), where researchers can target fundamentally different estimands on the individual-level or cluster-level. We numerically demonstrate that individual-pair and cluster-pair win estimands can substantially differ when cluster size is informative: where outcomes and/or treatment effects depend on cluster size. With such informative cluster sizes, individual-pair and cluster-pair win estimands can even yield opposite conclusions regarding treatment benefit. We describe consistent estimators for individual-pair and cluster-pair win estimands and propose a leave-one-cluster-out jackknife variance estimator for inference. Despite being consistent, our simulations highlight that some caution is needed when implementing individual-pair win estimators due to finite-sample bias. In contrast, cluster-pair win estimators are unbiased for their respective targets. Altogether, careful specification of the target estimand is essential when applying win estimators in CRTs. Failure to clearly define whether individual-pair or cluster-pair win estimands are of primary interest may result in answering a dramatically different question than intended.

stat.ME↗

A Bayesian approach to the survivor average causal effect in cluster-randomized crossover trials

In cluster-randomized crossover (CRXO) trials, groups of individuals are randomly assigned to two or more sequences of alternating treatments. Since clusters serve as their own control, the CRXO design is typically more statistically efficient than the usual parallel-arm design. CRXO trials are increasingly popular in many areas of health research where the number of available clusters is limited. Further, in trials among severely ill patients, researchers often want to assess the effect of treatments on secondary non-terminal outcomes, but frequently in these studies, there are patients who do not survive to have these measurements fully recorded. In this paper, we provide a causal inference framework and treatment effect estimation methods for addressing truncation by death in the setting of CRXO trials. We target the survivor average causal effect (SACE) estimand, a well-defined subgroup treatment effect obtained via principal stratification. We propose novel structural and standard modeling assumptions that enable estimating the SACE within a Bayesian paradigm. We evaluate the small-sample performance of our proposed Bayesian approach for estimation of the SACE in CRXO trial settings via simulation studies. We apply our methods to a previously conducted two-period cross-sectional CRXO study examining the impact of proton pump inhibitors compared to histamine-2 receptor blockers on length of hospitalization among adults requiring invasive mechanical ventilation.

stat.ME↗

Uncovering Treatment Effect Heterogeneity in Pragmatic Gerontology Trials

Detecting heterogeneity in treatment response enriches the interpretation of gerontologic trials. In aging research, estimating the effect of the intervention on clinically meaningful outcomes faces analytical challenges when it is truncated by death. For example, in the Whole Systems Demonstrator trial, a large cluster-randomized study evaluating telecare among older adults, the overall effect of the intervention on quality of life was found to be null. However, this marginal intervention estimate obscures potential heterogeneity of individuals responding to the intervention, particularly among those who survive to the end of follow-up. To explore this heterogeneity, we adopt a causal framework grounded in principal stratification, targeting the Survivor Average Causal Effect (SACE)-the treatment effect among "always-survivors," or those who would survive regardless of treatment assignment. We extend this framework using Bayesian Additive Regression Trees (BART), a nonparametric machine learning method, to flexibly model both latent principal strata and stratum-specific potential outcomes. This enables the estimation of the Conditional SACE (CSACE), allowing us to uncover variation in treatment effects across subgroups defined by baseline characteristics. Our analysis reveals that despite the null average effect, some subgroups experience distinct quality of life benefits (or lack thereof) from telecare, highlighting opportunities for more personalized intervention strategies. This study demonstrates how embedding machine learning methods, such as BART, within a principled causal inference framework can offer deeper insights into trial data with complex features including truncation by death and clustering-key considerations in analyzing pragmatic gerontology trials.

stat.AP↗

On Anticipation Effect in Stepped Wedge Cluster Randomized Trials

In stepped wedge cluster randomized trials (SW-CRTs), the intervention is rolled out to clusters over multiple periods. A standard approach for analyzing SW-CRTs utilizes the linear mixed model, where the treatment effect is only present after the treatment adoption, under the assumption of no anticipation. This assumption, however, may not always hold in practice because stakeholders, providers, or individuals who are aware of the treatment adoption timing (especially when blinding is challenging or infeasible) can inadvertently change their behaviors in anticipation of the forthcoming intervention. We provide an analytical framework to address the anticipation effect in SW-CRTs and study its impact. We derive expectations of the estimators based on a collection of linear mixed models and demonstrate that when the anticipation effect is ignored, these estimators give biased estimates of the treatment effect. We also provide updated sample size formulas that explicitly account for anticipation effects, exposure-time heterogeneity, or both in SW-CRTs and illustrate their impact on study power. Through simulation studies and empirical analyses, we compare the treatment effect estimators with and without adjusting for anticipation, and provide some practical considerations.

stat.ME↗

Evaluating Informative Cluster Size in Cluster Randomized Trials

In cluster randomized trials, the average treatment effect among individuals (i-ATE) can be different from the cluster average treatment effect (c-ATE) when informative cluster size is present, i.e., when treatment effects or participant outcomes depend on cluster size. In such scenarios, mixed-effects models and generalized estimating equations (GEEs) with exchangeable correlation structure are biased for both the i-ATE and c-ATE estimands, whereas GEEs with an independence correlation structure or analyses of cluster-level summaries are recommended in practice. However, when cluster size is non-informative, mixed-effects models and GEEs with exchangeable correlation structure can provide unbiased estimation and notable efficiency gains over other methods. Thus, hypothesis tests for informative cluster size would be useful to assess this key phenomenon under cluster randomization. In this work, we develop model-based, model-assisted, and randomization-based tests for informative cluster size in cluster randomized trials. We construct simulation studies to examine the operating characteristics of these tests, show they have appropriate Type I error control and meaningful power, and contrast them to existing model-based tests used in the observational study setting. The proposed tests are then applied to data from a recent cluster randomized trial, and practical recommendations for using these tests are discussed.

stat.ME↗

Semiparametric principal stratification analysis beyond monotonicity

Intercurrent events, common in clinical trials and observational studies, affect the existence or interpretation of final outcomes. Principal stratification addresses this challenge by defining local average treatment effect estimands within subpopulations, but often relies on restrictive assumptions such as monotonicity and counterfactual intermediate independence. To overcome these limitations, we propose a semiparametric framework for principal stratification analysis leveraging a margin-free, conditional odds ratio sensitivity parameter. Under principal ignorability, we derive nonparametric identification formulas and efficient estimation methods, including a conditionally doubly robust parametric estimator and a debiased machine learning estimator with data-adaptive nuisance learners. Our simulations show that incorrectly assuming monotonicity can frequently lead to biased inference, but incorrectly assuming non-monotonicity when monotonicity holds may maintain approximately valid inference. We demonstrate our methods in the context of a critical care trial, where monotonicity is unlikely to be valid.

stat.ME↗

Doubly robust estimation and sensitivity analysis with outcomes truncated by death in multi-arm clinical trials

In clinical trials, the observation of participant outcomes may frequently be hindered by death, leading to ambiguity in defining a scientifically meaningful final outcome for those who die. Principal stratification methods are valuable tools for addressing the average causal effect among always-survivors, i.e., the average treatment effect among a subpopulation defined as those who would survive regardless of treatment assignment. Although robust methods for the truncation-by-death problem in two-arm clinical trials have been previously studied, its expansion to multi-arm clinical trials remains elusive. In this article, we study the identification of a class of survivor average causal effect estimands with multiple treatments under monotonicity and principal ignorability, and first propose simple weighting and regression approaches for point estimation. As a further improvement, we derive the efficient influence function to motivate doubly robust estimators for the survivor average causal effects in multi-arm clinical trials. We also propose sensitivity methods under violations of key causal assumptions. Extensive simulations are conducted to investigate the finite-sample performance of the proposed methods against the existing methods, and a real data example is used to illustrate how to operationalize the proposed estimators and the sensitivity methods in practice.

stat.ME↗

Bayesian inference for cluster-randomized trials with multivariate outcomes subject to both truncation by death and missingness

Cluster-randomized trials (CRTs) on fragile populations frequently encounter complex attrition problems where the reasons for missing outcomes can be heterogeneous, with participants who are known alive, known to have died, or with unknown survival status, and with complex and distinct missing data mechanisms for each group. Although existing methods have been developed to address death truncation in CRTs, no existing methods can jointly accommodate participants who drop out for reasons unrelated to mortality or serious illnesses, or those with an unknown survival status. This paper proposes a Bayesian framework for estimating survivor average causal effects in CRTs while accounting for different types of missingness. Our approach uses a multivariate outcome that jointly estimates the causal effects, and in the posterior estimates, we distinguish the individual-level and the cluster-level survivor average causal effect. We perform simulation studies to evaluate the performance of our model and found low bias and high coverage on key parameters across several different scenarios. We use data from a geriatric CRT to illustrate the use of our model. Although our illustration focuses on the case of a bivariate continuous outcome, our model is straightforwardly extended to accommodate more than two endpoints as well as other types of endpoints (e.g., binary). Thus, this work provides a general modeling framework for handling complex missingness in CRTs and can be applied to a wide range of settings with aging and palliative care populations.

stat.ME↗

What is estimated in cluster randomized crossover trials with informative sizes? -- A survey of estimands and common estimators

In the analysis of cluster randomized trials (CRTs), previous work has defined two meaningful estimands: the individual-average treatment effect (iATE) and cluster-average treatment effect (cATE) estimand, to address individual and cluster-level hypotheses. In multi-period CRT designs, such as the cluster randomized crossover (CRXO) trial, additional weighted average treatment effect estimands help fully reflect the longitudinal nature of these trial designs, namely the cluster-period-average treatment effect (cpATE) and period-average treatment effect (pATE). We define different forms of informative sizes, where the treatment effects vary according to cluster, period, and/or cluster-period sizes, which subsequently cause these estimands to differ in magnitude. Under such conditions, we demonstrate which of the unweighted, inverse cluster-period size weighted, inverse cluster size weighted, and inverse period size weighted: (i.) independence estimating equation, (ii.) fixed effects model, (iii.) exchangeable mixed effects model, and (iv.) nested exchangeable mixed effects model treatment effect estimators are consistent for the aforementioned estimands in 2-period cross-sectional CRXO designs with continuous outcomes. We report a simulation study and conclude with a reanalysis of a CRXO trial testing different treatments on hospital length of stay among patients receiving invasive mechanical ventilation. Notably, with informative sizes, the unweighted and weighted nested exchangeable mixed effects model estimators are not consistent for any meaningful estimand and can yield biased results. In contrast, the unweighted and weighted independence estimating equation, and under specific scenarios, the fixed effects model and exchangeable mixed effects model, can yield consistent and empirically unbiased estimators for meaningful estimands in 2-period CRXO trials.

stat.ME↗

Time-varying treatment effect models in stepped-wedge cluster-randomized trials with multiple interventions

The traditional model specification of stepped-wedge cluster-randomized trials assumes a homogeneous treatment effect across time while adjusting for fixed-time effects. However, when treatment effects vary over time, the constant effect estimator may be biased. In the general setting of stepped-wedge cluster-randomized trials with multiple interventions, we derive the expected value of the constant effect estimator when the true treatment effects depend on exposure time periods. Applying this result to concurrent and factorial stepped wedge designs, we show that the estimator represents a weighted average of exposure-time-specific treatment effects, with weights that are not necessarily uniform across exposure periods. Extensive simulation studies reveal that ignoring time heterogeneity can result in biased estimates and poor coverage of the average treatment effect. In this study, we examine two models designed to accommodate multiple interventions with time-varying treatment effects: (1) a time-varying fixed treatment effect model, which allows treatment effects to vary by exposure time but remain fixed for each time point, and (2) a random treatment effect model, where the time-varying treatment effects are modeled as random deviations from an overall mean. In the simulations considered in this study, concurrent designs generally achieve higher power than factorial designs under a time-varying fixed treatment effect model, though the differences are modest. Finally, we apply the constant effect model and both time-varying treatment effect models to data from the Prognosticating Outcomes and Nudging Decisions in the Electronic Health Record (PONDER) trial. All three models indicate a lack of treatment effect for either intervention, though they differ in the precision of their estimates, likely due to variations in modeling assumptions.

stat.ME↗