SearcharxivSearch

arXiv subjects

James P. Hughes

Publications and source records attributed to James P. Hughes.

7 recordsLinked to original sources

Robust and Efficient Semiparametric Inference for the Stepped Wedge Design

Stepped wedge designs (SWDs) are increasingly used to evaluate longitudinal cluster-level interventions but pose substantial challenges for valid inference. Because crossover times are randomized, intervention effects are intrinsically confounded with secular time trends, while heterogeneity across clusters, complex correlation structures, baseline covariate imbalances, and small numbers of clusters further complicate inference. We propose a unified semiparametric framework for estimating possibly time-varying intervention effects in SWDs. Under a semiparametric model on treatment contrast, we develop a nonstandard semiparametric efficiency theory that accommodates correlated observations within clusters, varying cluster-period sizes, and weakly dependent treatment assignments. The resulting estimator is consistent and asymptotically normal even under misspecified covariance structure and control cluster-period means, and is efficient when both are correctly specified. To enable inference with few clusters, we exploit the permutation structure of treatment assignment to propose a standard error estimator that reflects finite-sample variability, with a leave-one-out correction to reduce plug-in bias. The framework also allows incorporation of effect modification and adjustment for imbalanced precision variables through design-based adjustment or double adjustment that additionally incorporates an outcome-based component. Simulations and application to a public health trial demonstrate the robustness and efficiency of the proposed method relative to standard approaches.

stat.ME

Which Small-Sample Correction Should Be Used When Analyzing Stepped-Wedge Designs with Time-Varying Treatment Effects?

Stepped-wedge cluster randomized trials (SW-CRTs) evaluate interventions rolled out across clusters over time. Standard analyses typically use immediate-treatment (IT) models, which assume effects begin at crossover and remain constant thereafter. When effects vary with exposure duration, IT models may misrepresent target effects. Exposure-time indicator (ETI) models address this by allowing treatment effects to differ by time since exposure and by targeting the time-averaged treatment effect (TATE) and long-term effect (LTE). Like IT models, ETI models require specification of a random-effects structure, which is often misspecified, and the performance of robust variance estimators (RVEs) in this setting is not well understood. We review RVEs for ETI models and evaluate them in simulation studies with continuous and binary outcomes under correctly specified (binary only) and misspecified random-effects structures. We compare the classic sandwich, Kauermann-Carroll (KC), Mancl-DeRouen (MD), and Morel-Bokossa-Neerchal (MBN) estimators for inference on the TATE and LTE. Our simulations show that under misspecified random-effects structures, model-based standard errors (SE) produced undercoverage, whereas RVEs improved performance. For continuous outcomes, MD with a t-distribution and degrees of freedom equal to the number of clusters minus two gave the most consistent coverage probabilities. For binary outcomes, MBN was the only consistently reliable option. MD, however, could be unstable in one-cluster-per-sequence designs because of data sparsity. Across scenarios, both model-based SE and RVE for LTE were unstable, indicating that greater caution is needed when targeting LTE under ETI models.

stat.ME

Factors affecting power in stepped wedge trials when the treatment effect varies with time

Stepped wedge cluster randomized trials (SW-CRTs) have historically been analyzed using immediate treatment (IT) models, which assume the effect of the treatment is immediate after treatment initiation and subsequently remains constant over time. However, recent research has shown that this assumption can lead to severely misleading results if treatment effects vary with exposure time, i.e. time since the intervention started. Models that account for time-varying treatment effects, such as the exposure time indicator (ETI) model, allow researchers to target estimands such as the time-averaged treatment effect (TATE) over an interval of exposure time, or the point treatment effect (PTE) representing a treatment contrast at one time point. However, this increased flexibility results in reduced power. In this paper, we use public power calculation software and simulation to characterize factors affecting SW-CRT power. Key elements include choice of estimand, study design considerations, and analysis model selection. or common SW-CRT designs, the sample size (clusters per sequence or individuals per cluster-period) must be increased substantially, commonly by a factor of 1.5 to 3, but often by much more, to maintain 90\% power when switching from an IT model to an ETI model (targeting the TATE over the study). However, the inflation factor is lower for TATE estimands over shorter periods that exclude longer exposure times. In general, SW-CRT designs (including the "staircase" variant) have much greater power for estimating "short-term effects" relative to "long-term effects". For an ETI model targeting a TATE estimand, substantial power can be gained by adding time points to the start of the study or increasing baseline sample size, but surprisingly little power is gained from adding time points to the end of the study. More restrictive choices for modeling the exposure... [truncated]

stat.ME

A discrete-time survival model to handle interval-censored covariates, with applications to HIV cohort studies

Methods are lacking to handle the problem of survival analysis in the presence of an interval-censored covariate, specifically the case in which the conditional hazard of the primary event of interest depends on the occurrence of a secondary event, the observation time of which is subject to interval censoring. We propose and study a flexible class of discrete-time parametric survival models that handle the censoring problem through simultaneous modeling of the interval-censored secondary event, the outcome, and the censoring mechanism. We apply this model to the research question that motivated the methodology, estimating the effect of HIV status on all-cause mortality in a prospective cohort study in South Africa. Our model has applicability for many open questions, including estimating the impact of policy decisions on population level HIV-related outcomes and determining causes of morbidity and mortality for which the HIV positive population may be at increased risk. Examples include determining how the large-scale transition from efavirenz-based to dolutegravir-based first-line ART impacted mortality for people living with HIV and determining whether HIV status is associated with increased risk of stroke, diabetes, hypertension, and other non-communicable diseases.

stat.ME

Adjusting for Incomplete Baseline Covariates in Randomized Controlled Trials: A Cross-World Imputation Framework

In randomized controlled trials, adjusting for baseline covariates is often applied to improve the precision of treatment effect estimation. However, missingness in covariates is common. Recently, Zhao & Ding (2022) studied two simple strategies, the single imputation method and missingness indicator method (MIM), to deal with missing covariates, and showed that both methods can provide efficiency gain. To better understand and compare these two strategies, we propose and investigate a novel imputation framework termed cross-world imputation (CWI), which includes single imputation and MIM as special cases. Through the lens of CWI, we show that MIM implicitly searches for the optimal CWI values and thus achieves optimal efficiency. We also derive conditions under which the single imputation method, by searching for the optimal single imputation values, can achieve the same efficiency as the MIM.

stat.ME

Analysis of stepped wedge cluster randomized trials in the presence of a time-varying treatment effect

Stepped wedge cluster randomized controlled trials are typically analyzed using models that assume the full effect of the treatment is achieved instantaneously. We provide an analytical framework for scenarios in which the treatment effect varies as a function of exposure time (time since the start of treatment) and define the "effect curve" as the magnitude of the treatment effect on the linear predictor scale as a function of exposure time. The "time-averaged treatment effect", (TATE) and "long-term treatment effect" (LTE) are summaries of this curve. We analytically derive the expectation of the estimator resulting from a model that assumes an immediate treatment effect and show that it can be expressed as a weighted sum of the time-specific treatment effects corresponding to the observed exposure times. Surprisingly, although the weights sum to one, some of the weights can be negative. This implies that the estimator may be severely misleading and can even converge to a value of the opposite sign of the true TATE or LTE. We describe several models that can be used to simultaneously estimate the entire effect curve, the TATE, and the LTE, some of which make assumptions about the shape of the effect curve. We evaluate these models in a simulation study to examine the operating characteristics of the resulting estimators and apply them to two real datasets.

stat.ME

Sample Size Calculation for Active-Arm Trial with Counterfactual Incidence Based on Recency Assay

The past decade has seen tremendous progress in the development of biomedical agents that are effective as pre-exposure prophylaxis (PrEP) for HIV prevention. To expand the choice of products and delivery methods, new medications and delivery methods are under development. Future trials of non-inferiority, given the high efficacy of ARV-based PrEP products as they become current or future standard of care, would require a large number of participants and long follow-up time that may not be feasible. This motivates the construction of a counterfactual estimate that approximates incidence for a randomized concurrent control group receiving no PrEP. We propose an approach that is to enroll a cohort of prospective PrEP users and augment screening for HIV with laboratory markers of duration of HIV infection to indicate recent infections. We discuss the assumptions under which these data would yield an estimate of the counterfactual HIV incidence and develop sample size and power calculations for comparisons to incidence observed on an investigational PrEP agent.

stat.ME