SearcharxivSearch

arXiv subjects

Suzie Cro

Publications and source records attributed to Suzie Cro.

9 recordsLinked to original sources

Rapid evaluation and calibration of Bayesian group sequential designs via conjugate-mixture semi-simulation

Bayesian group sequential designs (GSDs) extend frequentist GSDs with interpretable decision-making and external evidence borrowing, but their use is limited by the computational burden of design-stage operating-characteristic evaluation. Conventional methods simulate virtual trials with Markov chain Monte Carlo or approximate analytical posterior updates at each interim look, making joint calibration of decision thresholds and design skeletons impractical on commodity hardware. Here we introduce a semi-simulation framework with two innovations. First, finite conjugate-mixture priors replace the posterior computation for each look (``per-look'') with closed-form conjugate updates and low-dimensional numerical integration for decision-rule tail probabilities. Second, a precomputation strategy caches per-look posterior tail probabilities from a single Monte Carlo pass at the union of all candidate analysis times, and each design in the calibration grid is evaluated against the same cache by a sub-second sweep, with no further simulation cost. The framework supports posterior-probability decision rules with multiple efficacy and futility criteria under either binding or non-binding futility, and derives closed-form per-look updates for binary, continuous, count and time-to-event endpoints, with benchmarking here focused on the binary endpoint. When applied to re-design the ADRENAL trial, with up to nine analyses, the framework reproduces the operating characteristics of BATSS and adaptr within Monte Carlo error while running, per GSD, approximately $7\times$ to $16\times$ faster than adaptr at a matched budget (several hundredfold at the million-trial calibration budget) and $3{,}700\times$ to $6{,}600\times$ faster than BATSS. This brings routine Bayesian GSD calibration within computational reach for confirmatory trials.

stat.ME

From aggressive to conservative early stopping in Bayesian group sequential designs

Group sequential designs (GSDs) are widely used in confirmatory trials to allow interim monitoring while preserving control of the type I error rate. In the frequentist framework, O'Brien-Fleming-type stopping boundaries dominate practice because they impose highly conservative early stopping while allowing more liberal decisions as information accumulates. Bayesian GSDs, in contrast, are most often implemented using fixed posterior probability thresholds applied uniformly at all analyses. While such designs can be calibrated to control the overall type I error rate, they do not penalise early analyses and can therefore lead to substantially more aggressive early stopping. Such behaviour can risk premature conclusions and inflation of treatment effect estimates, raising concerns for confirmatory trials. We introduce two practically implementable refinements that restore conservative early stopping in Bayesian GSDs. The first introduces a two-phase structure for posterior probability thresholds, applying more stringent criteria in the early phase of the trial and relaxing them later to preserve power. The second replaces posterior probability monitoring at interim looks with predictive probability criteria, which naturally account for uncertainty in future data and therefore suppress premature stopping. Both strategies require only one additional tuning parameter and can be efficiently calibrated. In the HYPRESS setting, both approaches achieve higher power than the conventional Bayesian design while producing alpha-spending profiles closely aligned with O'Brien-Fleming-type behaviour at early looks. These refinements provide a principled and tractable way to align Bayesian GSDs with accepted frequentist practice and regulatory expectations, supporting their robust application in confirmatory trials.

stat.ME

Tools to help patients and other stakeholders' input into choice of estimand and intercurrent event strategy in randomised trials

Estimands can help to clarify the research questions being addressed in randomised trials. Because the choice of estimand can affect how relevant trial results are to patients and other stakeholders, such as clinicians or policymakers, it is important for them to be involved in these decisions. However, there are barriers to having these conversations. For instance, discussions around how intercurrent events should be addressed in the estimand definition typically involve complex concepts as well as technical language. We three tools to facilitate conversations between researchers and patients and other stakeholders about the choice of estimand and intercurrent event strategy: (i) a video explaining the concept of an estimand and the five different ways that intercurrent events can be incorporated into the estimand definition; (ii) an infographic outlining these five strategies; and (iii) an editable PowerPoint slide which can be completed with trial-specific details to facilitate conversations around choice of estimand for a particular trial. These resources can help to start conversations between the trial team and patients and other stakeholders about the best choice of estimand and intercurrent event strategies for a randomised trial.

stat.ME

Optimal scheduling of interim analyses in group sequential trials

Group sequential designs (GSDs) are well established and the most commonly used adaptive design in confirmatory clinical trials with interim analyses. However, they remain underutilised, and their implementation involves unique theoretical and practical decisions that demand careful consideration to optimise efficiency. A common practice is to schedule interim analyses at equal intervals based on calendar time or accumulated data. While straightforward, this approach does not completely exploit the potential sample size savings achievable with GSDs. To address this challenge, we develop OptimInterim, an R-based tool that can determine the optimal scheduling of interim analyses to minimise the expected sample size under the alternative hypothesis while controlling overall type I and type II errors. Our method accommodates trials with continuous or binary endpoints, allows multiple interim analyses and supports a range of stopping boundaries. Through extensive simulations, we demonstrate that optimally spaced interim analyses can yield substantial savings in expected sample size compared to equally spaced interim analyses, without compromising the maximum sample size, across various endpoint types, effect sizes, error rates and stopping rules. We illustrate its practical utility with two landmark trials evaluating steroid use in septic shock. Notably, for given type I and type II error rates, the optimal scheduling is independent of endpoint types and effect sizes, ensuring broad applicability across a wide range of trial contexts. To facilitate implementation, we offer a ready-to-use reference table of optimal schedules for up to eight interim analyses under commonly used error rates and stopping rules. Access OptimInterim at https://github.com/zhangyi-he/GSD_OptimInterim.

stat.ME

Multiple imputation of partially observed data after treatment-withdrawal

The ICH E9(R1) Addendum (International Council for Harmonization 2019) suggests treatment-policy as one of several strategies for addressing intercurrent events such as treatment withdrawal when defining an estimand. This strategy requires the monitoring of patients and collection of primary outcome data following termination of randomized treatment. However, when patients withdraw from a study before nominal completion this creates true missing data complicating the analysis. One possible way forward uses multiple imputation to replace the missing data based on a model for outcome on and off treatment prior to study withdrawal, often referred to as retrieved dropout multiple imputation. This article explores a novel approach to parameterizing this imputation model so that those parameters which may be difficult to estimate have mildly informative Bayesian priors applied during the imputation stage. A core reference-based model is combined with a compliance model, using both on- and off- treatment data to form an extended model for the purposes of imputation. This alleviates the problem of specifying a complex set of analysis rules to accommodate situations where parameters which influence the estimated value are not estimable or are poorly estimated, leading to unrealistically large standard errors in the resulting analysis.

stat.ME

Improving clinical trial interpretation with ACCEPT analyses

Effective decision making from randomised controlled clinical trials relies on robust interpretation of the numerical results. However, the language we use to describe clinical trials can cause confusion both in trial design and in comparing results across trials. ACceptability Curve Estimation using Probability Above Threshold (ACCEPT) aids comparison between trials (even where of different designs) by harmonising reporting of results, acknowledging different interpretations of the results may be valid in different situations, and moving the focus from comparison to a pre-specified value to interpretation of the trial data. ACCEPT can be applied to historical trials or incorporated into statistical analysis plans for future analyses. An online tool enables ACCEPT on up to three trials simultaneously.

stat.AP

Estimands and their Estimators for Clinical Trials Impacted by the COVID-19 Pandemic: A Report from the NISS Ingram Olkin Forum Series on Unplanned Clinical Trial Disruptions

The COVID-19 pandemic continues to affect the conduct of clinical trials globally. Complications may arise from pandemic-related operational challenges such as site closures, travel limitations and interruptions to the supply chain for the investigational product, or from health-related challenges such as COVID-19 infections. Some of these complications lead to unforeseen intercurrent events in the sense that they affect either the interpretation or the existence of the measurements associated with the clinical question of interest. In this article, we demonstrate how the ICH E9(R1) Addendum on estimands and sensitivity analyses provides a rigorous basis to discuss potential pandemic-related trial disruptions and to embed these disruptions in the context of study objectives and design elements. We introduce several hypothetical estimand strategies and review various causal inference and missing data methods, as well as a statistical method that combines unbiased and possibly biased estimators for estimation. To illustrate, we describe the features of a stylized trial, and how it may have been impacted by the pandemic. This stylized trial will then be re-visited by discussing the changes to the estimand and the estimator to account for pandemic disruptions. Finally, we outline considerations for designing future trials in the context of unforeseen disruptions.

stat.ME

How to design a pre-specified statistical analysis approach to limit p-hacking in clinical trials: the Pre-SPEC framework

Results from clinical trials can be susceptible to bias if investigators choose their analysis approach after seeing trial data, as this can allow them to perform multiple analyses and then choose the method that provides the most favourable result (commonly referred to as 'p-hacking'). Pre-specification of the planned analysis approach is essential to help reduce such bias, as it ensures analytical methods are chosen in advance of seeing the trial data. However, pre-specification is only effective if done in a way that does not allow p-hacking. For example, investigators may pre-specify a certain statistical method such as multiple imputation, but give little detail on how it will be implemented. Because there are many different ways to perform multiple imputation, this approach to pre-specification is ineffective, as it still allows investigators to analyse the data in different ways before deciding on a final approach. In this article we describe a five-point framework (the Pre-SPEC framework) for designing a pre-specified analysis approach that does not allow p-hacking. This framework is intended to be used in conjunction with the SPIRIT (Standard Protocol Items: Recommendations for Interventional Trials) statement and other similar guidelines to help investigators design the statistical analysis strategy for the trial's primary outcome in the trial protocol.

stat.ME

Information-Anchored Sensitivity Analysis: Theory and Application

Analysis of longitudinal randomised controlled trials is frequently complicated because patients deviate from the protocol. Where such deviations are relevant for the estimand, we are typically required to make an untestable assumption about post-deviation behaviour in order to perform our primary analysis and estimate the treatment effect. In such settings, it is now widely recognised that we should follow this with sensitivity analyses to explore the robustness of our inferences to alternative assumptions about post-deviation behaviour. Although there has been a lot of work on how to conduct such sensitivity analyses, little attention has been given to the appropriate loss of information due to missing data within sensitivity analysis. We argue more attention needs to be given to this issue, showing it is quite possible for sensitivity analysis to decrease and increase the information about the treatment effect. To address this critical issue, we introduce the concept of information-anchored sensitivity analysis. By this we mean sensitivity analysis in which the proportion of information about the treatment estimate lost due to missing data is the same as the proportion of information about the treatment estimate lost due to missing data in the primary analysis. We argue this forms a transparent, practical starting point for interpretation of sensitivity analysis. We then derive results showing that, for longitudinal continuous data, a broad class of controlled and reference-based sensitivity analyses performed by multiple imputation are information-anchored. We illustrate the theory with simulations and an analysis of a peer review trial, then discuss our work in the context of other recent work in this area. Our results give a theoretical basis for the use of controlled multiple imputation procedures for sensitivity analysis.

stat.ME