SearcharxivSearch

arXiv subjects

Laurent Billot

Publications and source records attributed to Laurent Billot.

4 recordsLinked to original sources

Rapid evaluation and calibration of Bayesian group sequential designs via conjugate-mixture semi-simulation

Bayesian group sequential designs (GSDs) extend frequentist GSDs with interpretable decision-making and external evidence borrowing, but their use is limited by the computational burden of design-stage operating-characteristic evaluation. Conventional methods simulate virtual trials with Markov chain Monte Carlo or approximate analytical posterior updates at each interim look, making joint calibration of decision thresholds and design skeletons impractical on commodity hardware. Here we introduce a semi-simulation framework with two innovations. First, finite conjugate-mixture priors replace the posterior computation for each look (``per-look'') with closed-form conjugate updates and low-dimensional numerical integration for decision-rule tail probabilities. Second, a precomputation strategy caches per-look posterior tail probabilities from a single Monte Carlo pass at the union of all candidate analysis times, and each design in the calibration grid is evaluated against the same cache by a sub-second sweep, with no further simulation cost. The framework supports posterior-probability decision rules with multiple efficacy and futility criteria under either binding or non-binding futility, and derives closed-form per-look updates for binary, continuous, count and time-to-event endpoints, with benchmarking here focused on the binary endpoint. When applied to re-design the ADRENAL trial, with up to nine analyses, the framework reproduces the operating characteristics of BATSS and adaptr within Monte Carlo error while running, per GSD, approximately $7\times$ to $16\times$ faster than adaptr at a matched budget (several hundredfold at the million-trial calibration budget) and $3{,}700\times$ to $6{,}600\times$ faster than BATSS. This brings routine Bayesian GSD calibration within computational reach for confirmatory trials.

stat.ME

Empirical comparison of win ratio and joint frailty models for recurrent event endpoints with applications in oncology and cardiology

Composite endpoints that combine recurrent non-fatal events with a terminal event are increasingly used in randomized clinical trials, yet conventional time-to-first event analyses may obscure clinically relevant information. We compared two statistical frameworks tailored to such endpoints: the joint frailty model (JFM) and the last-event assisted recurrent-event win ratio (LWR). The JFM specifies proportional hazards for the recurrent and terminal events linked through a shared frailty, yielding covariate-adjusted, component-specific hazard ratios that account for informative recurrences and dependence with death. The LWR is a nonparametric, prioritized pairwise comparison that incorporates all observed events over follow-up and summarizes a population-level benefit of treatment while respecting a pre-specified hierarchy between death and recurrences. We first assessed the performance of the methods using simulations that varied both the gamma-frailty variance and the event rates. We next illustrated both approaches using two clinical application examples in oncology and cardiology, highlighting how conclusions depend on whether treatment primarily affects recurrent events, mortality, or both. The JFM provided component-specific estimates, while the LWR led to a summary measure of treatment effect with direction. Power was systematically improved with JFM, which thus appeared as the most reliable approach for inference and sample size estimation. Methodological extensions of the LWR to appropriately handle censoring and to formalize causal estimands remain a promising direction for future research.

stat.ME

From aggressive to conservative early stopping in Bayesian group sequential designs

Group sequential designs (GSDs) are widely used in confirmatory trials to allow interim monitoring while preserving control of the type I error rate. In the frequentist framework, O'Brien-Fleming-type stopping boundaries dominate practice because they impose highly conservative early stopping while allowing more liberal decisions as information accumulates. Bayesian GSDs, in contrast, are most often implemented using fixed posterior probability thresholds applied uniformly at all analyses. While such designs can be calibrated to control the overall type I error rate, they do not penalise early analyses and can therefore lead to substantially more aggressive early stopping. Such behaviour can risk premature conclusions and inflation of treatment effect estimates, raising concerns for confirmatory trials. We introduce two practically implementable refinements that restore conservative early stopping in Bayesian GSDs. The first introduces a two-phase structure for posterior probability thresholds, applying more stringent criteria in the early phase of the trial and relaxing them later to preserve power. The second replaces posterior probability monitoring at interim looks with predictive probability criteria, which naturally account for uncertainty in future data and therefore suppress premature stopping. Both strategies require only one additional tuning parameter and can be efficiently calibrated. In the HYPRESS setting, both approaches achieve higher power than the conventional Bayesian design while producing alpha-spending profiles closely aligned with O'Brien-Fleming-type behaviour at early looks. These refinements provide a principled and tractable way to align Bayesian GSDs with accepted frequentist practice and regulatory expectations, supporting their robust application in confirmatory trials.

stat.ME

Optimal scheduling of interim analyses in group sequential trials

Group sequential designs (GSDs) are well established and the most commonly used adaptive design in confirmatory clinical trials with interim analyses. However, they remain underutilised, and their implementation involves unique theoretical and practical decisions that demand careful consideration to optimise efficiency. A common practice is to schedule interim analyses at equal intervals based on calendar time or accumulated data. While straightforward, this approach does not completely exploit the potential sample size savings achievable with GSDs. To address this challenge, we develop OptimInterim, an R-based tool that can determine the optimal scheduling of interim analyses to minimise the expected sample size under the alternative hypothesis while controlling overall type I and type II errors. Our method accommodates trials with continuous or binary endpoints, allows multiple interim analyses and supports a range of stopping boundaries. Through extensive simulations, we demonstrate that optimally spaced interim analyses can yield substantial savings in expected sample size compared to equally spaced interim analyses, without compromising the maximum sample size, across various endpoint types, effect sizes, error rates and stopping rules. We illustrate its practical utility with two landmark trials evaluating steroid use in septic shock. Notably, for given type I and type II error rates, the optimal scheduling is independent of endpoint types and effect sizes, ensuring broad applicability across a wide range of trial contexts. To facilitate implementation, we offer a ready-to-use reference table of optimal schedules for up to eight interim analyses under commonly used error rates and stopping rules. Access OptimInterim at https://github.com/zhangyi-he/GSD_OptimInterim.

stat.ME