SearcharxivSearch

arXiv subjects

Marcel Wolbers

Publications and source records attributed to Marcel Wolbers.

13 recordsLinked to original sources

Unified implementation and comparison of Bayesian shrinkage methods for treatment effect estimation in subgroups

Evaluating treatment effect heterogeneity across patient subgroups is a fundamental aspect of clinical trial analysis. These analyses have inherent limitations due to small sample sizes and the substantial number of subgroups investigated. There is a tendency to focus on extreme estimates, which may reflect random variation rather than true effects, potentially leading to spurious clinical conclusions. Statisticians in regulatory agencies and pharmaceutical companies have begun considering shrinkage methods grounded in Bayesian theory. These methods incorporate priors on treatment effect heterogeneity, which shrink subgroup estimates towards the overall treatment effect. Various shrinkage estimators have been proposed, yet it remains unclear which perform best. This work provides a unified presentation and software implementation of shrinkage methods. It also provides simulation comparisons of one-way and global shrinkage methods for two simulation set-ups. One-way models fit a separate shrinkage model for each subgrouping variable while global models include all subgroup indicators. Both can derive standardized subgroup-specific treatment effects. Across all simulation scenarios, shrinkage methods outperformed the standard subgroup estimator in terms of mean squared error. They were also more efficient in identifying a non-efficacious subgroup. Global shrinkage models tended to have smaller mean squared error and less dependence on hyperprior parameters than one-way models, but also exhibited slightly larger bias and worse frequentist coverage of credible intervals. For both models, hyperprior choices anchored in trial assumptions about the anticipated overall treatment effect size performed well. We conclude that some shrinkage is preferable to none and advocate routine inclusion of shrunken estimates in clinical forest plots to facilitate robust decision-making.

stat.ME

Communicating results in trials with multiple hypotheses or adaptive design features

Over time, clinical trials have increasingly incorporated complex design and analysis elements such as interim analyses, adaptations, multiple endpoints, and sophisticated multiplicity schemes for multiple endpoints and/or treatment arms following the paradigm of frequentist inference. In frequentist clinical trials multiplicity can come from (at least) four sources: multiple looks at the data, multiple endpoints, multiple populations, or multiple treatment comparisons. Normally, Type 1 error control across the multiple hypotheses is implemented to control chance of false positive decisions. To achieve this advanced techniques such as adaptive designs or graphical multiple testing procedures have been developed and are used in the design of clinical trials. However, these methods focus on hypothesis testing while subsequent estimation remains crucial to allow for a benefit-risk assessment and further use of the results by various stakeholders. Through examples, we illustrate challenges in estimation and transparent communication. In general, there are no simple solutions to this conceptual and communicational challenge. The purpose of this paper is to generate awareness of these issues and initiate a discussion about how to address them moving forward.

stat.ME

Bayesian analysis of the causal reference-based model for missing data in clinical trials, accommodating partially observed post-intercurrent event data

When treatment policy estimands are of interest, clinical trials often attempt to collect patient data after intercurrent events (ICEs), although such data are often limited. Retrieved dropout imputation methods, which use pre-ICE and available post-ICE data to impute missing post-ICE outcomes, are commonly applied but often yield treatment effect estimates with large standard errors (SEs) and may encounter convergence issues when post-ICE data are sparse. Reference-based imputation methods are also used, but they rely on strong assumptions about post-ICE outcomes, which can lead to biased estimates if these assumptions are incorrect. To address these limitations, we previously proposed the reference-based Bayesian causal model (BCM), which incorporates a prior on the maintained effect parameter to reflect uncertainty in reference-based assumptions for missing post-ICE data. Our earlier work assumed no post-ICE data were observed. Here, we extend the BCM to incorporate available post-ICE outcomes, providing an approach that mitigates limitations of both retrieved-dropout and standard reference-based methods. We propose both a fully Bayesian model and an imputation-based approach. A simulation study was conducted to evaluate the frequentist properties of the proposed methods in settings with partially observed post-ICE data and to compare performance with existing approaches. Retrieved-dropout methods produced higher estimated SEs than the BCM, particularly when post-ICE data were sparse. Under the BCM, treatment effect SEs increased as post-ICE data became more limited for both modelling approaches. Importantly, this increase can be controlled through the prior variance of the maintained effect parameter, with more informative priors stabilising estimation when post-ICE data are scarce.

stat.ME

"6 choose 4": A framework to understand and facilitate discussion of strategies for overall survival safety monitoring

Advances in anticancer therapies have significantly contributed to declining death rates in certain disease and clinical settings. However, they have also made it difficult to power a clinical trial in these settings with overall survival (OS) as the primary efficacy endpoint. Therefore, two approaches have been recently proposed for the pre-specified analysis of OS as a safety endpoint (Fleming et al., 2024; Rodriguez et al., 2024). In this paper, we provide a simple, unifying framework that includes the aforementioned approaches (and a couple others) as special cases. By highlighting each approach's focus, priority, tolerance for risk, and strengths or challenges for practical implementation, this framework can help to facilitate discussions between stakeholders on "fit-for-purpose OS data collection and assessment of harm" (American Association for Cancer Research, 2024). We apply this framework to a real clinical trial in large B-cell lymphoma to illustrate its application and value. Several recommendations and open questions are also raised.

stat.ME

Bayesian analysis of the causal reference-based model for missing data in clinical trials

The statistical analysis of clinical trials is often complicated by missing data. Patients sometimes experience intercurrent events (ICEs), which usually (although not always) lead to missing subsequent outcome measurements for such individuals. The reference-based imputation methods were proposed by Carpenter et al. (2013) and have been commonly adopted for handling missing data due to ICEs when estimating treatment policy strategy estimands. Conventionally, the variance for reference-based estimators was obtained using Rubin's rules. However, Rubin's rules variance estimator is biased compared to the repeated sampling variance of the point estimator, due to uncongeniality. Repeated sampling variance estimators were proposed as an alternative to variance estimation for reference-based estimators. However, these have the property that they decrease as the proportion of ICEs increases. White et al. (2019) introduced a causal model incorporating the concept of a 'maintained treatment effect' following the occurrence of ICEs and showed that this causal model included common reference-based estimators as special cases. Building on this framework, we propose introducing a prior distribution for the maintained effect parameter to account for uncertainty in this assumption. Our approach provides inference for reference-based estimators that explicitly reflects our uncertainty about how much treatment effects are maintained after the occurrence of ICEs. In trials where no or little post-ICE data are observed, our proposed Bayesian reference-based causal model approach can be used to estimate the treatment policy treatment effect, incorporating uncertainty about the reference-based assumption. We compare the frequentist properties of this approach with existing reference-based methods through simulations and by application to an antidepressant trial.

stat.ME

Balancing events, not patients, maximizes power of the logrank test: and other insights on unequal randomization in survival trials

We revisit the question of what randomization ratio (RR) maximizes power of the logrank test in event-driven survival trials under proportional hazards (PH). By comparing three approximations of the logrank test (Schoenfeld, Freedman, Rubinstein) to empirical simulations, we find that the RR that maximizes power is the RR that balances number of events across treatment arms at the end of the trial. This contradicts the common misconception implied by Schoenfeld's approximation that 1:1 randomization maximizes power. Besides power, we consider other factors that might influence the choice of RR (accrual, trial duration, sample size, etc.). We perform simulations to better understand how unequal randomization might impact these factors in practice. Altogether, we derive 6 insights to guide statisticians in the design of survival trials considering unequal randomization.

stat.ME

Using shrinkage methods to estimate treatment effects in overlapping subgroups in randomized clinical trials with a time-to-event endpoint

In randomized controlled trials, forest plots are frequently used to investigate the homogeneity of treatment effect estimates in subgroups. However, the interpretation of subgroup-specific treatment effect estimates requires great care due to the smaller sample size of subgroups and the large number of investigated subgroups. Bayesian shrinkage methods have been proposed to address these issues, but they often focus on disjoint subgroups while subgroups displayed in forest plots are overlapping, i.e., each subject appears in multiple subgroups. In our approach, we first build a flexible Cox model based on all available observations, including categorical covariates that identify the subgroups of interest and their interactions with the treatment group variable. We explore both penalized partial likelihood estimation with a lasso or ridge penalty for treatment-by-covariate interaction terms, and Bayesian estimation with a regularized horseshoe prior. One advantage of the Bayesian approach is the ability to derive credible intervals for shrunken subgroup-specific estimates. In a second step, the Cox model is marginalized to obtain treatment effect estimates for all subgroups. We illustrate these methods using data from a randomized clinical trial in follicular lymphoma and evaluate their properties in a simulation study. In all simulation scenarios, the overall mean-squared error is substantially smaller for penalized and shrinkage estimators compared to the standard subgroup-specific treatment effect estimator but leads to some bias for heterogeneous subgroups. We recommend that subgroup-specific estimators, which are typically displayed in forest plots, are more routinely complemented by treatment effect estimators based on shrinkage methods. The proposed methods are implemented in the R package bonsaiforest.

stat.ME

Estimation methods for estimands using the treatment policy strategy; a simulation study based on the PIONEER 1 Trial

Estimands using the treatment policy strategy for addressing intercurrent events are common in Phase III clinical trials. One estimation approach for this strategy is retrieved dropout whereby observed data following an intercurrent event are used to multiply impute missing data. However, such methods have had issues with variance inflation and model fitting due to data sparsity. This paper introduces likelihood-based versions of these approaches, investigating and comparing their statistical properties to the existing retrieved dropout approaches, simpler analysis models and reference-based multiple imputation. We use a simulation based upon the data from the PIONEER 1 Phase III clinical trial in Type II diabetics to present complex and relevant estimation challenges. The likelihood-based methods display similar statistical properties to their multiple imputation equivalents, but all retrieved dropout approaches suffer from high variance. Retrieved dropout approaches appear less biased than reference-based approaches, resulting in a bias-variance trade-off, but we conclude that the large degree of variance inflation is often more problematic than the bias. Therefore, only the simpler retrieved dropout models appear appropriate as a primary analysis in a clinical trial, and only where it is believed most data following intercurrent events will be observed. The jump-to-reference approach may represent a more promising estimation approach for symptomatic treatments due to its relatively high power and ability to fit in the presence of much missing data, despite its strong assumptions and tendency towards conservative bias. More research is needed to further develop how to estimate the treatment effect for a treatment policy strategy.

stat.AP

Standard and reference-based conditional mean imputation

Clinical trials with longitudinal outcomes typically include missing data due to missed assessments or structural missingness of outcomes after intercurrent events handled with a hypothetical strategy. Approaches based on Bayesian random multiple imputation and Rubin's rule for pooling results across multiple imputed datasets are increasingly used in order to align the analysis of these trials with the targeted estimand. We propose and justify deterministic conditional mean imputation combined with the jackknife for inference as an alternative approach. The method is applicable to imputations under a missing-at-random assumption as well as for reference-based imputation approaches. In an application and a simulation study, we demonstrate that it provides consistent treatment effect estimates with the Bayesian approach and reliable frequentist inference with accurate standard error estimation and type I error control. A further advantage of the method is that it does not rely on random sampling and is therefore replicable and unaffected by Monte Carlo error.

stat.ME

A Comparison of Estimand and Estimation Strategies for Clinical Trials in Early Parkinson's Disease

Parkinson's disease (PD) is a chronic, degenerative neurological disorder. PD cannot be prevented, slowed or cured as of today but highly effective symptomatic treatments are available. We consider relevant estimands and treatment effect estimators for randomized trials of a novel treatment which aims to slow down disease progression versus placebo in early, untreated PD. A commonly used endpoint in PD trials is the MDS-Unified Parkinson's Disease Rating Scale (MDS-UPDRS), which is longitudinally assessed at scheduled visits. The most important intercurrent events (ICEs) which affect the interpretation of the MDS-UPDRS are study treatment discontinuations and initiations of symptomatic treatment. Different estimand strategies are discussed and hypothetical or treatment policy strategies, respectively, for different types of ICEs seem most appropriate in this context. Several estimators based on multiple imputation which target these estimands are proposed and compared in terms of bias, mean-squared error, and power in a simulation study. The investigated estimators include methods based on a missing-at-random (MAR) assumption, with and without the inclusion of time-varying ICE-indicators, as well as reference-based imputation methods. Simulation parameters are motivated by data analyses of a cohort study from the Parkinson's Progression Markers Initiative (PPMI).

stat.AP

Principal Stratum Strategy: Potential Role in Drug Development

A randomized trial allows estimation of the causal effect of an intervention compared to a control in the overall population and in subpopulations defined by baseline characteristics. Often, however, clinical questions also arise regarding the treatment effect in subpopulations of patients, which would experience clinical or disease related events post-randomization. Events that occur after treatment initiation and potentially affect the interpretation or the existence of the measurements are called {\it intercurrent events} in the ICH E9(R1) guideline. If the intercurrent event is a consequence of treatment, randomization alone is no longer sufficient to meaningfully estimate the treatment effect. Analyses comparing the subgroups of patients without the intercurrent events for intervention and control will not estimate a causal effect. This is well known, but post-hoc analyses of this kind are commonly performed in drug development. An alternative approach is the principal stratum strategy, which classifies subjects according to their potential occurrence of an intercurrent event on both study arms. We illustrate with examples that questions formulated through principal strata occur naturally in drug development and argue that approaching these questions with the ICH E9(R1) estimand framework has the potential to lead to more transparent assumptions as well as more adequate analyses and conclusions. In addition, we provide an overview of assumptions required for estimation of effects in principal strata. Most of these assumptions are unverifiable and should hence be based on solid scientific understanding. Sensitivity analyses are needed to assess robustness of conclusions.

stat.AP

A pragmatic adaptive enrichment design for selecting the right target population for cancer immunotherapies

One of the challenges in the design of confirmatory trials is to deal with uncertainties regarding the optimal target population for a novel drug. Adaptive enrichment designs (AED) which allow for a data-driven selection of one or more pre-specified biomarker subpopulations at an interim analysis have been proposed in this setting but practical case studies of AEDs are still relatively rare. We present the design of an AED with a binary endpoint in the highly dynamic setting of cancer immunotherapy. The trial was initiated as a conventional trial in early triple-negative breast cancer but amended to an AED based on emerging data external to the trial suggesting that PD-L1 status could be a predictive biomarker. Operating characteristics are discussed including the concept of a minimal detectable difference, that is, the smallest observed treatment effect that would lead to a statistically significant result in at least one of the target populations at the interim or the final analysis, respectively, in the setting of AED.

stat.AP

Statistical Issues and Recommendations for Clinical Trials Conducted During the COVID-19 Pandemic

The COVID-19 pandemic has had and continues to have major impacts on planned and ongoing clinical trials. Its effects on trial data create multiple potential statistical issues. The scale of impact is unprecedented, but when viewed individually, many of the issues are well defined and feasible to address. A number of strategies and recommendations are put forward to assess and address issues related to estimands, missing data, validity and modifications of statistical analysis methods, need for additional analyses, ability to meet objectives and overall trial interpretability.

q-bio.OT