SearcharxivSearch

arXiv subjects

Kaspar Rufibach

Publications and source records attributed to Kaspar Rufibach.

At least 19 recordsLinked to original sources

Communicating results in trials with multiple hypotheses or adaptive design features

Over time, clinical trials have increasingly incorporated complex design and analysis elements such as interim analyses, adaptations, multiple endpoints, and sophisticated multiplicity schemes for multiple endpoints and/or treatment arms following the paradigm of frequentist inference. In frequentist clinical trials multiplicity can come from (at least) four sources: multiple looks at the data, multiple endpoints, multiple populations, or multiple treatment comparisons. Normally, Type 1 error control across the multiple hypotheses is implemented to control chance of false positive decisions. To achieve this advanced techniques such as adaptive designs or graphical multiple testing procedures have been developed and are used in the design of clinical trials. However, these methods focus on hypothesis testing while subsequent estimation remains crucial to allow for a benefit-risk assessment and further use of the results by various stakeholders. Through examples, we illustrate challenges in estimation and transparent communication. In general, there are no simple solutions to this conceptual and communicational challenge. The purpose of this paper is to generate awareness of these issues and initiate a discussion about how to address them moving forward.

stat.ME

Statistical Methodology Groups in the Pharmaceutical Industry

Research and Development is the largest budget position in the pharmaceutical industry, with clinical trials being a critical, yet costly and time-consuming component to inform decisions. Beyond drug efficacy, the probability of success and efficiency of research and development are highly dependent on the approaches used for designing, analyzing, and interpreting clinical trials. Deep understanding of statistical methodology and quantitative approaches is therefore essential. Consequently, dedicated methodology groups have emerged in mid-size and large pharmaceutical companies and CROs. Their remit is to lead the conception and implementation of innovative quantitative methodologies in order to improve drug development, often by addressing complexities or offering more efficient designs. To achieve this, they collaborate internally and externally (e.g., with academics, regulators) to identify common challenges and tear down silos in order to invest in methods with the highest impact on efficiency and value to the portfolio. Given the immense financial stakes of drug development -- where delays carry massive implications -- these groups represent a critical strategic investment. However, to realize this business impact, statistical innovations must be rigorously validated and seamlessly integrated. This manuscript explores the setup, remit, and value of dedicated methodology groups, alongside the critical organizational considerations and success factors required to maximize their impact on the speed, efficiency, and probability of success.

stat.OT

Exhausting the type I error level in event-driven group-sequential designs with a closed testing procedure for progression-free and overall survival

In oncological clinical trials, overall survival (OS) is the gold-standard endpoint, but long follow-up and treatment switching can delay or dilute detectable effects. Progression-free survival (PFS) often provides earlier evidence and is therefore frequently used together with OS as multiple primary endpoints. Since in certain scenarios trial success may be defined if one of the two hypotheses involved can be rejected, a correction for multiple testing may be deemed necessary. Because PFS and OS are generally highly dependent, their test statistics are typically correlated. Ignoring this dependency (e.g. via a simple Bonferroni correction) is not power optimal. We develop a group-sequential testing procedure for the multiple primary endpoints PFS and OS that fully exhausts the family-wise error rate (FWER) by exploiting their dependence. Specifically, we characterize the joint asymptotic distribution of log-rank statistics across endpoints and multiple event-driven analysis cutoffs. Furthermore, we show that we can consistently estimate the covariance structure. Embedding these results in a closed testing procedure, we can recalculate critical values of the test statistics in order to spend the available type I error optimally. An important extension to the current literature is that we allow for both interim and final analysis to be event-driven. Simulations based on illness-death multi-state models empirically confirm FWER control for moderate to large sample sizes. Compared with a simple Bonferroni correction, the proposed methods recover roughly two thirds of the power loss for OS, increase disjunctive and conjunctive power, and enable meaningful early stopping. In planning, these gains translate into about 5% fewer OS events required to reach the targeted power. We also discuss practical issues in the implementation of such designs and possible extensions of the introduced method.

stat.ME

"6 choose 4": A framework to understand and facilitate discussion of strategies for overall survival safety monitoring

Advances in anticancer therapies have significantly contributed to declining death rates in certain disease and clinical settings. However, they have also made it difficult to power a clinical trial in these settings with overall survival (OS) as the primary efficacy endpoint. Therefore, two approaches have been recently proposed for the pre-specified analysis of OS as a safety endpoint (Fleming et al., 2024; Rodriguez et al., 2024). In this paper, we provide a simple, unifying framework that includes the aforementioned approaches (and a couple others) as special cases. By highlighting each approach's focus, priority, tolerance for risk, and strengths or challenges for practical implementation, this framework can help to facilitate discussions between stakeholders on "fit-for-purpose OS data collection and assessment of harm" (American Association for Cancer Research, 2024). We apply this framework to a real clinical trial in large B-cell lymphoma to illustrate its application and value. Several recommendations and open questions are also raised.

stat.ME

Clinical trials with interim analyses: Standardizing Terminology to increase clarity

Interim analyses for group-sequential decision making are prevalent in clinical trials. Methodology is well established and has been routinely implemented over the last decades. Still, confusions and uncertainties on aspects of how to operationalize and interpret interim analyses exist for many stakeholders. In this paper, a team of statisticians from the pharmaceutical industry, academia, and regulatory agencies provide a multi-stakeholder perspective on the key concepts behind interim analyses, with the aim to introduce standard terminology to mitigate misunderstandings and facilitate clearer discussions.

stat.AP

Balancing events, not patients, maximizes power of the logrank test: and other insights on unequal randomization in survival trials

We revisit the question of what randomization ratio (RR) maximizes power of the logrank test in event-driven survival trials under proportional hazards (PH). By comparing three approximations of the logrank test (Schoenfeld, Freedman, Rubinstein) to empirical simulations, we find that the RR that maximizes power is the RR that balances number of events across treatment arms at the end of the trial. This contradicts the common misconception implied by Schoenfeld's approximation that 1:1 randomization maximizes power. Besides power, we consider other factors that might influence the choice of RR (accrual, trial duration, sample size, etc.). We perform simulations to better understand how unequal randomization might impact these factors in practice. Altogether, we derive 6 insights to guide statisticians in the design of survival trials considering unequal randomization.

stat.ME

Using shrinkage methods to estimate treatment effects in overlapping subgroups in randomized clinical trials with a time-to-event endpoint

In randomized controlled trials, forest plots are frequently used to investigate the homogeneity of treatment effect estimates in subgroups. However, the interpretation of subgroup-specific treatment effect estimates requires great care due to the smaller sample size of subgroups and the large number of investigated subgroups. Bayesian shrinkage methods have been proposed to address these issues, but they often focus on disjoint subgroups while subgroups displayed in forest plots are overlapping, i.e., each subject appears in multiple subgroups. In our approach, we first build a flexible Cox model based on all available observations, including categorical covariates that identify the subgroups of interest and their interactions with the treatment group variable. We explore both penalized partial likelihood estimation with a lasso or ridge penalty for treatment-by-covariate interaction terms, and Bayesian estimation with a regularized horseshoe prior. One advantage of the Bayesian approach is the ability to derive credible intervals for shrunken subgroup-specific estimates. In a second step, the Cox model is marginalized to obtain treatment effect estimates for all subgroups. We illustrate these methods using data from a randomized clinical trial in follicular lymphoma and evaluate their properties in a simulation study. In all simulation scenarios, the overall mean-squared error is substantially smaller for penalized and shrinkage estimators compared to the standard subgroup-specific treatment effect estimator but leads to some bias for heterogeneous subgroups. We recommend that subgroup-specific estimators, which are typically displayed in forest plots, are more routinely complemented by treatment effect estimators based on shrinkage methods. The proposed methods are implemented in the R package bonsaiforest.

stat.ME

Survival analysis for AdVerse events with VarYing follow-up times (SAVVY): summary of findings and a roadmap for the future of safety analyses in clinical trials

The SAVVY project aims to improve the analyses of adverse events (AEs) in clinical trials through the use of survival techniques appropriately dealing with varying follow-up times and competing events (CEs). This paper summarizes key features and conclusions from the various SAVVY papers. Through theoretical investigations using simulations and in an empirical study including randomized clinical trials from several sponsor organisations, biases from ignoring varying follow-up times or CEs are investigated. The bias of commonly used estimators of the absolute and relative AE risk is quantified. Furthermore, we provide a cursory assessment of how pertinent guidelines for the analysis of safety data deal with the features of varying follow-up time and CEs. SAVVY finds that for both, avoiding bias and categorization of evidence with respect to treatment effect on AE risk into categories, the choice of the estimator is key and more important than features of the underlying data such as percentage of censoring, CEs, amount of follow-up, or value of the gold-standard. The choice of the estimator of the cumulative AE probability and the definition of CEs are crucial. SAVVY recommends using the Aalen-Johansen estimator (AJE) with an appropriate definition of CEs whenever the risk for AEs is to be quantified. There is an urgent need to improve the guidelines of reporting AEs so that incidence proportions or one minus Kaplan-Meier estimators are finally replaced by the AJE with appropriate definition of CEs.

stat.AP

Oncology clinical trial design planning based on a multistate model that jointly models progression-free and overall survival endpoints

When planning an oncology clinical trial, the usual approach is to assume proportional hazards and even an exponential distribution for time-to-event endpoints. Often, besides the gold-standard endpoint overall survival (OS), progression-free survival (PFS) is considered as a second confirmatory endpoint. We use a survival multistate model to jointly model these two endpoints and find that neither exponential distribution nor proportional hazards will typically hold for both endpoints simultaneously. The multistate model provides a stochastic process approach to model the dependency of such endpoints neither requiring latent failure times nor explicit dependency modelling such as copulae. We use the multistate model framework to simulate clinical trials with endpoints OS and PFS and show how design planning questions can be answered using this approach. In particular, non-proportional hazards for at least one of the endpoints are naturally modelled as well as their dependency to improve planning. We consider an oncology trial on non-small-cell lung cancer as a motivating example from which we derive relevant trial design questions. We then illustrate how clinical trial design can be based on simulations from a multistate model. Key applications are co-primary endpoints and group-sequential designs. Simulations for these applications show that the standard simplifying approach may very well lead to underpowered or overpowered clinical trials. Our approach is quite general and can be extended to more complex trial designs, further endpoints, and other therapeutic areas. An R package is available on CRAN.

stat.AP

Quantification of follow-up time in oncology clinical trials with a time-to-event endpoint: Asking the right questions

For the analysis of a time-to-event endpoint in a single-arm or randomized clinical trial it is generally perceived that interpretation of a given estimate of the survival function, or the comparison between two groups, hinges on some quantification of the amount of follow-up. Typically, a median of some loosely defined quantity is reported. However, whatever median is reported, is typically not answering the question(s) trialists actually have in terms of follow-up quantification. In this paper, inspired by the estimand framework, we formulate a comprehensive list of relevant scientific questions that trialists have when reporting time-to-event data. We illustrate how these questions should be answered, and that reference to an unclearly defined follow-up quantity is not needed at all. In drug development, key decisions are made based on randomized controlled trials, and we therefore also discuss relevant scientific questions not only when looking at a time-to-event endpoint in one group, but also for comparisons. We find that different thinking about some of the relevant scientific questions around follow-up is required depending on whether a proportional hazards assumption can be made or other patterns of survival functions are anticipated, e.g. delayed separation, crossing survival functions, or the potential for cure. We conclude the paper with practical recommendations.

stat.ME

Survival analysis for AdVerse events with VarYing follow-up times (SAVVY) -- comparison of adverse event risks in randomized controlled trials

Analyses of adverse events (AEs) are an important aspect of the evaluation of experimental therapies. The SAVVY (Survival analysis for AdVerse events with Varying follow-up times) project aims to improve the analyses of AE data in clinical trials through the use of survival techniques appropriately dealing with varying follow-up times, censoring, and competing events (CE). In an empirical study including seventeen randomized clinical trials the effect of varying follow-up times, censoring, and competing events on comparisons of two treatment arms with respect to AE risks is investigated. The comparisons of relative risks (RR) of standard probability-based estimators to the gold-standard Aalen-Johansen estimator or hazard-based estimators to an estimated hazard ratio (HR) from Cox regression are done descriptively, with graphical displays, and using a random effects meta-analysis on AE level. The influence of different factors on the size of the bias is investigated in a meta-regression. We find that for both, avoiding bias and categorization of evidence with respect to treatment effect on AE risk into categories, the choice of the estimator is key and more important than features of the underlying data such as percentage of censoring, CEs, amount of follow-up, or value of the gold-standard RR. There is an urgent need to improve the guidelines of reporting AEs so that incidence proportions are finally replaced by the Aalen-Johansen estimator - rather than by Kaplan-Meier - with appropriate definition of CEs. For RRs based on hazards, the HR based on Cox regression has better properties than the ratio of incidence densities.

stat.AP

Applying the Estimand and Target Trial frameworks to external control analyses using observational data: a case study in the solid tumor setting

In causal inference, the correct formulation of the scientific question of interest is a crucial step. Here we apply the estimand framework to a comparison of the outcomes of patient-level clinical trials and observational data to help structure the clinical question. In addition, we complement the estimand framework with the target trial framework to address specific issues in defining the estimand attributes using observational data and discuss synergies and differences of the two frameworks. Whereas the estimand framework proves useful to address the challenge that in clinical trials and routine clinical practice patients may switch to subsequent systemic therapies after the initially assigned systematic treatment, the target trial framework supports addressing challenges around baseline confounding and the index date. We apply the combined framework to compare long-term outcomes of a pooled set of three previously reported randomized phase 3 trials studying patients with metastatic non-small cell lung cancer receiving front-line chemotherapy (randomized clinical trial cohort) and similar patients treated with front-line chemotherapy as part of routine clinical care (observational comparative cohort). We illustrate the process to define the estimand attributes and select the estimator to estimate the estimand of interest while accounting for key baseline confounders, index date, and receipt of subsequent therapies. The proposed combined framework provides more clarity on the causal contrast of interest and the estimator to adopt and thus facilitates design and interpretation of the analyses.

stat.ME

Principal Stratum Strategy: Potential Role in Drug Development

A randomized trial allows estimation of the causal effect of an intervention compared to a control in the overall population and in subpopulations defined by baseline characteristics. Often, however, clinical questions also arise regarding the treatment effect in subpopulations of patients, which would experience clinical or disease related events post-randomization. Events that occur after treatment initiation and potentially affect the interpretation or the existence of the measurements are called {\it intercurrent events} in the ICH E9(R1) guideline. If the intercurrent event is a consequence of treatment, randomization alone is no longer sufficient to meaningfully estimate the treatment effect. Analyses comparing the subgroups of patients without the intercurrent events for intervention and control will not estimate a causal effect. This is well known, but post-hoc analyses of this kind are commonly performed in drug development. An alternative approach is the principal stratum strategy, which classifies subjects according to their potential occurrence of an intercurrent event on both study arms. We illustrate with examples that questions formulated through principal strata occur naturally in drug development and argue that approaching these questions with the ICH E9(R1) estimand framework has the potential to lead to more transparent assumptions as well as more adequate analyses and conclusions. In addition, we provide an overview of assumptions required for estimation of effects in principal strata. Most of these assumptions are unverifiable and should hence be based on solid scientific understanding. Sensitivity analyses are needed to assess robustness of conclusions.

stat.AP

Conditional Power and Friends: The Why and How of (Un)planned, Unblinded Sample Size Recalculations in Confirmatory Trials

Adapting the final sample size of a trial to the evidence accruing during the trial is a natural way to address planning uncertainty. Designs with adaptive sample size need to account for their optional stopping to guarantee strict type-I error-rate control. A variety of different methods to maintain type-I error-rate control after unplanned changes of the initial sample size have been proposed in the literature. This makes interim analyses for the purpose of sample size recalculation feasible in a regulatory context. Since the sample size is usually determined via an argument based on the power of the trial, an interim analysis raises the question of how the final sample size should be determined conditional on the accrued information. Conditional power is a concept often put forward in this context. Since it depends on the unknown effect size, we take a strict estimation perspective and compare assumed conditional power, observed conditional power, and predictive power with respect to their properties as estimators of the unknown conditional power. We then demonstrate that pre-planning an interim analysis using methodology for unplanned interim analyses is ineffective and naturally leads to the concept of optimal two-stage designs. We conclude that unplanned design adaptations should only be conducted as reaction to trial-external new evidence, operational needs to violate the originally chosen design, or post hoc changes in the objective criterion. Finally, we show that commonly discussed sample size recalculation rules can lead to paradoxical outcomes and propose two alternative ways of reacting to newly emerging trial-external evidence.

stat.AP

Estimands in Hematologic Oncology Trials

The estimand framework included in the addendum to the ICH E9 guideline facilitates discussions to ensure alignment between the key question of interest, the analysis, and interpretation. Therapeutic knowledge and drug mechanism play a crucial role in determining the strategy and defining the estimand for clinical trial designs. Clinical trials in patients with hematological malignancies often present unique challenges for trial design due to complexity of treatment options and existence of potential curative but highly risky procedures, e.g. stem cell transplant or treatment sequence across different phases (induction, consolidation, maintenance). Here, we illustrate how to apply the estimand framework in hematological clinical trials and how the estimand framework can address potential difficulties in trial result interpretation. This paper is a result of a cross-industry collaboration to connect the International Conference on Harmonisation (ICH) E9 addendum concepts to applications. Three randomized phase 3 trials will be used to consider common challenges including intercurrent events in hematologic oncology trials to illustrate different scientific questions and the consequences of the estimand choice for trial design, data collection, analysis, and interpretation. Template language for describing estimand in both study protocols and statistical analysis plans is suggested for statisticians' reference.

q-bio.OT

Survival analysis for AdVerse events with VarYing follow-up times (SAVVY) -- estimation of adverse event risks

The SAVVY project aims to improve the analyses of adverse event (AE) data in clinical trials through the use of survival techniques appropriately dealing with varying follow-up times and competing events (CEs). Although statistical methodologies have advanced, in AE analyses often the incidence proportion, the incidence density, or a non-parametric Kaplan-Meier estimator (KME) are used, which either ignore censoring or CEs. In an empirical study including randomized clinical trials from several sponsor organisations, these potential sources of bias are investigated. The main aim is to compare the estimators that are typically used in AE analysis to the Aalen-Johansen estimator (AJE) as the gold-standard. Here, one-sample findings are reported, while a companion paper considers consequences when comparing treatment groups. Estimators are compared with descriptive statistics, graphical displays and with a random effects meta-analysis. The influence of different factors on the size of the bias is investigated in a meta-regression. Comparisons are conducted at the maximum follow-up time and at earlier evaluation time points. CEs definition does not only include death before AE but also end of follow-up for AEs due to events possibly related to the disease course or the treatment. Ten sponsor organisations provided 17 trials including 186 types of AEs. The one minus KME was on average about 1.2-fold larger than the AJE. Leading forces influencing bias were the amount of censoring and of CEs. As a consequence, the average bias using the incidence proportion was less than 5%. Assuming constant hazards using incidence densities was hardly an issue provided that CEs were accounted for. There is a need to improve the guidelines of reporting risks of AEs so that the KME and the incidence proportion are replaced by the AJE with an appropriate definition of CEs.

stat.AP

A review of Bayesian perspectives on sample size derivation for confirmatory trials

Sample size derivation is a crucial element of the planning phase of any confirmatory trial. A sample size is typically derived based on constraints on the maximal acceptable type I error rate and a minimal desired power. Here, power depends on the unknown true effect size. In practice, power is typically calculated either for the smallest relevant effect size or a likely point alternative. The former might be problematic if the minimal relevant effect is close to the null, thus requiring an excessively large sample size. The latter is dubious since it does not account for the a priori uncertainty about the likely alternative effect size. A Bayesian perspective on the sample size derivation for a frequentist trial naturally emerges as a way of reconciling arguments about the relative a priori plausibility of alternative effect sizes with ideas based on the relevance of effect sizes. Many suggestions as to how such `hybrid' approaches could be implemented in practice have been put forward in the literature. However, key quantities such as assurance, probability of success, or expected power are often defined in subtly different ways in the literature. Starting from the traditional and entirely frequentist approach to sample size derivation, we derive consistent definitions for the most commonly used `hybrid' quantities and highlight connections, before discussing and demonstrating their use in the context of sample size derivation for clinical trials.

stat.AP

Assessing the Impact of COVID-19 on the Objective and Analysis of Oncology Clinical Trials -- Application of the Estimand Framework

COVID-19 outbreak has rapidly evolved into a global pandemic. The impact of COVID-19 on patient journeys in oncology represents a new risk to interpretation of trial results and its broad applicability for future clinical practice. We identify key intercurrent events that may occur due to COVID-19 in oncology clinical trials with a focus on time-to-event endpoints and discuss considerations pertaining to the other estimand attributes introduced in the ICH E9 addendum. We propose strategies to handle COVID-19 related intercurrent events, depending on their relationship with malignancy and treatment and the interpretability of data after them. We argue that the clinical trial objective from a world without COVID-19 pandemic remains valid. The estimand framework provides a common language to discuss the impact of COVID-19 in a structured and transparent manner. This demonstrates that the applicability of the framework may even go beyond what it was initially intended for.

q-bio.OT