SearcharxivSearch

arXiv subjects

Ekkehard Glimm

Publications and source records attributed to Ekkehard Glimm.

At least 19 recordsLinked to original sources

Unbiased estimation in two-stage adaptive enrichment designs

Recent advances in biomedical research have identified an increasing number of biomarkers associated with heterogeneity in patient responses to medical treatments. When a treatment is suspected to benefit certain patient subpopulations, adaptive enrichment designs may be more efficient and ethical. In such designs, an interim analysis is incorporated during the trial to select patient subpopulations for which the experimental treatment appears promising, according to predefined subpopulation selection rules. However, data-dependent selection can induce selection bias, causing conventional maximum likelihood estimators (MLEs) to overestimate the treatment effect in the selected patient subgroup. Existing inference methods for addressing this bias are typically rule-specific, highlighting the need for an estimation framework that accommodate a broader class of subpopulation selection rules. In this work, we define a general class of subpopulation selection rules based on the sample space partition condition and provide a systematic derivation that yields a unified formula for the Uniformly Minimum Variance Conditional Unbiased Estimator (UMVCUE). This generality allows our formulation to encompass a wide spectrum of adaptive enrichment designs, eliminating the necessity for case-specific derivations for each new design. Extensive simulations confirm the unbiasedness of the proposed UMVCUE, ensuring that therapeutic benefits are not overestimated. By bridging the gap between flexible interim subpopulation selection and rigorous statistical inference, our framework has the potential to facilitate the implementation of diverse subpopulation selection rules with greater ease in real-world trials and promote more efficient and ethical drug development.

stat.ME

Confidence intervals for two-stage adaptive designs with subpopulation selection

We consider clinical trials in which an experimental treatment is compared with a control in pre-specified patient subpopulations. In such settings, adaptive enrichment designs allow the enrolled population to be modified at an interim analysis, with subpopulations selected according to preplanned rules. Since these interim decisions are data-dependent, valid statistical inference must account for them. We focus on constructing confidence intervals for the treatment effect in the selected population. Confidence interval methods that ignore the possibility of population modification may fail to achieve the desired coverage probability. We propose a new approach that constructs confidence intervals with exact nominal coverage conditional on the interim decision. Importantly, our method applies to a broad class of adaptive enrichment designs, rather than a single specific design. Our method involves deriving the distribution of the naive estimator of the treatment effect in the selected population conditional on the interim decision and inverting uniformly most accurate unbiased tests to obtain the confidence interval. We provide an efficient computational procedure and show through extensive simulations that the resulting confidence intervals satisfy the theoretical coverage guarantees.

stat.ME

Exact matching as an alternative to propensity score matching

The comparison of different medical treatments from observational studies or across different clinical studies is often biased by confounding factors such as systematic differences in patient demographics or in the inclusion criteria for the trials. Propensity score matching is a popular method to adjust for such confounding. It compares weighted averages of patient responses. The weights are calculated from logistic regression models with the intention to reduce differences between the confounders in the treatment groups. However, the groups are only "roughly matched" with no generally accepted principle to determine when a match is "good enough". In this manuscript, we propose an alternative approach to the matching problem by considering it as a constrained optimization problem. We investigate the conditions for exact matching in the sense that the average values of confounders are identical in the treatment groups after matching. Our approach is similar to the matching-adjusted indirect comparison approach by Signorovitch et al. (2010) but with two major differences: First, we do not impose any specific functional form on the matching weights; second, the proposed approach can be applied to individual patient data from several treatment groups as well as to a mix of individual patient and aggregated data.

stat.ME

Optimal allocation strategies in platform trials

Platform trials are randomized clinical trials that allow simultaneous comparison of multiple interventions, usually against a common control. Arms to test experimental interventions may enter and leave the platform over time. This implies that the number of experimental intervention arms in the trial may change over time. Determining optimal allocation rates to allocate patients to the treatment and control arms in platform trials is challenging because the change in treatment arms implies that also the optimal allocation rates will change when treatments enter or leave the platform. In addition, the optimal allocation depends on the analysis strategy used. In this paper, we derive optimal treatment allocation rates for platform trials with shared controls, assuming that a stratified estimation and testing procedure based on a regression model, is used to adjust for time trends. We consider both, analysis using concurrent controls only as well as analysis methods based on also non-concurrent controls and assume that the total sample size is fixed. The objective function to be minimized is the maximum of the variances of the effect estimators. We show that the optimal solution depends on the entry time of the arms in the trial and, in general, does not correspond to the square root of $k$ allocation rule used in the classical multi-arm trials. We illustrate the optimal allocation and evaluate the power and type 1 error rate compared to trials using one-to-one and square root of $k$ allocations by means of a case study.

stat.ME

Designing an exploratory phase 2b platform trial in NASH with correlated, co-primary binary endpoints

Non-alcoholic steatohepatitis (NASH) is the progressive form of nonalcoholic fatty liver disease (NAFLD) and a disease with high unmet medical need. Platform trials provide great benefits for sponsors and trial participants in terms of accelerating drug development programs. In this article, we describe some of the activities of the EU-PEARL consortium (EU Patient-cEntric clinicAl tRial pLatforms) regarding the use of platform trials in NASH, in particular the proposed trial design, decision rules and simulation results. For a set of assumptions, we present the results of a simulation study recently discussed with two health authorities and the learnings from these meetings from a trial design perspective. Since the proposed design uses co-primary binary endpoints, we furthermore discuss the different options and practical considerations for simulating correlated binary endpoints.

stat.AP

On model-based time trend adjustments in platform trials with non-concurrent controls

Platform trials can evaluate the efficacy of several treatments compared to a control. The number of treatments is not fixed, as arms may be added or removed as the trial progresses. Platform trials are more efficient than independent parallel-group trials because of using shared control groups. For arms entering the trial later, not all patients in the control group are randomised concurrently. The control group is then divided into concurrent and non-concurrent controls. Using non-concurrent controls (NCC) can improve the trial's efficiency, but can introduce bias due to time trends. We focus on a platform trial with two treatment arms and a common control arm. Assuming that the second treatment arm is added later, we assess the robustness of model-based approaches to adjust for time trends when using NCC. We consider approaches where time trends are modeled as linear or as a step function, with steps at times where arms enter or leave the trial. For trials with continuous or binary outcomes, we investigate the type 1 error (t1e) rate and power of testing the efficacy of the newly added arm under a range of scenarios. In addition to scenarios where time trends are equal across arms, we investigate settings with trends that are different or not additive in the model scale. A step function model fitted on data from all arms gives increased power while controlling the t1e, as long as the time trends are equal for the different arms and additive on the model scale. This holds even if the trend's shape deviates from a step function if block randomisation is used. But if trends differ between arms or are not additive on the model scale, t1e control may be lost. The efficiency gained by using step function models to incorporate NCC can outweigh potential biases. However, the specifics of the trial, plausibility of different time trends, and robustness of results should be considered

stat.ME

Familywise error rate control for block response-adaptive randomization

Response-adaptive randomization allows the probabilities of allocating patients to treatments in a clinical trial to change based on the previously observed response data, in order to achieve different experimental goals. One concern over the use of such designs in practice, particularly from a regulatory viewpoint, is controlling the type I error rate. To address this, Robertson and Wason (Biometrics, 2019) proposed methodology that guarantees familywise error rate control for a large class of response-adaptive designs. In this paper, we propose an improvement of their proposal that is conceptually simpler, in the specific context of block-randomised trials with a fixed allocation to the control arm. We show the modified method guarantees that there will never be negative weights for blocks of data, and can also provide a substantial power advantage in practice.

stat.ME

CohortPlat: Simulation of cohort platform trials investigating combination therapies

Platform trials have gained a lot of attention recently as a possible remedy for time-consuming classical two-arm randomized controlled trials, especially in early phase drug development. This short article illustrates how to use the CohortPlat R package to simulate a cohort platform trial, where each cohort consists of a combination treatment and the respective monotherapies and standard-of-care. The endpoint is always assumed to be binary. The package offers extensive flexibility with respect to both platform trial trajectories, as well as treatment effect scenarios and decision rules. As a special feature, the package provides a designated function for running multiple such simulations efficiently in parallel and saving the results in a concise manner. Many illustrations of code usage are provided.

stat.AP

Decision rules for identifying combination therapies in open-entry, randomized controlled platform trials

Platform trials have become increasingly popular for drug development programs, attracting interest from statisticians, clinicians and regulatory agencies. Many statistical questions related to designing platform trials - such as the impact of decision rules, sharing of information across cohorts, and allocation ratios on operating characteristics and error rates - remain unanswered. In many platform trials, the definition of error rates is not straightforward as classical error rate concepts are not applicable. For an open-entry, exploratory platform trial design comparing combination therapies to the respective monotherapies and standard-of-care, we define a set of error rates and operating characteristics and then use these to compare a set of design parameters under a range of simulation assumptions. When setting up the simulations, we aimed for realistic trial trajectories, such that e.g. a priori we do not know the exact number of treatments that will be included over time in a specific simulation run as this follows a stochastic mechanism. Our results indicate that the method of data sharing, exact specification of decision rules and a priori assumptions regarding the treatment efficacy all strongly contribute to the operating characteristics of the platform trial. Furthermore, different operating characteristics might be of importance to different stakeholders. Together with the potential flexibility and complexity of a platform trial, which also impact the achieved operating characteristics via e.g. the degree of efficiency of data sharing, this implies that utmost care needs to be given to evaluation of different assumptions and design parameters at the design stage.

stat.AP

Geometric approaches to assessing the numerical feasibility for conducting matching-adjusted indirect comparisons

We discuss how to handle matching-adjusted indirect comparison (MAIC) from a data analyst's perspective. We introduce several multivariate data analysis methods to assess the appropriateness of MAIC for a given data set. These methods focus on comparing the baseline variables used in the matching from a study that provides the summary statistics, or aggregated data (AD) and a study that provides individual patient level data (IPD). The methods identify situations when no numerical solutions are possible with the MAIC method. This helps to avoid misleading results being produced. Moreover, it has been observed that sometimes contradicting results are reported by two sets of MAIC analyses produced by two teams, each having their own IPD and applying MAIC using the AD published by the other team. We show that an intrinsic property of the MAIC estimated weights can be a contributing factor for this phenomenon.

stat.AP

Connecting Instrumental Variable methods for causal inference to the Estimand Framework

Causal inference methods are gaining increasing prominence in pharmaceutical drug development in light of the recently published addendum on estimands and sensitivity analysis in clinical trials to the E9 guideline of the International Council for Harmonisation. The E9 addendum emphasises the need to account for post-randomization or `intercurrent' events that can potentially influence the interpretation of a treatment effect estimate at a trial's conclusion. Instrumental Variables (IV) methods have been used extensively in economics, epidemiology and academic clinical studies for `causal inference', but less so in the pharmaceutical industry setting until now. In this tutorial paper we review the basic tools for causal inference, including graphical diagrams and potential outcomes, as well as several conceptual frameworks that an IV analysis can sit within. We discuss in detail how to map these approaches to the Treatment Policy, Principal Stratum and Hypothetical `estimand strategies' introduced in the E9 addendum, and provide details of their implementation using standard regression models. Specific attention is given to discussing the assumptions each estimation strategy relies on in order to be consistent, the extent to which they can be empirically tested and sensitivity analyses in which specific assumptions can be relaxed. We finish by applying the methods described to simulated data closely matching two recent pharmaceutical trials to further motivate and clarify the ideas

stat.ME

Confidence intervals with maximal average power

We propose a frequentist testing procedure that maintains a defined coverage and is optimal in the sense that it gives maximal power to detect deviations from a null hypothesis when the alternative to the null hypothesis is sampled from a pre-specified distribution (the prior distribution). Selecting a prior distribution allows to tune the decision rule. This leads to an increased power, if the true data generating distribution happens to be compatible with the prior. It comes at the cost of losing power, if the data generating distribution or the observed data are incompatible with the prior. We illustrate the proposed approach for a binomial experiment, which is sufficiently simple such that the decision sets can be illustrated in figures, which should facilitate an intuitive understanding. The potential beyond the simple example will be discussed: the approach is generic in that the test is defined based on the likelihood function and the prior only. It is comparatively simple to implement and efficient to execute, since it does not rely on Minimax optimization. Conceptually it is interesting to note that for constructing the testing procedure the Bayesian posterior probability distribution is used.

stat.AP

Blinded sample size re-estimation in equivalence testing

This paper investigates type I error violations that occur when blinded sample size reviews are applied in equivalence testing. We give a derivation which explains why such violations are more pronounced in equivalence testing than in the case of superiority testing. In addition, the amount of type I error inflation is quantified by simulation as well as by some theoretical considerations. Non-negligible type I error violations arise when blinded interim re-assessments of sample sizes are performed particularly if sample sizes are small, but within the range of what is practically relevant.

stat.AP

Group sequential designs for negative binomial outcomes

Count data and recurrent events in clinical trials, such as the number of lesions in magnetic resonance imaging in multiple sclerosis, the number of relapses in multiple sclerosis, the number of hospitalizations in heart failure, and the number of exacerbations in asthma or in chronic obstructive pulmonary disease (COPD) are often modeled by negative binomial distributions. In this manuscript we study planning and analyzing clinical trials with group sequential designs for negative binomial outcomes. We propose a group sequential testing procedure for negative binomial outcomes based on Wald statistics using maximum likelihood estimators. The asymptotic distribution of the proposed group sequential tests statistics are derived. The finite sample size properties of the proposed group sequential test for negative binomial outcomes and the methods for planning the respective clinical trials are assessed in a simulation study. The simulation scenarios are motivated by clinical trials in chronic heart failure and relapsing multiple sclerosis, which cover a wide range of practically relevant settings. Our research assures that the asymptotic normal theory of group sequential designs can be applied to negative binomial outcomes when the hypotheses are tested using Wald statistics and maximum likelihood estimators. We also propose two methods, one based on Student's t-distribution and one based on resampling, to improve type I error rate control in small samples. The statistical methods studied in this manuscript are implemented in the R package \textit{gscounts}, which is available for download on the Comprehensive R Archive Network (CRAN).

stat.AP

Optimal exact tests for multiple binary endpoints

In confirmatory clinical trials with small sample sizes, hypothesis tests based on asymptotic distributions are often not valid and exact non-parametric procedures are applied instead. However, the latter are based on discrete test statistics and can become very conservative, even more so, if adjustments for multiple testing as the Bonferroni correction are applied. We propose improved exact multiple testing procedures for the setting where two parallel groups are compared in multiple binary endpoints. Based on the joint conditional distribution of test statistics of Fisher's exact tests, optimal rejection regions for intersection hypotheses tests are constructed. To efficiently search the large space of possible rejection regions, we propose an optimization algorithm based on constrained optimization and integer linear programming. Depending on the optimization objective, the optimal test yields maximal power under a specific alternative, maximal exhaustion of the nominal type I error rate, or the largest possible rejection region controlling the type I error rate. Applying the closed testing principle, we construct optimized multiple testing procedures with strong familywise error rate control. Furthermore, we propose a greedy algorithm for nearly optimal tests, which is computationally more efficient. We numerically compare the unconditional power of the optimized procedure with alternative approaches and illustrate the optimal tests with a clinical trial example in a rare disease.

stat.ME

On weighted parametric tests

We describe a general framework for weighted parametric multiple test procedures based on the closure principle. We utilize general weighting strategies that can reflect complex study objectives and include many procedures in the literature as special cases. The proposed weighted parametric tests bridge the gap between rejection rules using either adjusted significance levels or adjusted $p$-values. This connection is possible by allowing intersection hypotheses to be tested at level smaller than $α$, which may be needed for certain study considerations. For such cases we introduce a subclass of exact $α$-level parametric tests which satisfy the consonance property. When only subsets of test statistics are correlated, a new procedure is proposed to fully utilize the parametric assumptions within each subset. We illustrate the proposed weighted parametric tests using a clinical trial example.

stat.ME

Model-based dose finding under model uncertainty using general parametric models

Statistical methodology for the design and analysis of clinical Phase II dose response studies, with related software implementation, are well developed for the case of a normally distributed, homoscedastic response considered for a single timepoint in parallel group study designs. In practice, however, binary, count, or time-to-event endpoints are often used, typically measured repeatedly over time and sometimes in more complex settings like crossover study designs. In this paper we develop an overarching methodology to perform efficient multiple comparisons and modeling for dose finding, under uncertainty about the dose-response shape, using general parametric models. The framework described here is quite general and covers dose finding using generalized non-linear models, linear and non-linear mixed effects models, Cox proportional hazards (PH) models, etc. In addition to the core framework, we also develop a general purpose methodology to fit dose response data in a computationally and statistically efficient way. Several examples, using a variety of different statistical models, illustrate the breadth of applicability of the results. For the analyses we developed the R add-on package DoseFinding, which provides a convenient interface to the general approach adopted here.

stat.ME

Some Notes on Blinded Sample Size Re-Estimation

This note investigates a number of scenarios in which unadjusted testing following a blinded sample size re-estimation leads to type I error violations. For superiority testing, this occurs in certain small-sample borderline cases. We discuss a number of alternative approaches that keep the type I error rate. The paper also gives a reason why the type I error inflation in the superiority context might have been missed in previous publications and investigates why it is more marked in case of non-inferiority testing.

stat.ME