SearcharxivSearch

arXiv subjects

Werner Brannath

Publications and source records attributed to Werner Brannath.

At least 19 recordsLinked to original sources

Optimal monotone conditional error functions

This paper presents a general method that provides optimal monotone conditional error functions for confirmatory adaptive two-stage designs with conditional power based sample size recalculations. The presented method builds on a previously developed general theory for optimal adaptive two-stage designs where sample sizes are reassessed for a specific conditional power and the goal is to minimize the expected sample size. The previous theory can easily lead to a non-monotonous conditional error function, which is highly undesirable for logical reasons and, as we show, can harm type I error rate control for composite null hypotheses. We also show that type I error control is generally guaranteed with a conditional error function (CEF) that is non-increasing in the first stage p-value. We present a method that extends the existing theory by introducing an intermediate monotonising steps that can easily be implemented and provides a non-increasing conditional error function. We show mathematically that the monotonising step provides the optimal non-increasing conditional error function. We illustrate the method with several examples using optconerrf, an R package implemented for this paper.

stat.ME

Blinded sample size review for McNemar's test based on primary and surrogate endpoints

We develop blinded sample size re-estimation strategies for McNemar's test based on paired binary primary and secondary short-term surrogate endpoints. The development is motivated by a prospective randomized clinical trial on childhood glaucoma. A conditional power expression for McNemar's test given the primary endpoint at an interim analysis is derived and complemented by a sample size re-estimation rule. We show that this procedure preserves the type I error rate while allowing the second-stage sample size to be chosen to attain a prespecified target power. In the case where for some patients only a short-term surrogate endpoint is available at interim, we introduce a surrogate-based re-estimation approach that conditions on all possible numbers of primary-endpoint discordant pairs using transition rates from the surrogate to the primary outcome. We show how these transition rates can be estimated from data on a subsample for which both surrogate and primary endpoint are available. We derive the resulting surrogate endpoint-based conditional power and sample size rule and illustrate their use with the example of the motivating trial.

stat.ME

LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models

The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedented flexibility of NNs. In order to preserve interpretability, it is usually necessary to restrict the NN components to prevent them from dominating the model. However, existing methods that enforce structural constraints on their NN components severely limit their models' flexibility; in contrast, methods that only enforce weak, indirect constraints lose meaningful interpretability. The method we propose therefore leverages invertible residual neural networks (i-ResNets) to equip generalized linear models with both nonlinear parameter estimation and a flexible correction of their distributional assumptions while always retaining stochastic monotonicity of the modeled distribution in the (formerly linear) predictor. The i-ResNets correspond to a controlled deviation from identity and by constraining their Lipschitz constant one can rigorously limit and quantify how far the hybrid model deviates from its traditional counterpart. This enables a user-specifiable compromise between flexibility and interpretability without limiting the structure of nonlinear and interaction effects that can be learned. Furthermore, we develop specific inherent interpretation techniques for our model and enforce model identifiability through an adapted post-hoc orthogonalization.

stat.ML

Multiple type I error concepts for clinical trials with overlapping populations

The population-wise error rate (PWER) was introduced as a more liberal alternative to the family-wise error rate (FWER) for clinical trials with multiple, overlapping patient populations. These trials are particularly relevant in personalized medicine, which aims to find therapies tailored to specific patient subgroups. By controlling an average multiple type I error probability over all population strata, the PWER can substantially improve statistical power. However, one disadvantage of this concept is that the error probability for a given population can strongly depend on the presence or absence of other populations included in the analysis. To address this issue, we propose two modifications of the PWER that enforce individual error control either for all target populations, or for all possible unions of target populations. We call these approaches the PWER over the populations (PWER-P) and the PWER over population unions (PWER-U). We investigate the properties of these new error rates and compare them with the PWER and FWER in terms of type I error control and power.

stat.ME

Principles in harmony: Closed testing meets the partitioning principle for computational efficiency

We explore and utilize the algorithmic relationship between the closed testing principle for multiple tests with family-wise error rate (FWER) control and the partitioning principle for the construction of simultaneous confidence intervals. Starting with the simple observation that a multiple test with FWER control is formally equivalent to a one-sided simultaneous confidence interval for the vector of binary parameter indicating whether the null or alternative hypothesis is true, we show that the closed testing and partitioning principles follow the same computational approach. We will then utilise this relationship to extend concepts of consonance for closed tests to the partitioning principle, with the aim of deriving computationally feasible and efficient algorithms for the calculation of simultaneous confidence intervals. We will also utilize the relationship between closed testing and partitioning principle to extend common closed testing procedures to simultaneous confidence intervals, referencing the existing literature on informative simultaneous confidence intervals. The relationships and extensions will be illustrated by simple, instructive examples.

stat.ME

Anytime-valid testing with e-values and confirmatory adaptive designs

Confirmatory adaptive designs were introduced more than 30 years ago and enable for example sample size re-assessments and the selection of treatments, endpoints as well as subpopulations during the course of a clinical trial. Recently, sequential tests based on e-values for an anytime-valid inference have been developed, promising seemingly similar or even more flexibility and utility. In this note, we compare these two independently developed concepts, shedding light on their formal and methodological connections and differences. Specifically, we show that adaptive design tools like conditional error functions and combination tests are formally equivalent to e-value based, anytime-valid sequential tests. However, in spite of their common fundamental intention to bring flexibility into statistical inference, they have quite different emphases: While hypothesis testing with combination tests and conditional error function usually intent to exhaust type I error rates under the offered flexibility, e-value based testing aims on the additional flexibility with regard to optional continuation, the chosen level and, in recent extensions, in the loss functions to be controlled. We also indicate how recent e-value achievements could enrich clinical trial methodology and adaptive design methodology could inspire and improve e-value based testing.

stat.ME

Informative Simultaneous Confidence Intervals for Graphical Group Sequential Test Procedures

Test procedures for multiple hypotheses in a group sequential clinical trial that control the family-wise error rate are considered. Several graphical group sequential tests suggested in the literature, which are special cases of Bonferroni-closure tests, are discussed. The focus is on the question of whether to consider at the current stage only the evidence of the current repeated p-value or the evidence over all repeated p-values from the previous stages. A new test strategy controlling the family-wise error rate is introduced that consistently works across all hypotheses, with the evidence (i.e., repeated p-value) from the current stage. The strategy is more powerful than similar previously suggested test procedures. This is achieved by using the evidence from previous stages to increase the significance levels. For the test procedures, corresponding compatible simultaneous confidence intervals are presented, having the disadvantage of often not providing additional information on the treatment effects. For this reason, we extend previous work about informative simultaneous confidence intervals for one-stage graphical tests to graphical group sequential trials. Iterative algorithms are introduced that calculate these informative bounds that have a small power loss compared to the original graphical group sequential test. The boundaries can be calculated after each stage. In addition, previous work is extended by a criterion to estimate the accuracy of the numerically calculated boundaries. The suggested informative bounds can be used to provide median-conservative, i.e., reliable estimators, for estimating the treatment effects in a group sequential test with multiple hypotheses.

stat.ME

A prediction interval for the population-wise error rate

We construct an asymptotic prediction interval for the population-wise error rate (PWER), which is a multiple type I error criterion for clinical trials with overlapping patient populations. The PWER is the probability that a randomly selected patient will receive an ineffective treatment. It must usually be estimated due to unknown population strata sizes, such that only an estimate can be controlled at the given significance level. We apply the delta method to find a prediction interval for the resulting true PWER, we demonstrate by simulations that the interval has the required coverage probability, and illustrate the approach with real data examples.

stat.ME

Family-wise error rate control in clinical trials with overlapping populations

We consider clinical trials with multiple, overlapping patient populations, that test multiple treatment policies specifically tailored to these populations. Such designs may lead to multiplicity issues, as false statements will affect several populations. For type I error control, often the family-wise error rate (FWER) is controlled, which is the probability to reject at least one true null hypothesis. If the joint distribution of the test statistics is known, the FWER level can be exhausted by determining critical values or adjusted $α$-levels. The adjustment is typically done under the common ANOVA assumptions. However, the performed tests are then only valid under the rather strong assumption of homogeneous null effects, i.e., when the null hypothesis applies to all subpopulations and their intersections. We show that under cancelling null effects, when heterogeneous effects cancel out in some or all subpopulations, this procedure does not provide FWER control. We also suggest different alternatives and compare them in terms of FWER control and their power.

stat.ME

Adaptive Designs in Fast-Track Registration Processes for Digital Health Applications

Fast-track procedures play an important role in the context of conditional registration of medical devices, such as listing processes for digital health applications. They offer the potential for earlier patient access to innovative products and involve two registration steps. The applicants can apply first for conditional registration. A successful conditional registration provides a limited funding or approval period and time to prepare the application for permanent registration (the second registration step). For conditional registration, products have to fulfill only a part of the requirements necessary for permanent registration. There is interest in valid and efficient study designs for fast-track procedures. This will be addressed in this paper. A motivating example is the German fast-track registration process of digital health applications (DiGA) for reimbursement by statutory health insurances. The main focus of the paper is the systematic statistical investigation of the utility of adaptive designs in the context of fast-track registration processes like the DiGA fast-track. We demonstrate that, in most cases, such designs are much more efficient than the current standard of two separate studies. A careful statistical discussion of the registration requirements and their consequences is also included. The results are based on numerical calculations supported by mathematical arguments.

stat.ME

Asymptotic Online FWER Control for Dependent Test Statistics

In online multiple testing, an a priori unknown number of hypotheses are tested sequentially, i.e. at each time point a test decision for the current hypothesis has to be made using only the data available so far. Although many powerful test procedures have been developed for online error control in recent years, most of them are designed solely for independent or at most locally dependent test statistics. In this work, we provide a new framework for deriving online multiple test procedures which ensure asymptotical (with respect to the sample size) control of the familywise error rate (FWER), regardless of the dependence structure between test statistics. In this context, we give a few concrete examples of such test procedures and discuss their properties. Furthermore, we conduct a simulation study in which the type I error control of these test procedures is also confirmed for a finite sample size and a gain in power is indicated.

stat.ME

The effect of estimating prevalences on the population-wise error rate

The population-wise error rate (PWER) is a type I error rate for clinical trials with multiple target populations. In such trials, a treatment is tested for its efficacy in each population. The PWER is defined as the probability that a randomly selected, future patient will be exposed to an inefficient treatment based on the study results. It can be understood and computed as an average of strata-specific family wise error rates and involves the prevalences of these strata. A major issue of this concept is that the prevalences are usually unknown in practice, so that the PWER cannot be directly controlled. Instead, one could use an estimator based on the given sample, like their maximum-likelihood estimator under a multinomial distribution. In this article, we demonstrate through simulations that this does not substantially inflate the true PWER. We differentiate between the expected PWER, which is almost perfectly controlled, and study-specific values of the PWER which are conditioned on all subgroup sample sizes and vary within a narrow range. Thereby, we consider up to eight different overlapping populations and moderate to large sample sizes. In these settings, we also consider the maximum strata-wise family wise error rate, which is found to be, on average, at least bounded by twice the significance level used for PWER control.

stat.ME

ADDIS-Graphs for online error control with application to platform trials

In contemporary research, online error control is often required, where an error criterion, such as familywise error rate (FWER) or false discovery rate (FDR), shall remain under control while testing an a priori unbounded sequence of hypotheses. The existing online literature mainly considered large-scale designs and constructed blackbox-like algorithms for these. However, smaller studies, such as platform trials, require high flexibility and easy interpretability to take study objectives into account and facilitate the communication. Another challenge in platform trials is that due to the shared control arm some of the p-values are dependent and significance levels need to be prespecified before the decisions for all the past treatments are available. We propose ADDIS-Graphs with FWER control that due to their graphical structure perfectly adapt to such settings and provably uniformly improve the state-of-the-art method. We introduce several extensions of these ADDIS-Graphs, including the incorporation of information about the joint distribution of the p-values and a version for FDR control.

stat.ME

Informative Simultaneous Confidence Intervals for Graphical Test Procedures

Simultaneous confidence intervals (SCIs) that are compatible with a given closed test procedure are often non-informative. More precisely, for a one-sided null hypothesis, the bound of the SCI can stick to the border of the null hypothesis, irrespective of how far the point estimate deviates from the null hypothesis. This has been illustrated for the Bonferroni-Holm and fall-back procedures, for which alternative SCIs have been suggested, that are free of this deficiency. These informative SCIs are not fully compatible with the initial multiple test, but are close to it and hence provide similar power advantages. They provide a multiple hypothesis test with strong family-wise error rate control that can be used in replacement of the initial multiple test. The current paper extends previous work for informative SCIs to graphical test procedures. The information gained from the newly suggested SCIs is shown to be always increasing with increasing evidence against a null hypothesis. The new SCIs provide a compromise between information gain and the goal to reject as many hypotheses as possible. The SCIs are defined via a family of dual graphs and the projection method. A simple iterative algorithm for the computation of the intervals is provided. A simulation study illustrates the results for a complex graphical test procedure.

stat.ME

An exhaustive ADDIS principle for online FWER control

In this paper we consider online multiple testing with familywise error rate (FWER) control, where the probability of committing at least one type I error shall remain under control while testing a possibly infinite sequence of hypotheses over time. Currently, Adaptive-Discard (ADDIS) procedures seem to be the most promising online procedures with FWER control in terms of power. Now, our main contribution is a uniform improvement of the ADDIS principle and thus of all ADDIS procedures. This means, the methods we propose reject as least as much hypotheses as ADDIS procedures and in some cases even more, while maintaining FWER control. In addition, we show that there is no other FWER controlling procedure that enlarges the event of rejecting any hypothesis. Finally, we apply the new principle to derive uniform improvements of the ADDIS-Spending and ADDIS-Graph.

stat.ME

The Online Closure Principle

The closure principle is fundamental in multiple testing and has been used to derive many efficient procedures with familywise error rate control. However, it is often unsuitable for modern research, which involves flexible multiple testing settings where not all hypotheses are known at the beginning of the evaluation. In this paper, we focus on online multiple testing where a possibly infinite sequence of hypotheses is tested over time. At each step, it must be decided on the current hypothesis without having any information about the hypotheses that have not been tested yet. Our main contribution is a general and stringent mathematical definition of online multiple testing and a new online closure principle which ensures that the resulting closed procedure can be applied in the online setting. We prove that any familywise error rate controlling online procedure can be derived by this online closure principle and provide admissibility results. In addition, we demonstrate how short-cuts of these online closed procedures can be obtained under a suitable consonance property.

stat.ME

Statistical optimization of expensive multi-response black-box functions

Assume that a set of $P$ process parameters $p_i$, $i=1,\dots,P$, determines the outcome of a set of $D$ descriptor variables $d_j$, $j=1,\dots,D$, via an unknown functional relationship $ϕ: \mathbf{p} \mapsto \mathbf{d}, \, \mathbb{R}^{P} \to \mathbb{R}^{D}$, where $\mathbf{p}=(p_1,\dots,p_{P})$, $\mathbf{d}=(d_1,\dots,d_{D})$. It is desired to find appropriate values $\mathbf{\hat p} = ({\hat p}_1,\dots, {\hat p}_P)$ for the process parameters such that the corresponding values of the descriptor variables $ϕ(\mathbf {\hat p})$ are close to a given target $\mathbf d^*=(d^*_1,\dots,d^*_D)$, assuming that at least one exact solution exists. A sequential approach using dimension reduction techniques has been developed to achieve this. In a simulation study, results of the suggested approach and the algorithms NSGA-II, SMS-EMOA and MOEA/D are compared.

stat.CO

Post-Selection Confidence Bounds for Prediction Performance

In machine learning, the selection of a promising model from a potentially large number of competing models and the assessment of its generalization performance are critical tasks that need careful consideration. Typically, model selection and evaluation are strictly separated endeavors, splitting the sample at hand into a training, validation, and evaluation set, and only compute a single confidence interval for the prediction performance of the final selected model. We however propose an algorithm how to compute valid lower confidence bounds for multiple models that have been selected based on their prediction performances in the evaluation set by interpreting the selection problem as a simultaneous inference problem. We use bootstrap tilting and a maxT-type multiplicity correction. The approach is universally applicable for any combination of prediction models, any model selection strategy, and any prediction performance measure that accepts weights. We conducted various simulation experiments which show that our proposed approach yields lower confidence bounds that are at least comparably good as bounds from standard approaches, and that reliably reach the nominal coverage probability. In addition, especially when sample size is small, our proposed approach yields better performing prediction models than the default selection of only one model for evaluation does.

stat.ML