SearcharxivSearch

arXiv subjects

Frank Konietschke

Publications and source records attributed to Frank Konietschke.

18 recordsLinked to original sources

A New Approach to the Nonparametric Behrens-Fisher Problem with Compatible Confidence Intervals

We propose a new test to address the nonparametric Behrens-Fisher problem involving different distribution functions in the two samples. Our procedure tests the null hypothesis $\mathcal{H}_0: θ= \frac{1}{2}$, where $θ= P(X<Y) + \frac{1}{2}P(X=Y)$ denotes the Mann-Whitney effect. No restrictions on the underlying distributions of the data are imposed with the trivial exception of one-point distributions. The method is based on evaluating the ratio of the variance $σ_N^2$ of the Mann-Whitney effect estimator $\widehatθ$ to its theoretical maximum, as derived from the Birnbaum-Klose inequality. Through simulations, we demonstrate that the proposed test effectively controls the type-I error rate under various conditions, including small sample sizes, unbalanced designs, and different data-generating mechanisms. Notably, it provides better control of the type-1 error rate compared to the widely used Brunner-Munzel test, particularly at small significance levels such as $α\in \{0.01, 0.005\}$. Additionally, we derive range-preserving compatible confidence intervals, showing that they offer improved coverage over those compatible to the Brunner-Munzel test. Finally, we illustrate the application of our method in a clinical trial example.

stat.ME

Single CASANOVA? Not in multiple comparisons

When comparing multiple groups in clinical trials, we are not only interested in whether there is a difference between any groups but rather the location. Such research questions lead to testing multiple individual hypotheses. To control the familywise error rate (FWER), we must apply some corrections or introduce tests that control the FWER by design. In the case of time-to-event data, a Bonferroni-corrected log-rank test is commonly used. This approach has two significant drawbacks: (i) it loses power when the proportional hazards assumption is violated [1] and (ii) the correction generally leads to a lower power, especially when the test statistics are not independent [2]. We propose two new tests based on combined weighted log-rank tests. One as a simple multiple contrast test of weighted log-rank tests and one as an extension of the so-called CASANOVA test [3]. The latter was introduced for factorial designs. We propose a new multiple contrast test based on the CASANOVA approach. Our test promises to be more powerful under crossing hazards and eliminates the need for additional p-value correction. We assess the performance of our tests through extensive Monte Carlo simulation studies covering both proportional and non-proportional hazard scenarios. Finally, we apply the new and reference methods to a real-world data example. The new approaches control the FWER and show reasonable power in all scenarios. They outperform the adjusted approaches in some non-proportional settings in terms of power.

stat.ME

An unbiased rank-based estimator of the Mann-Whitney variance including the case of ties

Many estimators of the variance of the well-known unbiased and uniform most powerful estimator $\htheta$ of the Mann-Whitney effect, $θ= P(X < Y) + \nfrac12 P(X=Y)$, are considered in the literature. Some of these estimators are only valid in case of no ties or are biased in case of small sample sizes where the amount of the bias is not discussed. Here we derive an unbiased estimator that is based on different rankings, the so-called 'placements' (Orban and Wolfe, 1980), and is therefore easy to compute. This estimator does not require the assumption of continuous \dfs\ and is also valid in the case of ties. Moreover, it is shown that this estimator is non-negative and has a sharp upper bound which may be considered an empirical version of the well-known Birnbaum-Klose inequality. The derivation of this estimator provides an option to compute the biases of some commonly used estimators in the literature. Simulations demonstrate that, for small sample sizes, the biases of these estimators depend on the underlying \dfs\ and thus are not under control. This means that in the case of a biased estimator, simulation results for the type-I error of a test or the coverage probability of a \ci\ do not only depend on the quality of the approximation of $\htheta$ by a normal \db\ but also an additional unknown bias caused by the variance estimator. Finally, it is shown that this estimator is $L_2$-consistent.

stat.ME

Unlocking Insights: Enhanced Analysis of Covariance in General Factorial Designs through Multiple Contrast Tests under Variance Heteroscedasticity

A common goal in clinical trials is to conduct tests on estimated treatment effects adjusted for covariates such as age or sex. Analysis of Covariance (ANCOVA) is often used in these scenarios to test the global null hypothesis of no treatment effect using an $F$-test. However, in several samples, the $F$-test does not provide any information about individual null hypotheses and has strict assumptions such as variance homoscedasticity. We extend the method proposed by Konietschke et al. (2021) to a multiple contrast test procedure (MCTP), which allows us to test arbitrary linear hypotheses and provides information about the global as well as the individual null hypotheses. Further, we can calculate compatible simultaneous confidence intervals for the individual effects. We derive a small sample size approximation of the distribution of the test statistic via a multivariate t-distribution. As an alternative, we introduce a Wild-bootstrap method. Extensive simulations show that our methods are applicable even when sample sizes are small. Their application is further illustrated within a real data example.

stat.ME

A studentized permutation test in group sequential designs

In group sequential designs, where several data looks are conducted for early stopping, we generally assume the vector of test statistics from the sequential analyses follows (at least approximately or asymptotially) a multivariate normal distribution. However, it is well-known that test statistics for which an asymptotic distribution is derived may suffer from poor small sample approximation. This might become even worse with an increasing number of data looks. The aim of this paper is to improve the small sample behaviour of group sequential designs while maintaining the same asymptotic properties as classical group sequential designs. This improvement is achieved through the application of a modified permutation test. In particular, this paper shows that the permutation distribution approximates the distribution of the test statistics not only under the null hypothesis but also under the alternative hypothesis, resulting in an asymptotically valid permutation test. An extensive simulation study shows that the proposed permutation test better controls the Type I error rate than its competitors in the case of small sample sizes.

math.ST

Incompletely observed nonparametric factorial designs with repeated measurements: A wild bootstrap approach

In many life science experiments or medical studies, subjects are repeatedly observed and measurements are collected in factorial designs with multivariate data. The analysis of such multivariate data is typically based on multivariate analysis of variance (MANOVA) or mixed models, requiring complete data, and certain assumption on the underlying parametric distribution such as continuity or a specific covariance structure, e.g., compound symmetry. However, these methods are usually not applicable when discrete data or even ordered categorical data are present. In such cases, nonparametric rank-based methods that do not require stringent distributional assumptions are the preferred choice. However, in the multivariate case, most rank-based approaches have only been developed for complete observations. It is the aim of this work is to develop asymptotic correct procedures that are capable of handling missing values, allowing for singular covariance matrices and are applicable for ordinal or ordered categorical data. This is achieved by applying a wild bootstrap procedure in combination with quadratic form-type test statistics. Beyond proving their asymptotic correctness, extensive simulation studies validate their applicability for small samples. Finally, two real data examples are analyzed.

stat.ME

Simultaneous inference for partial areas under receiver operating curves -- with a view towards efficiency

We propose new simultaneous inference methods for diagnostic trials with elaborate factorial designs. Instead of the commonly used total area under the receiver operating characteristic (ROC) curve, our parameters of interest are partial areas under ROC curve segments that represent clinically relevant biomarker cut-off values. We construct a nonparametric multiple contrast test for these parameters and show that it asymptotically controls the family-wise type one error rate. Finite sample properties of this test are investigated in a series of computer experiments. We provide empirical and theoretical evidence supporting the conjecture that statistical inference about partial areas under ROC curves is more efficient than inference about the total areas.

math.ST

Using meta-analytic priors to incorporate external information for study evaluation

Background: The COVID-19 pandemic has had a profound impact on health, everyday life and economics around the world. An important complication that can arise in connection with a COVID-19 infection is acute kidney injury. A recent observational cohort study of COVID-19 patients treated at multiple sites of a tertiary care center in Berlin, Germany identified risk factors for the development of (severe) acute kidney injury. Since inferring results from a single study can be tricky, we validate these findings and potentially adjust results by including external information from other studies on acute kidney injury and COVID-19. Methods: We synthesize the results of the main study with other trials via a Bayesian meta-analysis. The external information is used to construct a predictive distribution and to derive posterior estimates for the study of interest. We focus on various important potential risk factors for acute kidney injury development such as mechanical ventilation, use of vasopressors, hypertension, obesity, diabetes, gender and smoking. Results: Our results show that depending on the degree of heterogeneity in the data the estimated effect sizes may be refined considerably with inclusion of external data. Our findings confirm that mechanical ventilation and use of vasopressors are important risk factors for the development of acute kidney injury in COVID-19 patients. Hypertension also appears to be a risk factor that should not be ignored. Shrinkage weights depended to a large extent on the estimated heterogeneity in the model. Conclusions: Our work shows how external information can be used to adjust the results from a primary study, using a Bayesian meta-analytic approach. How much information is borrowed from external studies will depend on the degree of heterogeneity present in the model.

stat.ME

Statistical Review of Animal trials -- A Guideline

Examinations of any experiment involving living organisms require justifications of the need and moral defensibleness of the study. Statistical planning, design and sample size calculation of the experiment are no less important review criteria than general medical and ethical points to consider. Errors made in the statistical planning and data evaluation phase can have severe consequences on both results and conclusions. They might proliferate and thus impact future trials-an unintended outcome of fundamental research with profound ethical consequences. Therefore, any trial must be efficient in both a medical and statistical way in answering the questions of interests to be considered as approvable. Unified statistical standards are currently missing for animal review boards in Germany. In order to accompany, we developed a biometric form to be filled and handed in with the proposal at the local authority on animal welfare. It addresses relevant points to consider for biostatistical planning of animal experiments and can help both the applicants and the reviewers in overseeing the entire experiment(s) planned. Furthermore, the form might also aid in meeting the current standards set by the 3+3R's principle of animal experimentation Replacement, Reduction, Refinement as well as Robustness, Registration and Reporting. The form has already been in use by the local authority of animal welfare in Berlin, Germany. In addition, we provide reference to our user guide giving more detailed explanation and examples for each section of the biometric form. Unifying the set of biostatistical aspects will help both the applicants and the reviewers to equal standards and increase quality of preclinical research projects, also for translational, multicenter, or international studies.

stat.AP

The Behrens-Fisher Problem with Covariates and Baseline Adjustments

The Welch-Satterthwaite t-test is one of the most prominent and often used statistical inference method in applications. The method is, however, not flexible with respect to adjustments for baseline values or other covariates, which may impact the response variable. Existing analysis of covariance methods are typically based on the assumption of equal variances across the groups. This assumption is hard to justify in real data applications and the methods tend to not control the type-1 error rate satisfactorily under variance heteroscedasticity. In the present paper, we tackle this problem and develop unbiased variance estimators of group specific variances, and especially of the variance of the estimated adjusted treatment effect in a general analysis of covariance model. These results are used to generalize the Welch-Satterthwaite t-test to covariates adjustments. Extensive simulation studies show that the method accurately controls the nominal type-1 error rate, even for very small sample sizes, moderately skewed distributions and under variance heteroscedasticity. A real data set motivates and illustrates the application of the proposed methods.

stat.ME

Ranks and Pseudo-Ranks - Paradoxical Results of Rank Tests -

Rank-based inference methods are applied in various disciplines, typically when procedures relying on standard normal theory are not justifiable, for example when data are not symmetrically distributed, contain outliers, or responses are even measured on ordinal scales. Various specific rank-based methods have been developed for two and more samples, and also for general factorial designs (e.g., Kruskal-Wallis test, Jonckheere-Terpstra test). It is the aim of the present paper (1) to demonstrate that traditional rank-procedures for several samples or general factorial designs may lead to paradoxical results in case of unbalanced samples, (2) to explain why this is the case, and (3) to provide a way to overcome these disadvantages of traditional rankbased inference. Theoretical investigations show that the paradoxical results can be explained by carefully considering the non-centralities of the test statistics which may be non-zero for the traditional tests in unbalanced designs. These non-centralities may even become arbitrarily large for increasing sample sizes in the unbalanced case. A simple solution is the use of socalled pseudo-ranks instead of ranks. As a special case, we illustrate the effects in sub-group analyses which are often used when dealing with rare diseases.

math.ST

Multiplication-Combination Tests for Incomplete Paired Data

We consider statistical procedures for hypothesis testing of real valued functionals of matched pairs with missing values. In order to improve the accuracy of existing methods, we propose a novel multiplication combination procedure. Dividing the observed data into dependent (completely observed) pairs and independent (incompletely observed) components, it is based on combining separate results of adequate tests for the two sub datasets. Our methods can be applied for parametric as well as semi- and nonparametric models and make efficient use of all available data. In particular, the approaches are flexible and can be used to test different hypotheses in various models of interest. This is exemplified by a detailed study of mean- as well as rank-based apporaches. Extensive simulations show that the proposed procedures are more accurate than existing competitors. A real data set illustrates the application of the methods.

math.ST

Analysis of Multivariate Data and Repeated Measures Designs with the R Package MANOVA.RM

The numerical availability of statistical inference methods for a modern and robust analysis of longitudinal- and multivariate data in factorial experiments is an essential element in research and education. While existing approaches that rely on specific distributional assumptions of the data (multivariate normality and/or characteristic covariance matrices) are implemented in statistical software packages, there is a need for user-friendly software that can be used for the analysis of data that do not fulfill the aforementioned assumptions and provide accurate p-value and confidence interval estimates. Therefore, newly developed statistical methods for the analysis of repeated measures designs and multivariate data that neither assume multivariate normality nor specific covariance matrices have been implemented in the freely available R-package MANOVA.RM. The package is equipped with a graphical user interface for plausible applications in academia and other educational purpose. Several motivating examples illustrate the application of the methods.

stat.CO

Wild Bootstrapping Rank-Based Procedures: Multiple Testing in Nonparametric Split-Plot Designs

Split-plot or repeated measures designs are frequently used for planning experiments in the life or social sciences. Typical examples include the comparison of different treatments over time, where both factors may possess an additional factorial structure. For such designs, the statistical analysis usually consists of several steps. If the global null is rejected, multiple comparisons are usually performed. Usually, general factorial repeated measures designs are inferred by classical linear mixed models. Common underlying assumptions, such as normality or variance homogeneity are often not met in real data. Furthermore, to deal even with, e.g., ordinal or ordered categorical data, adequate effect sizes should be used. Here, multiple contrast tests and simultaneous confidence intervals for general factorial split-plot designs are developed and equipped with a novel asymptotically correct wild bootstrap approach. Because the regulatory authorities typically require the calculation of confidence intervals, this work also provides simultaneous confidence intervals for single contrasts and for the ratio of different contrasts in meaningful effects. Extensive simulations are conducted to foster the theoretical findings. Finally, two different datasets exemplify the applicability of the novel procedure.

math.ST

A studentized permutation test for three-arm trials in the 'gold standard' design

The 'gold standard' design for three-arm trials refers to trials with an active control and a placebo control in addition to the experimental treatment group. This trial design is recommended when being ethically justifiable and it allows the simultaneous comparison of experimental treatment, active control, and placebo. Parametric testing methods have been studied plentifully over the past years. However, these methods often tend to be liberal or conservative when distributional assumptions are not met particularly with small sample sizes. In this article, we introduce a studentized permutation test for testing non-inferiority and superiority of the experimental treatment compared to the active control in three-arm trials in the `gold standard' design. The performance of the studentized permutation test for finite sample sizes is assessed in a Monte-Carlo simulation study under various parameter constellations. Emphasis is put on whether the studentized permutation test meets the target significance level. For comparison purposes, commonly used Wald-type tests are included in the simulation study. The simulation study shows that the presented studentized permutation test for assessing non-inferiority in three-arm trials in the 'gold standard' design outperforms its competitors for count data. The methods discussed in this paper are implemented in the R package ThreeArmedTrials which is available on the comprehensive R archive network (CRAN).

stat.AP

Rank-Based Procedures in Factorial Designs: Hypotheses about Nonparametric Treatment Effects

Existing tests for factorial designs in the nonparametric case are based on hypotheses formulated in terms of distribution functions. Typical null hypotheses, however, are formulated in terms of some parameters or effect measures, particularly in heteroscedastic settings. Here this idea is extended to nonparametric models by introducing a novel nonparametric ANOVA-type-statistic based on ranks which is suitable for testing hypotheses formulated in meaningful nonparametric treatment effects in general factorial designs. This is achieved by a careful in-depth study of the common distribution of rank-based estimators for the treatment effects. Since the statistic is asymptotically not a pivotal quantity we propose three different approximation techniques, discuss their theoretic properties and compare them in extensive simulations together with two additionalWald-type tests. An extension of the presented idea to general repeated measures designs is briefly outlined. The proposed rank-based procedures maintain the pre-assigned type-I error rate quite accurately, also in unbalanced and heteroscedastic models.

stat.ME

Using EEG, SPECT, and Multivariate Resampling Methods to Differentiate Between Alzheimer's and other Cognitive Impairments

The incidence of Alzheimer's disease (AD) and other forms of dementia is increasing in most western countries. For a precise and early diagnosis, several examination modalities exist, among them single-photon emission computed tomography (SPECT) and the electroencephalogram (EEG). The latter is highly available, free of radiation hazards, and non-invasive. Thus, its diagnostic utility regarding different stages of dementia is of great interest in neurological research, along with the question of whether its utility depends on age or sex of the person being examined. However, SPECT or EEG measurements are intrinsically multivariate, and there has been a shortage of sufficiently general inferential techniques for the analysis of multivariate data in factorial designs when neither multivariate normality nor equality of covariance matrices across groups should be assumed. We adapt an asymptotic model based (parametric) bootstrap approach to this situation, demonstrate its ability, and use it for a truly multivariate analysis of the EEG and SPECT measurements, taking into account demographic factors such as age and sex. These multivariate results are supplemented by marginal effects bootstrap inference whose theoretical properties can be derived analogously to the multivariate methods. Both inference approaches can have advantages in particular situations, as illustrated in the data analysis.

stat.AP