Searcharxiv⌕ Search

arXiv subjects

Ludwig A. Hothorn

Publications and source records attributed to Ludwig A. Hothorn.

At least 19 recordsLinked to original sources

Simultaneous comparisons of the variances of k treatments with that of a control: a Levene-Dunnett type procedure

There are some global tests for heterogeneity of variance in k-sample one-way layouts, but few consider pairwise comparisons between treatment levels. For experimental designs with a control, comparisons of the variances between the treatment levels and the control are of interest - in analogy to the location parameter with the Dunnett (1955) procedure. Such a many-to-one approach for variances is proposed using the Levene transformation, a kind of residuals. Its properties are characterized with simulation studies and corresponding data examples are evaluated with R code.

stat.ME↗

Unlocking Insights: Enhanced Analysis of Covariance in General Factorial Designs through Multiple Contrast Tests under Variance Heteroscedasticity

A common goal in clinical trials is to conduct tests on estimated treatment effects adjusted for covariates such as age or sex. Analysis of Covariance (ANCOVA) is often used in these scenarios to test the global null hypothesis of no treatment effect using an $F$-test. However, in several samples, the $F$-test does not provide any information about individual null hypotheses and has strict assumptions such as variance homoscedasticity. We extend the method proposed by Konietschke et al. (2021) to a multiple contrast test procedure (MCTP), which allows us to test arbitrary linear hypotheses and provides information about the global as well as the individual null hypotheses. Further, we can calculate compatible simultaneous confidence intervals for the individual effects. We derive a small sample size approximation of the distribution of the test statistic via a multivariate t-distribution. As an alternative, we introduce a Wild-bootstrap method. Extensive simulations show that our methods are applicable even when sample sizes are small. Their application is further illustrated within a real data example.

stat.ME↗

Bartholomew's trend test -- approximated by a multiple contrast test

Bartholomew's trend test belongs to the broad class of isotonic regression models, specifically with a single qualitative factor, e.g. dose levels. Using the approximation of the ANOVA F-test by the maximum contrast test against grand mean and pool-adjacent-violator estimates under order restriction, an easier to use approximation is proposed.

stat.ME↗

Tests for strict monotonic trend in bio-medical dose-response relationships (respective concentration-response or exposure-response relationships) -- a biostatistical perspective

Evidence of a global trend in dose-response dependencies is commonly used in bio-medicine and epidemiology, especially because this represents a causality criterion. However, conventional trend tests indicate a significant trend even when dependence is in the opposite direction for low doses when the high dose alone has a superior effect. Here we present a trend test for a strictly monotonic increasing (or decreasing) trend, evaluate selected sample data for it, and provide corresponding R code using CRAN packages.

stat.AP↗

Consistent ANOVA-type tests for various effect sizes

Analysis of variance (ANOVA) reveals some disadvantages, such as non-robustness against heteroscedastic or non-normal errors and using difference to overall mean as effect sizes only. As an alternative the multiple contrast test comparing to the overall mean is proposed for 7 effect sizes: ratio-to-OM, quantiles for both ratio or differences, odds ratios for continuous data, odds ratio for proportions, risk ratio/differences, relative effect size for continuous up to discrete data, and hazard ratio. Using CRAN packages the related analysis is simple.

stat.ME↗

The Dunnett procedure with possibly heterogeneous variances

Most comparisons of treatments or doses against a control are performed by the original Dunnett single step procedure \cite{Dunnett1955} providing both adjusted p-values and simultaneous confidence intervals for differences to the control. Motivated by power arguments, unbalanced designs with higher sample size in the control are recommended. When higher variance occur in the treatment of interest or in the control, the related per-pairs power is reduced, as expected. However, if the variance is increased in a non-affected treatment group, e.g. in the highest dose (which is highly significant), the per-pairs power is also reduced in the remaining treatment groups of interest. I.e., decisions about the significance of certain comparisons may be seriously distorted. To avoid this nasty property, three modifications for heterogeneous variances are compared by a simulation study with the original Dunnett procedure. For small and medium sample sizes, a Welch-type modification can be recommended. For medium to high sample sizes, the use of a sandwich estimator instead of the common mean square estimator is useful. Related CRAN packages are provided. Summarizing we recommend not to use the original Dunnett procedure in routine and replace it by a robust modification. Particular care is needed in small sample size studies.

stat.ME↗

The Kruskal Wallis test can not be recommended

Although the Kruskal-Wallis (KW) test is widely used, it should not be recommended: it is not robust to arbitrary alternatives, it is only a global test without confidence intervals for the marginal hypotheses, it is inherently defined for two-sided hypotheses, it is not very suitable for pre/post hoc test combinations and hard to modified for factorial designs or the analysis of covariance. As an alternative a double maximum test is proposed: a maximum over multiple contrasts against the grand mean (approximating global power as a linear test statistics) and a maximum over three rank scores, sensitive for location, scale and shape effects. The joint distribution of this new test is achieved by the multiple marginal models approach. Related R-code is provided.

stat.ME↗

Seven recommendations for alternatives to the common analysis of variance (ANOVA) with application in the life sciences -- using R

Standard ANOVA is among the most widely used tests in the life sciences and beyond. Several alternatives are proposed to provide simultaneous confidence intervals, ensure tight control of FWER, be robust to variance heterogeneity, avoid pre-testing for global effect (for one-way designs) or irrelevant interaction (for multi-way designs) prior to multiple comparison procedures. On the basis of selected examples, the R programs are provided for this purpose.

stat.ME↗

Evaluation of histological findings with severity grade, to analyze toxicology in-vivo studies

In-vivo toxicological studies are characterized by multiple primary endpoints with quite different scales. Whereas guidelines and publications provide various statistical tests for normally distributed endpoints (such as organ weights) and proportions (such as tumor rates), few approaches are available for graded histopathological findings, such as 0, +, ++, +++. This represents a basic contradiction of the statistical analysis because these graded findings sometimes show a high predictive value for potential toxic effects. Here we discuss different methods comparatively, especially from the viewpoints of i) designs for very small sample sizes and ii) interpretability by toxicologists. A new approach is recommended where a simultaneous test is performed over all class combinations of score levels, such as (0, +) vs (++, +++). Corresponding R code is provided by way of a data example.

stat.AP↗

Simultaneous inference of correlated marginal tests using intersection-union or union-intersection test principle

Two main approaches in simultaneous inference are intersection-union tests and union-intersection tests. For intersection-union hypotheses, the classical IUT based on marginal p-values and the all-in-alternative UIT are compared. Depending on correlation, number of marginal tests and patterns of the alternative the inherent power loss of the aiaUIT seems to be acceptable, considering its advantage, namely the availability of simple-to-interpret simultaneous confidence interval.

stat.ME↗

Simultaneous comparisons of treatments versus control (Dunnett-type tests) for location-scale alternatives

Commonly, the comparisons of treatment groups versus a control is performed for location effects only where possible scale effects are considered as disturbing. Sometimes scale effects are also relevant, as a kind of early indicator for changes. Here several approaches for Dunnett-type tests for location or scale effects are proposed and compared by a simulation study. Two real data examples are analysed accordingly and the related R-code is available in the Appendix.

stat.ME↗

A statistical method for estimating the no-observed-adverse-event-level

In toxicological risk assessment the benchmark dose (BMD) is recommended instead of the no-observed-adverse effect-level (NOAEL). Still a simple test procedure to estimate NOAEL is proposed here, explaining its advantages and disadvantages. Versatile applicability is illustrated using four different data examples of selected in vivo toxicity bioassays.

stat.AP↗

Comparisons of proportions in k dose groups against a negative control assuming order restriction: Williams-type test vs. closed test procedures

The comparison of proportions is considered in the asymptotic generalized linear model with the odds ratio as effect size. When several doses are compared with a control assuming an order restriction, a Williams-type trend test can be used. As an alternative, two variants of the closed testing approach are considered, one using global Williams-tests in the partition hypotheses, one with pairwise contrasts. Their advantages in terms of power and simplicity are demonstrated. Related R-code is provided.

stat.AP↗

Statistical evaluation of in-vivo bioassays in regulatory toxicology considering males and females

The separate evaluation for males and females is the recent standard in in-vivo toxicology for dose or treatment effects using Dunnett tests. The alternative pre-test for sex-by-treatment interaction is problematic. Here a joint test is proposed considering the two sex-specific and the pooled Dunnett-type comparisons. The calculation of either simultaneous confidence intervals or adjusted p-values with the R-package multcomp is demonstrated using a real data example.

stat.AP↗