SearcharxivSearch

arXiv subjects

Paavo Sattler

Publications and source records attributed to Paavo Sattler.

15 recordsLinked to original sources

Generalized multivariate Mann-Whitney-$U$ tests and confidence regions for relative effects under random missingness

Marginal Mann-Whitney effects are widely used across various fields of research, and extensions of this estimand have been developed in many directions in statistical methodology. In this paper, we focus on an extensions for repeated measurements and factorial designs subject to randomly missing data. In a previous work by Rubarth et al. (2022a), asymptotically correct tests were developed under the assumption of deterministic missing indicators. In contrast, the approach in the present paper accounts for the stochastic nature of missing values under realistic mechanisms. Thus, the involved covariance matrix incorporates the true variability of missing data. The combination with a randomization procedure using random permutations within each data point yields asymptotically exact tests and a generally improved type-I error control. Additionally, the tests control the type-I error for finite sample sizes in the special case of exchangeable sampling distributions. Simulations across a wide range of settings demonstrate the benefits of the proposed method in small samples, also for different missingness mechanisms. A real data analysis about school children learning math illustrates several practical aspects of the tests' application.

stat.ME

Covariance Correction for Permutation Statistics in Multiple Testing Problems

In qualitative statistics, permutation tests are very popular, mainly because of their finite-sample exactness under exchangeability. However, in non-exchangeable settings, the covariance structure of permuted statistics typically differs from that of the original statistic. A common solution is studentization, which restores asymptotic correctness for general hypotheses while preserving exactness under exchangeability. In multiple testing settings, however, standard studentization fails to provide the correct joint limiting distribution. Existing solutions such as prepivoting address this issue but are computationally expensive and therefore rarely used in practice. We propose a general, computationally more efficient methodology that overcomes this fundamental limitation. By appropriately correcting the covariance matrix of multiple permutation statistics, our approach restores the correct joint asymptotic dependence structure, enabling asymptotically valid permutation tests in broad multiple testing frameworks. The proposed method is highly flexible: it accommodates singular covariance structures and is not tied to specific parameters, test statistics, or permutation schemes. This generality makes it applicable across a wide range of problems. Extensive simulation studies demonstrate that our approach results in reliable inference and outperforms existing methods across diverse settings.

stat.ME

Inference for high dimensional repeated measure designs with the R package hdrm

Repeated-measure designs allow comparisons within a group as well as between groups, and are commonly referred to as split-plot designs. While originating in agricultural experiments, they are now widely used in medical research, psychology, and the life sciences, where repeated observations on the same subject are essential. Modern data collection often produces observation vectors with dimension $d$ comparable to or exceeding the sample size $N$. Although this can be advantageous in terms of cost efficiency, ethical considerations, and the study of rare diseases, it poses substantial challenges for statistical inference. Parametric methods based on multivariate normality provide a flexible framework that avoids restrictive assumptions on covariance structures or on the asymptotic relationship between $d$ and $N$. Within this framework, the freely available R-package hdrm enables the analysis of a wide range of hypotheses concerning expectation vectors in high-dimensional repeated-measure designs, covering both single-group and multi-group settings with homogeneous or heterogeneous covariance matrices. This paper describes the implemented tests, demonstrates their use through examples, and discusses their applicability in practical high-dimensional data scenarios. To address computational challenges arising for large $d$, the package incorporates efficient estimators and subsampling strategies that substantially reduce computation time while preserving statistical validity.

stat.CO

Testing Hypotheses regarding Covariance and Correlation matrices with the R package CovCorTest

In addition to the commonly analyzed measures of location, dispersion measurements such as variance and correlation provide many valuable information. Consequently, they play a crucial role in multivariate statistics, which leads to tests regarding covariance and correlation matrices. Furthermore, also the structure of these matrices leads to important hypotheses of interest, since it contains substantial information about the underlying model. In fact, assumptions regarding the structures of covariance and correlation matrices are often fundamental in statistical modelling and testing. In this context, semi-parametric settings with minimal distributional assumptions and very general hypotheses are essential for enabling manifold usage. The free available package CovCorTest provides suitable tests addressing all aforementioned issues, using bootstrap and similar techniques to achieve good performance, particularly in small samples. Additionally, the package offers flexible specification options for the hypotheses under investigation in two central tests, accommodating users with varying levels of expertise, which results in high flexibility and user-friendliness at the same time. This paper also presents the application of \textbf{CovCorTest} for various issues, illustrated by multiple examples, where the tests are applied to a real-world dataset.

stat.CO

Multivariate and Multiple Contrast Testing in General Covariate-adjusted Factorial Designs

Evaluating intervention effects on multiple outcomes is a central research goal in a wide range of quantitative sciences. It is thereby common to compare interventions among each other and with a control across several, potentially highly correlated, outcome variables. In this context, researchers are interested in identifying effects at both, the global level (across all outcome variables) and the local level (for specific variables). At the same time, potential confounding must be accounted for. This leads to the need for powerful multiple contrast testing procedures (MCTPs) capable of handling multivariate outcomes and covariates. Given this background, we propose an extension of MCTPs within a semiparametric MANCOVA framework that allows applicability beyond multivariate normality, homoscedasticity, or non-singular covariance structures. We illustrate our approach by analysing multivariate psychological intervention data, evaluating joint physiological and psychological constructs such as heart rate variability.

stat.ME

Quadratic Form based Multiple Contrast Tests for Comparison of Group Means

Comparing the mean vectors across different groups is a cornerstone in the realm of multivariate statistics, with quadratic forms commonly serving as test statistics. However, when the overall hypothesis is rejected, identifying specific vector components or determining the groups among which differences exist requires additional investigations. Conversely, employing multiple contrast tests (MCT) allows conclusions about which components or groups contribute to these differences. However, they come with a trade-off, as MCT lose some benefits inherent to quadratic forms. In this paper, we combine both approaches to get a quadratic form based multiple contrast test that leverages the advantages of both. To understand its theoretical properties, we investigate its asymptotic distribution in a semiparametric model. We thereby focus on two common quadratic forms - the Wald-type statistic and the Anova-type statistic - although our findings are applicable to any quadratic form. Furthermore, we employ Monte-Carlo and resampling techniques to enhance the test's performance in small sample scenarios. Through an extensive simulation study, we assess the performance of our proposed tests against existing alternatives, highlighting their advantages.

stat.ME

Testing for patterns and structures in covariance and correlation matrices

Covariance matrices of random vectors contain information that is crucial for modelling. Specific structures and patterns of the covariances (or correlations) may be used to justify parametric models, e.g., autoregressive models. Until now, there have been only a few approaches for testing such covariance structures and most of them can only be used for one particular structure. In the present paper, we propose a systematic and unified testing procedure working among others for the large class of linear covariance structures. Our approach requires only weak distributional assumptions. It covers common structures such as diagonal matrices, Toeplitz matrices and compound symmetry, as well as the more involved autoregressive matrices. We exemplify the approach for all these structures. We prove the correctness of these tests for large sample sizes and use bootstrap techniques for a better small-sample approximation. Moreover, the proposed tests invite adaptations to other covariance patterns by choosing the hypothesis matrix appropriately. With the help of a simulation study, we also assess the small sample properties of the tests. Finally, we illustrate the procedure in an application to a real data set.

stat.ME

Resampling NANCOVA: Nonparametric Analysis of Covariance in Small Samples

Analysis of covariance is a crucial method for improving precision of statistical tests for factor effects in randomized experiments. However, existing solutions suffer from one or more of the following limitations: (i) they are not suitable for ordinal data (as endpoints or explanatory variables); (ii) they require semiparametric model assumptions; (iii) they are inapplicable to small data scenarios due to often poor type-I error control; or (iv) they provide only approximate testing procedures and (asymptotically) exact test are missing. In this paper, we investigate a resampling approach to the NANCOVA framework, which is a fully nonparametric model based on relative effects that allows for an arbitrary number of covariates and groups, where both outcome variable (endpoint) and covariates can be metric or ordinal. Thereby, we evaluate novel NANCOVA tests and a nonparametric competitor test without covariate adjustment in extensive simulations. Unlike approximate tests in the NANCOVA framework, our resampling version showed good performance in small sample scenarios and maintained the nominal type-I error well. Resampling NANCOVA also provided consistently high power: up to 26% higher than the test without covariate adjustment in a small sample scenario with 4 groups and two covariates. Moreover, we prove that resampling NANCOVA provides an asymptotically exact testing procedure, which makes it the first one in the NANCOVA framework. In summary, resampling NANCOVA can be considered a viable tool for analysis of covariance that overcomes issues (i) - (iv).

stat.ME

Choice of the hypothesis matrix for using the Anova-type-statistic

Initially developed in Brunner et al. (1997), the Anova-type-statistic (ATS) is one of the most used quadratic forms for testing multivariate hypotheses for a variety of different parameter vectors $\boldsymbolθ\in\mathbb{R}^d$. Such tests can be based on several versions of ATS and in most settings, they are preferable over those based on other quadratic forms, as for example the Wald-type-statistic (WTS). However, the same null hypothesis $\boldsymbol{H}\boldsymbolθ=\boldsymbol{y}$ can be expressed by a multitude of hypothesis matrices $\boldsymbol{H}\in\mathbb{R}^{m\times d}$ and corresponding vectors $\boldsymbol{y}\in\mathbb{R}^m$, which leads to different values of the test statistic, as it can be seen in simple examples. Since this can entail distinct test decisions, it remains to investigate under which conditions tests using different hypothesis matrices coincide. Here, the dimensions of the different hypothesis matrices can be substantially different, which has exceptional potential to save computation effort. In this manuscript, we show that for the Anova-type-statistic and some versions thereof, it is possible for each hypothesis $\boldsymbol{H}\boldsymbolθ=\boldsymbol{y}$ to construct a companion matrix $\boldsymbol{L}$ with a minimal number of rows, which not only tests the same hypothesis but also always yields the same test decisions. This allows a substantial reduction of computation time, which is investigated in several conducted simulations.

stat.ME

Choice of the hypothesis matrix for using the Wald-type-statistic

A widely used formulation for null hypotheses in the analysis of multivariate $d$-dimensional data is $\mathcal{H}_0: \boldsymbol{H} \boldsymbolθ =\boldsymbol{y}$ with $\boldsymbol{H}$ $\in\mathbb{R}^{m\times d}$, $\boldsymbolθ$ $\in \mathbb{R}^d$ and $\boldsymbol{y}\in\mathbb{R}^m$, where $m\leq d$. Here the unknown parameter vector $\boldsymbolθ$ can, for example, be the expectation vector $\boldsymbolμ$, a vector $\boldsymbolβ $ containing regression coefficients or a quantile vector $\boldsymbol{q}$. Also, the vector of nonparametric relative effects $\boldsymbol{p}$ or an upper triangular vectorized covariance matrix $\textbf{v}$ are useful choices. However, even without multiplying the hypothesis with a scalar $γ\neq 0$, there is a multitude of possibilities to formulate the same null hypothesis with different hypothesis matrices $\boldsymbol{H}$ and corresponding vectors $\boldsymbol{y}$. Although it is a well-known fact that in case of $\boldsymbol{y}=\boldsymbol{0}$ there exists a unique projection matrix $\boldsymbol{P}$ with $\boldsymbol{H}\boldsymbolθ=\boldsymbol{0}\Leftrightarrow \boldsymbol{P}\boldsymbolθ=\boldsymbol{0}$, for $\boldsymbol{y}\neq \boldsymbol{0}$ such a projection matrix does not necessarily exist. Moreover, since such hypotheses are often investigated using a quadratic form as the test statistic, the corresponding projection matrices often contain zero rows; so, they are not even effective from a computational aspect. In this manuscript, we show that for the Wald-type-statistic (WTS), which is one of the most frequently used quadratic forms, the choice of the concrete hypothesis matrix does not affect the test decision. Moreover, some simulations are conducted to investigate the possible influence of the hypothesis matrix on the computation time.

math.ST

Testing Hypotheses about Correlation Matrices in General MANOVA Designs

Correlation matrices are an essential tool for investigating the dependency structures of random vectors or comparing them. We introduce an approach for testing a variety of null hypotheses that can be formulated based upon the correlation matrix. Examples cover MANOVA-type hypothesis of equal correlation matrices as well as testing for special correlation structures such as, e.g., sphericity. Apart from existing fourth moments, our approach requires no other assumptions, allowing applications in various settings. To improve the small sample performance, a bootstrap technique is proposed and theoretically justified. Based on this, we also present a procedure to simultaneously test the hypotheses of equal correlation and equal covariance matrices. The performance of all new test statistics is compared with existing procedures through extensive simulations.

math.ST

Inference for high-dimensional split-plot designs with different dimensions between groups

In repeated Measure Designs with multiple groups, the primary purpose is to compare different groups in various aspects. For several reasons, the number of measurements and therefore the dimension of the observation vectors can depend on the group, making the usage of existing approaches impossible. We develop an approach which can be used not only for a possibly increasing number of groups $a$, but also for group-depending dimension $d_i$, which is allowed to go to infinity. This is a unique high-dimensional asymptotic framework impressing through its variety and do without usual conditions on the relation between sample size and dimension. It especially includes settings with fixed dimensions in some groups and increasing dimensions in other ones, which can be seen as semi-high-dimensional. To find a appropriate statistic test new and innovative estimators are developed, which can be used under these diverse settings on $a,d_i$ and $n_i$ without any adjustments. We investigated the asymptotic distribution of a quadratic-form-based test statistic and developed an asymptotic correct test. Finally, an extensive simulation study is conducted to investigate the role of the single group's dimension.

math.ST

Testing Hypotheses about Covariance Matrices in General MANOVA Designs

We introduce a unified approach to testing a variety of rather general null hypotheses that can be formulated in terms of covariances matrices. These include as special cases, for example, testing for equal variances, equal traces, or for elements of the covariance matrix taking certain values. The proposed method only requires very few assumptions and thus promises to be of broad practical use. Two test statistics are defined, and their asymptotic or approximate sampling distributions are derived. In order to improve particularly the small-sample behavior of the resulting tests, two bootstrap-based methods are developed and theoretically justified. Several simulations shed light on the performance of the proposed tests. The analysis of a real data set illustrates the application of the procedures.

math.ST

Manifold Asymptotics of Quadratic-Form-Based Inference in Repeated Measures Designs

Split-Plot or Repeated Measures Designs with multiple groups occur naturally in sciences. Their analysis is usually based on the classical Repeated Measures ANOVA. Roughly speaking, the latter can be shown to be asymptotically valid for large sample sizes $n_i$ assuming a fixed number of groups $a$ and time points $d$. However, for high-dimensional settings with $d>n_i$ this argument breaks down and statistical tests are often based on (standardized) quadratic forms. Furthermore analysis of their limit behaviour is usually based on certain assumptions on how $d$ converges to $\infty$ with respect to $n_i$. As this may be hard to argue in practice, we do not want to make such restrictions. Moreover, sometimes also the number of groups $a$ may be large compared to $d$ or $n_i$. To also have an impression about the behaviour of (standardized) quadratic forms as test statistic, we analyze their asymptotics under diverse settings on $a$, $d$ and $n_i$. In fact, we combine all kind of combinations, where they diverge or are bounded in a unified framework. Studying the limit distributions in detail, we follow Sattler and Pauly (2018) and propose an approximation to obtain critical values. The resulting test together with their approximation approach are investigated in an extensive simulation study with a focus on the exceptional asymptotic frameworks which are the main focus of this work.

math.ST

Inference For High-Dimensional Split-Plot-Designs: A Unified Approach for Small to Large Numbers of Factor Levels

Statisticians increasingly face the problem to reconsider the adaptability of classical inference techniques. In particular, divers types of high-dimensional data structures are observed in various research areas; disclosing the boundaries of conventional multivariate data analysis. Such situations occur, e.g., frequently in life sciences whenever it is easier or cheaper to repeatedly generate a large number $d$ of observations per subject than recruiting many, say $N$, subjects. In this paper we discuss inference procedures for such situations in general heteroscedastic split-plot designs with $a$ independent groups of repeated measurements. These will, e.g., be able to answer questions about the occurrence of certain time, group and interactions effects or about particular profiles. The test procedures are based on standardized quadratic forms involving suitably symmetrized U-statistics-type estimators which are robust against an increasing number of dimensions $d$ and/or groups $a$. We then discuss its limit distributions in a general asymptotic framework and additionally propose improved small sample approximations. Finally its small sample performance is investigated in simulations and the applicability is illustrated by a real data analysis.

math.ST