SearcharxivSearch

arXiv subjects

Lubna Amro

Publications and source records attributed to Lubna Amro.

6 recordsLinked to original sources

Generalized multivariate Mann-Whitney-$U$ tests and confidence regions for relative effects under random missingness

Marginal Mann-Whitney effects are widely used across various fields of research, and extensions of this estimand have been developed in many directions in statistical methodology. In this paper, we focus on an extensions for repeated measurements and factorial designs subject to randomly missing data. In a previous work by Rubarth et al. (2022a), asymptotically correct tests were developed under the assumption of deterministic missing indicators. In contrast, the approach in the present paper accounts for the stochastic nature of missing values under realistic mechanisms. Thus, the involved covariance matrix incorporates the true variability of missing data. The combination with a randomization procedure using random permutations within each data point yields asymptotically exact tests and a generally improved type-I error control. Additionally, the tests control the type-I error for finite sample sizes in the special case of exchangeable sampling distributions. Simulations across a wide range of settings demonstrate the benefits of the proposed method in small samples, also for different missingness mechanisms. A real data analysis about school children learning math illustrates several practical aspects of the tests' application.

stat.ME

Incompletely observed nonparametric factorial designs with repeated measurements: A wild bootstrap approach

In many life science experiments or medical studies, subjects are repeatedly observed and measurements are collected in factorial designs with multivariate data. The analysis of such multivariate data is typically based on multivariate analysis of variance (MANOVA) or mixed models, requiring complete data, and certain assumption on the underlying parametric distribution such as continuity or a specific covariance structure, e.g., compound symmetry. However, these methods are usually not applicable when discrete data or even ordered categorical data are present. In such cases, nonparametric rank-based methods that do not require stringent distributional assumptions are the preferred choice. However, in the multivariate case, most rank-based approaches have only been developed for complete observations. It is the aim of this work is to develop asymptotic correct procedures that are capable of handling missing values, allowing for singular covariance matrices and are applicable for ordinal or ordered categorical data. This is achieved by applying a wild bootstrap procedure in combination with quadratic form-type test statistics. Beyond proving their asymptotic correctness, extensive simulation studies validate their applicability for small samples. Finally, two real data examples are analyzed.

stat.ME

Asymptotic based bootstrap approach for matched pairs with missingness in a single-arm

The issue of missing values is an arising difficulty when dealing with paired data. Several test procedures are developed in the literature to tackle this problem. Some of them are even robust under deviations and control type-I error quite accurately. However, most these methods are not applicable when missing values are present only in a single arm. For this case, we provide asymptotic correct resampling tests that are robust under heteroscedasticity and skewed distributions. The tests are based on a clever restructuring of all observed information in a quadratic form-type test statistic. An extensive simulation study is conducted exemplifying the tests for finite sample sizes under different missingness mechanisms. In addition, an illustrative data example based on a breast cancer gene study is analyzed.

stat.ME

A cautionary tale on using imputation methods for inference in matched pairs design

Imputation procedures in biomedical fields have turned into statistical practice, since further analyses can be conducted ignoring the former presence of missing values. In particular, non-parametric imputation schemes like the random forest or a combination with the stochastic gradient boosting have shown favorable imputation performance compared to the more traditionally used MICE procedure. However, their effect on valid statistical inference has not been analyzed so far. This paper closes this gap by investigating their validity for inferring mean differences in incompletely observed pairs while opposing them to a recent approach that only works with the given observations at hand. Our findings indicate that machine learning schemes for (multiply) imputing missing values may inflate type-I-error or result in comparably low power in small to moderate matched pairs, even after modifying the test statistics using Rubin's multiple imputation rule. In addition to an extensive simulation study, an illustrative data example from a breast cancer gene study has been considered.

stat.AP

Multiplication-Combination Tests for Incomplete Paired Data

We consider statistical procedures for hypothesis testing of real valued functionals of matched pairs with missing values. In order to improve the accuracy of existing methods, we propose a novel multiplication combination procedure. Dividing the observed data into dependent (completely observed) pairs and independent (incompletely observed) components, it is based on combining separate results of adequate tests for the two sub datasets. Our methods can be applied for parametric as well as semi- and nonparametric models and make efficient use of all available data. In particular, the approaches are flexible and can be used to test different hypotheses in various models of interest. This is exemplified by a detailed study of mean- as well as rank-based apporaches. Extensive simulations show that the proposed procedures are more accurate than existing competitors. A real data set illustrates the application of the methods.

math.ST

Permuting Incomplete Paired Data: A Novel Exact and Asymptotic Correct Randomization Test

Various statistical tests have been developed for testing the equality of means in matched pairs with missing values. However, most existing methods are commonly based on certain distributional assumptions such as normality, 0-symmetry or homoscedasticity of the data. The aim of this paper is to develop a statistical test that is robust against deviations from such assumptions and also leads to valid inference in case of heteroscedasticity or skewed distributions. This is achieved by applying a novel randomization approach. The resulting test procedure is not only shown to be asymptotically correct but is also finitely exact if the distribution of the data is invariant with respect to the considered randomization group. Its small sample performance is further studied in an extensive simulation study and compared to existing methods. Finally, an illustrative data example is analyzed.

math.ST