SearcharxivSearch

arXiv subjects

Anna Vesely

Publications and source records attributed to Anna Vesely.

10 recordsLinked to original sources

Post-Selection Inference for Multiverse Analysis in Mixed-Effects Models (PIMAX)

Sign-flipping score tests provide robust inference in generalized linear models under variance misspecification and form the basis of two recent inferential frameworks: post-selection inference in multiverse analysis (PIMA) and the sign-flipping score-based two-stage summary-statistics approach (flip2sss). PIMA provides asymptotically valid inference across a multiverse of model specifications, whereas flip2sss extends sign-flipping score testing to longitudinal and hierarchical data through cluster-level summary statistics. In this paper, we combine these two approaches to develop PIMAX, a multiverse inferential framework for clustered observations. The resulting method extends post-selection inference to clustered-data settings, accommodating heteroscedasticity, unbalanced designs, and within-cluster dependence. Given a multiverse of candidate specifications, PIMAX provides a global p-value for testing whether any specification exhibits a non-zero effect (weak control of the family-wise error rate, FWER), lower confidence bounds on the number of true discoveries, and multiplicity-adjusted p-values for identifying the specific contributing specifications (strong FWER control). By avoiding inference based on a fully specified random-effects covariance structure, PIMAX solves a key source of type I error inflation due to random-effects misspecification while enabling inference across a multiverse of fixed-effects specifications.

stat.ME

Inference on multiple quantiles in regression models by a rank-score approach

This paper tackles the challenge of performing multiple quantile regressions across different quantile levels and the associated problem of controlling the familywise error rate, an issue that is generally overlooked in practice. We propose a multivariate extension of the rank-score test and embed it within a closed-testing procedure to efficiently account for multiple testing. Then we further generalize the multivariate test to enhance statistical power against alternatives in selected directions. Theoretical foundations and simulation studies demonstrate that our method effectively controls the familywise error rate while achieving higher power than traditional corrections, such as Bonferroni.

stat.ME

Partial Conjunction Analysis in Neuroimaging: A Comparative Study

Replicability is a cornerstone of science. The partial conjunction (PC) hypothesis testing framework objectively quantifies replicability across disciplines. Although several statistical methodologies for testing PC hypotheses exist, it is not clear which method performs well under which circumstances. In this paper, we consider the PC hypothesis testing problem from a neuroimaging perspective. Identifying the brain regions activated by a specific cognitive task constitutes a central challenge in neuroimaging. This problem becomes complex when the objective is to evaluate whether activation patterns are consistent across different cognitive tasks or subjects. In this paper, we cast this question as a PC hypothesis testing problem, assessing, for each location in the brain, whether it is activated in at least $\gamma$ subjects, for a pre-specified granularity $\gamma$. In our comparative study, we consider three methods, namely: adaFilter, CoFilter, and a method proposed by Benjamini, Heller, and Yekutieli (BHY). In equi-correlated simulated data, the BHY procedure tends to outperform the competing methods for high values of $\gamma$, while CoFilter performs well for low values of $\gamma$. In the real-data analysis, CoFilter dominates the other methods for intermediate values of $\gamma$.

stat.ME

Selective inference for fMRI cluster-wise analysis, issues, and recommendations for critical vector selection: A comment on Blain et al

Two permutation-based methods for simultaneous inference on the proportion of active voxels in cluster-wise brain imaging analysis have recently been published: Notip (Blain et al. 2022) and pARI (Andreella et al. 2023). Both rely on the definition of a critical vector of ordered p-values, chosen from a family of candidate vectors, but differ in how the family is defined: computed from randomization of external data for Notip and determined a priori for pARI. These procedures were compared to other proposals in the literature, but an extensive comparison between the two methods is missing due to their parallel publication. We provide such a comparison and find that pARI outperforms Notip if both methods are applied under their recommended settings. However, each method carries different advantages and drawbacks.

stat.AP

Utilizing Multiple Testing for Grouping in Singular Spectrum Analysis

A key step in separating signal from noise in time series by means of singular spectrum analysis (SSA) is grouping. We present a multiple testing method for the grouping step in SSA. As separability criterion, we utilize the weighted correlation between the signal and the noise component of the (reconstructed) time series, and we test whether this weighted correlation is equal to zero. This test has to be performed for several possible groupings, resulting in a multiple test problem. The null distributions of the corresponding test statistics are approximated by a wild bootstrap procedure. The performance of our proposed method is assessed in a simulation study, and we illustrate its practical application with an analysis of real world data.

stat.ME

Confidence bounds for the true discovery proportion based on the exact distribution of the number of rejections

In multiple hypotheses testing it has become widely popular to make inference on the true discovery proportion (TDP) of a set $\mathcal{M}$ of null hypotheses. This approach is useful for several application fields, such as neuroimaging and genomics. Several procedures to compute simultaneous lower confidence bounds for the TDP have been suggested in prior literature. Simultaneity allows for post-hoc selection of $\mathcal{M}$. If sets of interest are specified a priori, it is possible to gain power by removing the simultaneity requirement. We present an approach to compute lower confidence bounds for the TDP if the set of null hypotheses is defined a priori. The proposed method determines the bounds using the exact distribution of the number of rejections based on a step-up multiple testing procedure under independence assumptions. We assess robustness properties of our procedure and apply it to real data from the field of functional magnetic resonance imaging.

stat.ME

Permutation-Based True Discovery Guarantee by Sum Tests

Sum-based global tests are highly popular in multiple hypothesis testing. In this paper we propose a general closed testing procedure for sum tests, which provides lower confidence bounds for the proportion of true discoveries (TDP), simultaneously over all subsets of hypotheses. These simultaneous inferences come for free, i.e., without any adjustment of the alpha-level, whenever a global test is used. Our method allows for an exploratory approach, as simultaneity ensures control of the TDP even when the subset of interest is selected post hoc. It adapts to the unknown joint distribution of the data through permutation testing. Any sum test may be employed, depending on the desired power properties. We present an iterative shortcut for the closed testing procedure, based on the branch and bound algorithm, which converges to the full closed testing results, often after few iterations; even if it is stopped early, it controls the TDP. We compare the properties of different choices for the sum test through simulations, then we illustrate the feasibility of the method for high dimensional data on brain imaging and genomics data.

stat.ME

Procrustes-based distances for exploring between-matrices similarity

The statistical shape analysis called Procrustes analysis minimizes the distance between matrices by similarity transformations. The method returns a set of optimal orthogonal matrices, which project each matrix into a common space. This manuscript presents two types of distances derived from Procrustes analysis for exploring between-matrices similarity. The first one focuses on the residuals from the Procrustes analysis, i.e., the residual-based distance metric. In contrast, the second one exploits the fitted orthogonal matrices, i.e., the rotational-based distance metric. Thanks to these distances, similarity-based techniques such as the multidimensional scaling method can be applied to visualize and explore patterns and similarities among observations. The proposed distances result in being helpful in functional magnetic resonance imaging (fMRI) data analysis. The brain activation measured over space and time can be represented by a matrix. The proposed distances applied to a sample of subjects -- i.e., matrices -- revealed groups of individuals sharing patterns of neural brain activation.

stat.AP

Post-selection Inference in Multiverse Analysis (PIMA): an inferential framework based on the sign flipping score test

When analyzing data researchers make some decisions that are either arbitrary, based on subjective beliefs about the data generating process, or for which equally justifiable alternative choices could have been made. This wide range of data-analytic choices can be abused, and has been one of the underlying causes of the replication crisis in several fields. Recently, the introduction of multiverse analysis provides researchers with a method to evaluate the stability of the results across reasonable choices that could be made when analyzing data. Multiverse analysis is confined to a descriptive role, lacking a proper and comprehensive inferential procedure. Recently, specification curve analysis adds an inferential procedure to multiverse analysis, but this approach is limited to simple cases related to the linear model, and only allows researchers to infer whether at least one specification rejects the null hypothesis, but not which specifications should be selected. In this paper we present a Post-selection Inference approach to Multiverse Analysis (PIMA) which is a flexible and general inferential approach that accounts for all possible models, i.e., the multiverse of reasonable analyses. The approach allows for a wide range of data specifications (i.e. pre-processing) and any generalized linear model; it allows testing the null hypothesis of a given predictor not being associated with the outcome, by merging information from all reasonable models of multiverse analysis, and provides strong control of the family-wise error rate such that it allows researchers to claim that the null-hypothesis can be rejected for each specification that shows a significant effect. The inferential proposal is based on a conditional resampling procedure. To be continued...

stat.ME

Resampling-Based Multisplit Inference for High-Dimensional Regression

We propose a novel resampling-based method to construct an asymptotically exact test for any subset of hypotheses on coefficients in high-dimensional linear regression. It can be embedded into any multiple testing procedure to make confidence statements on relevant predictor variables. The method constructs permutation test statistics for any individual hypothesis by means of repeated splits of the data and a variable selection technique; then it defines a test for any subset by suitably aggregating its variables' test statistics. The resulting procedure is extremely flexible, as it allows different selection techniques and several combining functions. We present it in two ways: an exact method and an approximate one, that requires less memory usage and shorter computation time, and can be scaled up to higher dimensions. We illustrate the performance of the method with simulations and the analysis of real gene expression data.

stat.ME