SearcharxivSearch

arXiv subjects

Thorsten Dickhaus

Publications and source records attributed to Thorsten Dickhaus.

At least 19 recordsLinked to original sources

On E-Backtesting: Generalizations and Sample Size Determination

We present an approach for determining sample sizes required to detect underestimations of the expected shortfall with a prescribed power when applying the recently proposed e-backtesting procedure. We consider scenarios in which the value-at-risk at level $p$ is always estimated correctly, while the difference between the true expected shortfall and the value-at-risk is underestimated by a given factor $r$. We show that exploiting the structure of the backtest e-statistic proposed for backtesting the expected shortfall at level $p$ enables the derivation of approximate lower bounds for the required sample sizes by considering a sequence of independent and identically distributed Bernoulli random variables. We also discuss potential limitations of this approximation and compare the resulting sample size requirements with those obtained in practical applications using Monte Carlo simulations. Furthermore, we present generalizations of the e-backtesting procedure, in particular to risk measures which constitute Bayes pairs.

stat.ME

Estimation of the complexity of a network under a Gaussian graphical model

The proportion of edges in a Gaussian graphical model (GGM) characterizes the complexity of its conditional dependence structure. Since edge presence corresponds to a nonzero entry of the precision matrix, estimation of this proportion can be formulated as a large-scale multiple testing problem. We propose an estimator that combines p-values from simultaneous edge-wise tests, conducted under false discovery rate control, with Storey's estimator of the proportion of true null hypotheses. We establish weak dependence conditions on the precision matrix under which the empirical cumulative distribution function of the p-values converges to its population counterpart. These conditions cover high-dimensional regimes, including those arising in genetic association studies. Under such dependence, we characterize the asymptotic bias of the Schweder--Spj{\o}tvoll estimator, showing that it is upward biased and thus slightly underestimates the true edge proportion. Simulation studies across a variety of models confirm accurate recovery of graph complexity.

stat.ME

Partial Conjunction Analysis in Neuroimaging: A Comparative Study

Replicability is a cornerstone of science. The partial conjunction (PC) hypothesis testing framework objectively quantifies replicability across disciplines. Although several statistical methodologies for testing PC hypotheses exist, it is not clear which method performs well under which circumstances. In this paper, we consider the PC hypothesis testing problem from a neuroimaging perspective. Identifying the brain regions activated by a specific cognitive task constitutes a central challenge in neuroimaging. This problem becomes complex when the objective is to evaluate whether activation patterns are consistent across different cognitive tasks or subjects. In this paper, we cast this question as a PC hypothesis testing problem, assessing, for each location in the brain, whether it is activated in at least $\gamma$ subjects, for a pre-specified granularity $\gamma$. In our comparative study, we consider three methods, namely: adaFilter, CoFilter, and a method proposed by Benjamini, Heller, and Yekutieli (BHY). In equi-correlated simulated data, the BHY procedure tends to outperform the competing methods for high values of $\gamma$, while CoFilter performs well for low values of $\gamma$. In the real-data analysis, CoFilter dominates the other methods for intermediate values of $\gamma$.

stat.ME

Identifying rapid changes in the hemodynamic response in event-related functional magnetic resonance imaging

The hemodynamic response (HR) in event-related functional magnetic resonance imaging is typically assumed to be stationary. While there are some approaches in the literature to model nonstationary HRs, few focus on rapid changes. In this work, we propose two procedures to investigate rapid changes in the HR. Both procedures make inference on the existence of rapid changes for multi-subject data. We allow the change point locations to vary between subjects, conditions and brain regions. The first procedure utilizes available information about the change point locations to compare multiple shape parameters of the HR over time. In the second procedure, the change point locations are determined for each subject separately. To account for the estimation of the change point locations, we propose the notion of post selection variance. The power of the proposed procedures is assessed in simulation studies. We apply the procedure for pre-specified change point locations to data from a category learning experiment.

stat.ME

Utilizing Multiple Testing for Grouping in Singular Spectrum Analysis

A key step in separating signal from noise in time series by means of singular spectrum analysis (SSA) is grouping. We present a multiple testing method for the grouping step in SSA. As separability criterion, we utilize the weighted correlation between the signal and the noise component of the (reconstructed) time series, and we test whether this weighted correlation is equal to zero. This test has to be performed for several possible groupings, resulting in a multiple test problem. The null distributions of the corresponding test statistics are approximated by a wild bootstrap procedure. The performance of our proposed method is assessed in a simulation study, and we illustrate its practical application with an analysis of real world data.

stat.ME

Confidence bounds for the true discovery proportion based on the exact distribution of the number of rejections

In multiple hypotheses testing it has become widely popular to make inference on the true discovery proportion (TDP) of a set $\mathcal{M}$ of null hypotheses. This approach is useful for several application fields, such as neuroimaging and genomics. Several procedures to compute simultaneous lower confidence bounds for the TDP have been suggested in prior literature. Simultaneity allows for post-hoc selection of $\mathcal{M}$. If sets of interest are specified a priori, it is possible to gain power by removing the simultaneity requirement. We present an approach to compute lower confidence bounds for the TDP if the set of null hypotheses is defined a priori. The proposed method determines the bounds using the exact distribution of the number of rejections based on a step-up multiple testing procedure under independence assumptions. We assess robustness properties of our procedure and apply it to real data from the field of functional magnetic resonance imaging.

stat.ME

Integrating Transformations in Probabilistic Circuits

This study addresses the predictive limitation of probabilistic circuits and introduces transformations as a remedy to overcome it. We demonstrate this limitation in robotic scenarios. We motivate that independent component analysis is a sound tool to preserve the independence properties of probabilistic circuits. Our approach is an extension of joint probability trees, which are model-free deterministic circuits. By doing so, it is demonstrated that the proposed approach is able to achieve higher likelihoods while using fewer parameters compared to the joint probability trees on seven benchmark data sets as well as on real robot data. Furthermore, we discuss how to integrate transformations into tree-based learning routines. Finally, we argue that exact inference with transformed quantile parameterized distributions is not tractable. However, our approach allows for efficient sampling and approximate inference.

stat.ML

Regionalization approaches for the spatial analysis of extremal dependence

The impact of an extreme climate event depends strongly on its geographical scale. Max-stable processes can be used for the statistical investigation of climate extremes and their spatial dependencies on a continuous area. Most existing parametric models of max-stable processes assume spatial stationarity and are therefore not suitable for the application to data that cover a large and heterogeneous area. For this reason, it has recently been proposed to use a clustering algorithm to divide the area of investigation into smaller regions and to fit parametric max-stable processes to the data within those regions. We investigate this clustering algorithm further and point out that there are cases in which it results in regions on which spatial stationarity is not a reasonable assumption. We propose an alternative clustering algorithm and demonstrate in a simulation study that it can lead to improved results.

stat.ME

On the closed-loop Volterra method for analyzing time series

The main focus of this paper is to approximate time series data based on the closed-loop Volterra series representation. Volterra series expansions are a valuable tool for representing, analyzing, and synthesizing nonlinear dynamical systems. However, a major limitation of this approach is that as the order of the expansion increases, the number of terms that need to be estimated grows exponentially, posing a considerable challenge. This paper considers a practical solution for estimating the closed-loop Volterra series in stationary nonlinear time series using the concepts of Reproducing Kernel Hilbert Spaces (RKHS) and polynomial kernels. We illustrate the applicability of the suggested Volterra representation by means of simulations and real data analysis. Furthermore, we apply the Kolmogorov-Smirnov Predictive Accuracy (KSPA) test, to determine whether there exists a statistically significant difference between the distribution of estimated errors for concurring time series models, and secondly to determine whether the estimated time series with the lower error based on some loss function also has exhibits a stochastically smaller error than estimated time series from a competing method. The obtained results indicate that the closed-loop Volterra method can outperform the ARFIMA, ETS, and Ridge regression methods in terms of both smaller error and increased interpretability.

stat.ME

Multiple testing of composite null hypotheses for discrete data using randomized $p$-values

$P$-values that are derived from continuously distributed test statistics are typically uniformly distributed on $(0,1)$ under least favorable parameter configurations (LFCs) in the null hypothesis. Conservativeness of a $p$-value $P$ (meaning that $P$ is under the null hypothesis stochastically larger than a random variable which is uniformly distributed on $(0,1)$) can occur if the test statistic from which $P$ is derived is discrete, or if the true parameter value under the null is not an LFC. To deal with both of these sources of conservativeness, we present two approaches utilizing randomized $p$-values, namely single-stage and two-stage randomization. We illustrate their effectiveness for testing a composite null hypothesis under a binomial model. We also give an example of how the proposed $p$-values can be used to test a composite null in group testing designs. Similar to previous findings, we find that the proposed randomized $p$-values are less conservative compared to non-randomized $p$-values under the null hypothesis, but that they are stochastically not smaller under the alternative. The problem of establishing the validity of randomized $p$-values is not trivial and has received attention in previous literature. We show that our proposed randomized $p$-values are valid under various discrete statistical models which are such that the distribution of the corresponding test statistic belongs to an exponential family. The behaviour of the power function for the tests based on the proposed randomized $p$-values as a function of the sample size is also investigated. Simulations and a real data analysis are used to compare the different considered $p$-values.

stat.ME

Long-term temporal evolution of extreme temperature in a warming Earth

We present a new approach to modeling the future development of extreme temperatures globally and on a long time-scale by using non-stationary generalized extreme value distributions in combination with logistic functions. This approach is applied to data from the fully coupled climate model AWI-ESM. It enables us to investigate how extremes will change depending on the geographic location not only in terms of the magnitude, but also in terms of the timing of the changes. We observe that in general, changes in extremes are stronger and more rapid over land masses than over oceans. In addition, our models differentiate between changes in mean, in variability and in distributional shape, allowing for developments in these statistics to take place independently and at different times. Different models are presented and the Bayesian Information Criterion is used for model selection. It turns out that in most regions, changes in mean and variance take place simultaneously while the shape parameter of the distribution is predicted to stay constant. In the Arctic region, however, a different picture emerges: There, climate variability drastically and abruptly increases around 2050 due to the melting of ice, whereas changes in the mean values take longer and come into effect later.

physics.ao-ph

Combining Multiple Testing with Multivariate Singular Spectrum Analysis

Appropriate preprocessing is a fundamental prerequisite for analyzing a noisy dataset. The purpose of this paper is to apply a nonparametric preprocessing method, called Singular Spectrum Analysis (SSA), to a variety of datasets which are subsequently analyzed by means of multiple statistical hypothesis tests. SSA is a nonparametric preprocessing method which has recently been utilized in the context of many life science problems. In the present work, SSA is compared with three other state-of-the-art preprocessing methods in terms of goodness of denoising and in terms of the statistical power of the subsequent multiple test. These other methods are either parametric or nonparametric. Our findings demonstrate that (multivariate) SSA can be taken into account as a promising method to reduce noise, to extract the main signal from noisy data, and to detect statistically significant signal components.

stat.ME

Multiple multi-sample testing under arbitrary covariance dependency

Modern high-throughput biomedical devices routinely produce data on a large scale, and the analysis of high-dimensional datasets has become commonplace in biomedical studies. However, given thousands or tens of thousands of measured variables in these datasets, extracting meaningful features poses a challenge. In this article, we propose a procedure to evaluate the strength of the associations between a nominal (categorical) response variable and multiple features simultaneously. Specifically, we propose a framework of large-scale multiple testing under arbitrary correlation dependency among test statistics. First, marginal multinomial regressions are performed for each feature individually. Second, we use an approach of multiple marginal models for each baseline-category pair to establish asymptotic joint normality of the stacked vector of the marginal multinomial regression coefficients. Third, we estimate the (limiting) covariance matrix between the estimated coefficients from all marginal models. Finally, our approach approximates the realized false discovery proportion of a thresholding procedure for the marginal p-values, for each baseline-category pair. The proposed approach offers a sensible trade-off between the expected numbers of true and false rejections. Furthermore, we demonstrate a practical application of the method on hyperspectral imaging data. This dataset is obtained by a matrix-assisted laser desorption/ionization (MALDI) instrument. MALDI demonstrates tremendous potential for clinical diagnosis, particularly for cancer research. In our application, the nominal response categories represent cancer subtypes.

stat.ME

A procedure for multiple testing of partial conjunction hypotheses based on a hazard rate inequality

The partial conjunction null hypothesis is tested in order to discover a signal that is present in multiple studies. The standard approach of carrying out a multiple test procedure on the partial conjunction (PC) $p$-values can be extremely conservative. We suggest alleviating this conservativeness, by eliminating many of the conservative PC $p$-values prior to the application of a multiple test procedure. This leads to the following two step procedure: first, select the set with PC $p$-values below a selection threshold; second, within the selected set only, apply a family-wise error rate or false discovery rate controlling procedure on the conditional PC $p$-values. The conditional PC $p$-values are valid if the null p-values are uniform and the combining method is Fisher. The proof of their validity is based on a novel inequality in hazard rate order of partial sums of order statistics which may be of independent interest. We also provide the conditions for which the false discovery rate controlling procedures considered will be below the nominal level. We demonstrate the potential usefulness of our novel method, CoFilter (conditional testing after filtering), for analyzing multiple genome wide association studies of Crohn's disease.

stat.ME

Multiple two-sample testing under arbitrary covariance dependency with an application in imaging mass spectrometry

Large-scale hypothesis testing has become a ubiquitous problem in high-dimensional statistical inference, with broad applications in various scienfitic disciplines. One relevant application is constituted by imaging mass spectrometry (IMS) association studies, where a large number of tests are performed simultaneously in order to identify molecular masses that are associated with a particular phenotype, e. g., a cancer subtype. Mass spectra obtained from Matrix-assisted laser desorption/ionization (MALDI) experiments are dependent, when considered as statistical quantities. False discovery proportion (FDP) control under arbitrary dependency structure among test statistics is an active topic in modern multiple testing research. In this context, we are concerned with the evaluation of associations between the binary outcome variable (describing the phenotype) and multiple predictors derived from MALDI measurements. We propose an inference procedure in which the correlation matrix of the test statistics is utilized. The approach is based on multiple marginal models (MMM). Specifically, we fit a marginal logistic regression model for each predictor individually. Asymptotic joint normality of the stacked vector of the marginal regression coefficients is established under standard regularity assumptions, and their (limiting) correlation matrix is estimated. The proposed method extracts common factors from the resulting empirical correlation matrix. Finally, we estimate the realized FDP of a thresholding procedure for the marginal $p$-values. We demonstrate a practical application of the proposed workflow to MALDI IMS data in an oncological context.

stat.ME

Standard Curves for Empirical Likelihood Ratio Tests of Means

We present simulated standard curves for the calibration of empirical likelihood ratio (ELR) tests of means. With the help of these curves, the nominal significance level of the ELR test can be adjusted in order to achieve (quasi-) exact type I error rate control for a given, finite sample size. By theoretical considerations and by computer simulations, we demonstrate that the adjusted significance level depends most crucially on the skewness and on the kurtosis of the parent distribution. For practical purposes, we tabulate adjusted critical values under several prototypical statistical models.

stat.ME

Combining independent p-values in replicability analysis: A comparative study

Given a family of null hypotheses $H_{1},\ldots,H_{s}$, we are interested in the hypothesis $H_{s}^γ$ that at most $γ-1$ of these null hypotheses are false. Assuming that the corresponding $p$-values are independent, we are investigating combined $p$-values that are valid for testing $H_{s}^γ$. In various settings in which $H_{s}^γ$ is false, we determine which combined $p$-value works well in which setting. Via simulations, we find that the Stouffer method works well if the null $p$-values are uniformly distributed and the signal strength is low, and the Fisher method works better if the null $p$-values are conservative, i.e. stochastically larger than the uniform distribution. The minimum method works well if the evidence for the rejection of $H_{s}^γ$ is focused on only a few non-null $p$-values, especially if the null $p$-values are conservative. Methods that incorporate the combination of $e$-values work well if the null hypotheses $H_{1},\ldots,H_{s}$ are simple.

stat.AP