SearcharxivSearch

arXiv subjects

Dominic Edelmann

Publications and source records attributed to Dominic Edelmann.

13 recordsLinked to original sources

Approximating the null distribution of generalized distance covariance

The null distribution of distance covariance is usually approximated by permutation, which is prohibitive when very small p-values are needed, or by matching a few moments to a parametric family, which is inaccurate in the tails. A third option is to approximate the limiting distribution, a weighted sum of chi-square variables, directly through the spectra of the doubly centred distance matrices. This is used for kernel-based tests but has lacked a rigorous justification. We prove that the empirical spectra give a uniformly consistent approximation of the limiting null distribution, and hence an asymptotically valid test, for a general class of distances of negative type on separable metric spaces. The result covers the Hilbert-Schmidt independence criterion as a special case. We also give an adaptive algorithm that brackets the p-value from a partial eigendecomposition, reducing the cost from $O(n^3)$ to $O(k n^2)$, and a shrinkage correction matching the first two moments. In simulations, the proposed tests are the only non-Monte-Carlo procedures whose empirical type I error converges to the nominal level.

stat.ME

Omnibus Goodness-of-Fit Testing for Distributions on Stiefel Manifolds

In this article, a comprehensive framework for goodness-of-fit testing for distributions on Stiefel manifolds is developed. The approach is based on integrals of the squared differences between empirical and theoretical characteristic functions, yielding test statistics that are consistent against all fixed alternatives. For the Fisher-Bingham family of distributions, explicit computable forms of the test statistic are derived. Simplified expressions for important special cases, including the matrix Fisher, matrix Bingham, and uniform distributions are provided. In the case of testing uniformity on hyperspheres, we obtain the complete asymptotic distribution of the test statistic, enabling computationally efficient asymptotic testing. For general Fisher-Bingham distributions, we establish theoretically justified Monte Carlo testing procedures for both simple and composite hypotheses. Simulation studies demonstrate accurate Type I error control and strong power across a wide range of alternatives. The practical relevance of the proposed methodology is illustrated by an application to data on the orbits of comets.

math.ST

Effect measures for comparing paired event times

The progression-free survival ratio (PFSr) is a widely used measure in personalized oncology trials. It evaluates the effectiveness of treatment by comparing two consecutive event times - one under standard therapy and one under an experimental treatment. However, most proposed tests based on the PFSr cannot control the nominal type I error rate, even under mild assumptions such as random right-censoring. Consequently the results of these tests are often unreliable. As a remedy, we propose to estimate the relevant probabilities related to the PFSr by adapting recently developed methodology for the relative treatment effect between paired event times. As an additional alternative, we develop inference procedures based on differences and ratios of restricted mean survival times. An extensive simulation study confirms that the proposed novel methodology provides reliable inference, whereas previously proposed techniques break down in many realistic settings. The utility of our methods is further illustrated through an analysis of real data from a molecularly aided tumor trial.

stat.ME

A generalized distance covariance framework for genome-wide association studies

When testing for the association of a single SNP with a phenotypic response, one usually considers an additive genetic model, assuming that the mean of of the response for the heterozygous state is the average of the means for the two homozygous states. However, this simplification often does not hold. In this paper, we present a novel framework for testing the association of a single SNP and a phenotype. Different from the predominant standard approach, our methodology is guaranteed to detect all dependencies expressed by classical genetic association models. The asymptotic distribution under mild regularity assumptions is derived. Moreover, the finite sample distribution under Gaussianity is provided in which the exact p-value can be efficiently evaluated via the classical Appell hypergeometric series. Both results are extended to a regression-type setting with nuisance covariates, enabling hypotheses testing in a wide range of scenarios. A connection of our approach to score tests is explored, leading to intuitive interpretations as locally most powerful tests. A simulation study demonstrates the computational efficiency and excellent statistical performance of the proposed methodology. A real data example is provided.

stat.ME

Tests for categorical data beyond Pearson: A distance covariance and energy distance approach

Categorical variables are of uttermost importance in biomedical research. When two of them are considered, it is often the case that one wants to test whether or not they are statistically dependent. We show weaknesses of classical methods -- such as Pearson's and the G-test -- and we propose testing strategies based on distances that lack those drawbacks. We first develop this theory for classical two-dimensional contingency tables, within the context of distance covariance, an association measure that characterizes general statistical independence of two variables. We then apply the same fundamental ideas to one-dimensional tables, namely to the testing for goodness of fit to a discrete distribution, for which we resort to an analogous statistic called energy distance. We prove that our methodology has desirable theoretical properties, and we show that we can calibrate the null distribution of our test statistics without resampling. We illustrate all this in simulations, as well as with some real data examples, demonstrating the adequate performance of our approach for biostatistical practice.

stat.ME

A Basic Treatment of the Distance Covariance

The distance covariance of Sz\'ekely, et al. [23] and Sz\'ekely and Rizzo [21], a powerful measure of dependence between sets of multivariate random variables, has the crucial feature that it equals zero if and only if the sets are mutually independent. Hence the distance covariance can be applied to multivariate data to detect arbitrary types of non-linear associations between sets of variables. We provide in this article a basic, albeit rigorous, introductory treatment of the distance covariance. Our investigations yield an approach that can be used as the foundation for presentation of this important and timely topic even in advanced undergraduate- or junior graduate-level courses on mathematical statistics.

math.ST

Product Inequalities for Multivariate Gaussian, Gamma, and Positively Upper Orthant Dependent Distributions

The Gaussian product inequality is an important conjecture concerning the moments of Gaussian random vectors. While all attempts to prove the Gaussian product inequality in full generality have been unsuccessful to date, numerous partial results have been derived in recent decades and we provide here further results on the problem. Most importantly, we establish a strong version of the Gaussian product inequality for multivariate gamma distributions in the case of nonnegative correlations, thereby extending a result recently derived by Genest and Ouimet [5]. Further, we show that the Gaussian product inequality holds with nonnegative exponents for all random vectors with positive components whenever the underlying vector is positively upper orthant dependent. Finally, we show that the Gaussian product inequality with negative exponents follows directly from the Gaussian correlation inequality.

math.PR

Testing for genetic interactions in complex disease with distance correlation

Understanding epistasis (genetic interaction) may shed some light on the genomic basis of common diseases, including disorders of maximum interest due to their high socioeconomic burden, like schizophrenia. Distance correlation is an association measure that characterises general statistical independence between random variables, not only the linear one. Here, we propose distance correlation as a novel tool for the detection of epistasis from case-control data of single-nucleotide polymorphisms (SNPs). On the methodological side, we highlight the derivation of the explicit asymptotic distribution of the test statistic. We show that this is the only way to obtain enough computational speed for the method to be used in practice, in a scenario where the resampling techniques found in the literature are impractical. Our simulations show satisfactory calibration of significance, as well as comparable or better power than existing methodology. We conclude with the application of our technique to a schizophrenia genetics dataset, obtaining biologically sound insights.

math.ST

An Updated Literature Review of Distance Correlation and its Applications to Time Series

The concept of distance covariance/correlation was introduced recently to characterize dependence among vectors of random variables. We review some statistical aspects of distance covariance/correlation function and we demonstrate its applicability to time series analysis. We will see that the auto-distance covariance/correlation function is able to identify nonlinear relationships and can be employed for testing the i.i.d.\ hypothesis. Comparisons with other measures of dependence are included.

stat.ME

The Distance Standard Deviation

The distance standard deviation, which arises in distance correlation analysis of multivariate data, is studied as a measure of spread. The asymptotic distribution of the empirical distance standard deviation is derived under the assumption of finite second moments. Applications are provided to hypothesis testing on a data set from materials science and to multivariate statistical quality control. The distance standard deviation is compared to classical scale measures for inference on the spread of heavy-tailed distributions. Inequalities for the distance variance are derived, proving that the distance standard deviation is bounded above by the classical standard deviation and by Gini's mean difference. New expressions for the distance standard deviation are obtained in terms of Gini's mean difference and the moments of spacings of order statistics. It is also shown that the distance standard deviation satisfies the axiomatic properties of a measure of spread.

math.ST

Distance Correlation Coefficients for Lancaster Distributions

We consider the problem of calculating distance correlation coefficients between random vectors whose joint distributions belong to the class of Lancaster distributions. We derive under mild convergence conditions a general series representation for the distance covariance for these distributions. To illustrate the general theory, we apply the series representation to derive explicit expressions for the distance covariance and distance correlation coefficients for the bivariate normal distribution and its generalizations of Lancaster type, the multivariate normal distributions, and the bivariate gamma, Poisson, and negative binomial distributions which are of Lancaster type.

math.ST

A Generalization of an Integral Arising in the Theory of Distance Correlation

We generalize an integral which arises in several areas in probability and statistics and which is at the core of the field of distance correlation, a concept developed by Sz\'ekely, Rizzo and Bakirov (2007) to measure dependence between random variables. Let $m$ be a positive integer and let ${\cos_m}(u)$, $u \in \mathbb{R}$, be the truncated Maclaurin expansion of ${\cos}(u)$, where the expansion is truncated at the $m$th summand. For $t, x \in \mathbb{R}^d$, let $\langle t,x\rangle$ and $\|x\|$ denote the standard Euclidean inner product and norm, respectively. We establish the integral formula: For $\alpha \in \mathbb{C}$ and $x \in \mathbb{R}^d$, $\int_{{\mathbb{R}}^d} [\cos_m(\langle t,x\rangle) - \cos(\langle t,x\rangle)] \,{\rm d}t/{\|t\|^{d+\alpha}} = C(d,\alpha) \, \|x\|^{\alpha}$, with absolute convergence if and only if $2(m-1) < \Re(\alpha) < 2m$. Moreover, the constant $C(d,\alpha)$ does not depend on $m$.

math.ST

The affinely invariant distance correlation

Sz\'{e}kely, Rizzo and Bakirov (Ann. Statist. 35 (2007) 2769-2794) and Sz\'{e}kely and Rizzo (Ann. Appl. Statist. 3 (2009) 1236-1265), in two seminal papers, introduced the powerful concept of distance correlation as a measure of dependence between sets of random variables. We study in this paper an affinely invariant version of the distance correlation and an empirical version of that distance correlation, and we establish the consistency of the empirical quantity. In the case of subvectors of a multivariate normally distributed random vector, we provide exact expressions for the affinely invariant distance correlation in both finite-dimensional and asymptotic settings, and in the finite-dimensional case we find that the affinely invariant distance correlation is a function of the canonical correlation coefficients. To illustrate our results, we consider time series of wind vectors at the Stateline wind energy center in Oregon and Washington, and we derive the empirical auto and cross distance correlation functions between wind vectors at distinct meteorological stations.

math.ST