SearcharxivSearch

arXiv subjects

Johannes Heiny

Publications and source records attributed to Johannes Heiny.

At least 19 recordsLinked to original sources

Quadratic form of heavy-tailed self-normalized random vector with applications in $α$-heavy Marčenko--Pastur law

Let $\mathbf{x}$ be a random vector with $n$ i.i.d.\ real-valued components in the domain attraction of an $α$-stable law with $α\in(0,2)$, and let $\mathbf{y}=\mathbf{x}/\|\mathbf{x}\|_2$ be the associated self-normalized vector on the unit sphere. For a (possibly random) Hermitian matrix $\mathbf{A}_n=\big(a_{ij}^{(n)}\big)$ independent of $\mathbf{y}$, we study the asymptotic law of the quadratic form $\mathbf{y}^\top \mathbf{A}_n \mathbf{y}$. Building on the sharp separation between diagonal and off-diagonal contributions in this heavy-tailed setting, we show that under a mild assumption on the Frobenius norm of the off-diagonal part of $\mathbf{A}_n$ the limiting law is solely governed by the empirical distribution of the diagonal entries and the index $α$. More precisely, if $n^{-1}\sum_{i=1}^n δ_{a^{(n)}_{ii}}$ converges weakly almost surely to a deterministic $ν$, then $Q_n$ converges in distribution to a non-degenerate law $μ_{ν,α}$ characterized through its Stieltjes transform. The law $μ_{ν,α}$ is shown to be atom-free (provided that $ν$ is non-degenerate) with an explicit density and tractable tail behavior. As an application in random matrix theory, we derive an implicit resolvent-based representation of the $α$-heavy Marčenko--Pastur law $H_{α,γ}$ for heavy-tailed sample correlation matrices and prove that $H_{α,γ}$ has no atoms except possibly at the origin. For comparison with the light-tailed setting, we also provide a Hanson--Wright-type concentration inequality for $\mathbf{y}^\top \mathbf{A}_n \mathbf{y}$ when the components of $\mathbf{x}$ are sub-Gaussian.

math.PR

Characteristic polynomial of self-normalized random matrices

We study the characteristic polynomial of self-normalized random matrices, whose rows are independent and normalized to have unit $\mathrm{L}^2$ norm. The entries before normalization are assumed to have regularly varying tails with tail index $α\in [0,2]$. We prove that, outside the unit disk, the characteristic polynomial converges to a random analytic function $F_α$. We identify $F_α$ as a multiplicative chaos described in terms of Poisson point processes. The family of limiting functions $(F_α)_{α\in[0,2]}$ interpolates between two universal regimes: Poisson multiplicative chaos at $α=0$ and Gaussian multiplicative chaos at the boundary $α=2$. Thus, self-normalization provides a matrix model where one observes the transition between Poissonian and Gaussian regimes for the limiting characteristic polynomial as the tail index varies. A similar transition is found for the fluctuations of the traces of self-normalized matrices. As an application of our results, we derive that the spectral radius of self-normalized matrices is asymptotically bounded above by one in probability for any symmetric entry distribution.

math.PR

Ties, Tails and Spectra: On Rank-Based Dependency Measures in High Dimensions

This work is concerned with the limiting spectral distribution of rank-based dependency measures in high dimensions. We provide distribution-free results for multivariate empirical versions of Kendall's $τ$ and Spearman's $ρ$ in a setting where the dimension $p$ grows at most proportionally to the sample size $n$. Although rank-based measures are known to be well suited for discrete and heavy-tailed data, previous works in the field focused mostly on the continuous and light-tailed case. We close this gap by imposing mild assumptions and allowing for general types of distributions. Interestingly, our analysis reveals that a non-trivial adjustment of classical Kendall's $τ$ is needed to obtain a pivotal limiting distribution in the presence of tied data. The proof for Spearman's $ρ$ is facilitated by a result regarding the limiting eigenvalue distribution of a general class of random matrices with rows on the Euclidean unit sphere, which is of independent interest. For instance, this finding can be used to derive the limiting spectral distribution of sample correlation matrices, which, in contrast to most existing works, accommodates heavy-tailed data.

math.ST

Non-Gaussian fluctuations for traces of squared sample correlation matrices in high dimensions

We provide limit theory for the trace of the squared sample correlation matrix $\mathbf R$, constructed from $n$ observations of a $p$-dimensional random vector with iid components. If the entries have finite fourth moment and $p$ and $n$ grow proportionally, it is known that $\operatorname{tr}({\mathbf R}^2)$ satisfies a central limit theorem (CLT) and the centering and scaling sequences are universal in the sense that they do not depend on the entry distribution. Under symmetry and regular variation assumption with index $α$ and any growth rate of the dimension, we prove that the universal CLT remains valid for $α>3$. For $α<3$, we identify a critical dimension growth at which the fluctuations of $\operatorname{tr}({\mathbf R}^2)$ become non-Gaussian. Moreover, if the dimension $p$ grows faster and $α\le 3$ we establish a non-universal CLT with norming sequences depending on the value of $α$. Our findings are illustrated in a simulation study.

math.ST

Maximum of sparsely equicorrelated Gaussian fields and applications

We investigate the extreme values of a sparse and equicorrelated Gaussian field on a triangle: the correlations on every vertical or horizontal line are all equal to a parameter $r \in [0,1/2]$ and are zero everywhere else. This problem is closely linked with various problems in high-dimensional statistics and extreme-value theory. We identify the threshold for $r$ at which the standard Gumbel law breaks down. Our result is based on a subtle application of the Chen-Stein method for Poisson approximation. As applications, we discuss the implication of our results on multiple testing and resolve several questions that were left open in \cite{heiny2024maximum}, \cite{tang2022asymptotic} and \cite{Jiang19}.

math.PR

Joint extremes of inversions and descents of random permutations

We provide asymptotic theory for the joint distribution of $X_{\mathrm{inv}}$ and $X_{\mathrm{des}}$, the numbers of inversions and descents of random permutations. Recently, Dörr & Kahle (2022) proved that $X_{\mathrm{inv}}$, respectively, $X_{\mathrm{des}}$ is in the maximum domain of attraction of the Gumbel distribution. To tackle the dependency between these two permutation statistics, we use Hájek projections and a suitable quantitative Gaussian approximation. We show that $(X_{\mathrm{inv}}, X_{\mathrm{des}})$ is in the maximum domain of attraction of the two-dimensional Gumbel distribution with independent margins. This result can be stated in the broader combinatorial framework of finite Coxeter groups, on which our method also yields the central limit theorem for $(X_{\mathrm{inv}}, X_{\mathrm{des}})$ and various other permutation statistics as a novel contribution. In particular, signed permutation groups with random biased signs and products of classical Weyl groups are investigated.

math.PR

Testing for practically significant dependencies in high dimensions via bootstrapping maxima of U-statistics

This paper takes a different look on the problem of testing the mutual independence of the components of a high-dimensional vector. Instead of testing if all pairwise associations (e.g. all pairwise Kendall's $τ$) between the components vanish, we are interested in the (null)-hypothesis that all pairwise associations do not exceed a certain threshold in absolute value. The consideration of these hypotheses is motivated by the observation that in the high-dimensional regime, it is rare, and perhaps impossible, to have a null hypothesis that can be exactly modeled by assuming that all pairwise associations are precisely equal to zero. The formulation of the null hypothesis as a composite hypothesis makes the problem of constructing tests non-standard and in this paper we provide a solution for a broad class of dependence measures, which can be estimated by $U$-statistics. In particular we develop an asymptotic and a bootstrap level $α$-test for the new hypotheses in the high-dimensional regime. We also prove that the new tests are minimax-optimal and investigate their finite sample properties by means of a small simulation study and a data example.

math.ST

Maximum interpoint distance of high-dimensional random vectors

A limit theorem for the largest interpoint distance of $p$ independent and identically distributed points in $\mathbb{R}^n$ to the Gumbel distribution is proved, where the number of points $p=p_n$ tends to infinity as the dimension of the points $n\to\infty$. The theorem holds under moment assumptions and corresponding conditions on the growth rate of $p$. We obtain a plethora of ancillary results such as the joint convergence of maximum and minimum interpoint distances. Using the inherent sum structure of interpoint distances, our result is generalized to maxima of dependent random walks with non-decaying correlations and we also derive point process convergence. An application of the maximum interpoint distance to testing the equality of means for high-dimensional random vectors is presented. Moreover, we study the largest off-diagonal entry of a sample covariance matrix. The proofs are based on the Chen-Stein Poisson approximation method and Gaussian approximation to large deviation probabilities.

math.PR

Point process convergence for symmetric functions of high-dimensional random vectors

The convergence of a sequence of point processes with dependent points, defined by a symmetric function of iid high-dimensional random vectors, to a Poisson random measure is proved. This also implies the convergence of the joint distribution of a fixed number of upper order statistics. As applications of the result a generalization of maximum convergence to point process convergence is given for simple linear rank statistics, rank-type U-statistics and the entries of sample covariance matrices.

math.PR

The volume of random simplices from elliptical distributions in high dimension

Random simplices and more general random convex bodies of dimension $p$ in $\mathbb{R}^n$ with $p\leq n$ are considered, which are generated by random vectors having an elliptical distribution. In the high-dimensional regime, that is, if $p\to\infty$ and $n\to\infty$ in such a way that $p/n\toγ\in(0,1)$, a central and a stable limit theorem for the logarithmic volume of random simplices and random convex bodies is shown. The result follows from a related central limit theorem for the log-determinant of $p\times n$ random matrices whose rows are copies of a random vector with an elliptical distribution, which is established as well.

math.PR

Asymptotic independence of point process and Frobenius norm of a large sample covariance matrix

A joint limit theorem for the point process of the off-diagonal entries of a sample covariance matrix $\mathbf{S}$, constructed from $n$ observations of a $p$-dimensional random vector with iid components, and the Frobenius norm of $\mathbf{S}$ is proved. In particular, assuming that $p$ and $n$ tend to infinity we obtain a central limit theorem for the Frobenius norm in the case of finite fourth moment of the components and an infinite variance stable law in the case of infinite fourth moment. Extending a theorem of Kallenberg, we establish asymptotic independence of the point process and the Frobenius norm of $\mathbf{S}$. To the best of our knowledge, this is the first result about joint convergence of a point process of dependent points and their sum in the non-Gaussian case.

math.PR

Logarithmic law of large random correlation matrices

Consider a random vector $\mathbf{y}=\mathbfΣ^{1/2}\mathbf{x}$, where the $p$ elements of the vector $\mathbf{x}$ are i.i.d. real-valued random variables with zero mean and finite fourth moment, and $\mathbfΣ^{1/2}$ is a deterministic $p\times p$ matrix such that the spectral norm of the population correlation matrix $\mathbf{R}$ of $\mathbf{y}$ is uniformly bounded. In this paper, we find that the log determinant of the sample correlation matrix $\hat{\mathbf{R}}$ based on a sample of size $n$ from the distribution of $\mathbf{y}$ satisfies a CLT (central limit theorem) for $p/n\to γ\in (0, 1]$ and $p\leq n$. Explicit formulas for the asymptotic mean and variance are provided. In case the mean of $\mathbf{y}$ is unknown, we show that after recentering by the empirical mean the obtained CLT holds with a shift in the asymptotic mean. This result is of independent interest in both large dimensional random matrix theory and high-dimensional statistical literature of large sample correlation matrices for non-normal data. At last, the obtained findings are applied for testing of uncorrelatedness of $p$ random variables. Surprisingly, in the null case $\mathbf{R}=\mathbf{I}$, the test statistic becomes completely pivotal and the extensive simulations show that the obtained CLT also holds if the moments of order four do not exist at all, which conjectures a promising and robust test statistic for heavy-tailed high-dimensional data.

math.ST

Log determinant of large correlation matrices under infinite fourth moment

In this paper, we show the central limit theorem for the logarithmic determinant of the sample correlation matrix $\mathbf{R}$ constructed from the $(p\times n)$-dimensional data matrix $\mathbf{X}$ containing independent and identically distributed random entries with mean zero, variance one and infinite fourth moments. Precisely, we show that for $p/n\to γ\in (0,1)$ as $n,p\to \infty$ the logarithmic law \begin{equation*} \frac{\log \det \mathbf{R} -(p-n+\frac{1}{2})\log(1-p/n)+p-p/n}{\sqrt{-2\log(1-p/n)- 2 p/n}} \overset{d}{\rightarrow} N(0,1)\, \end{equation*} is still valid if the entries of the data matrix $\mathbf{X}$ follow a symmetric distribution with a regularly varying tail of index $α\in (3,4)$. The latter assumptions seem to be crucial, which is justified by the simulations: if the entries of $\mathbf{X}$ have the infinite absolute third moment and/or their distribution is not symmetric, the logarithmic law is not valid anymore. The derived results highlight that the logarithmic determinant of the sample correlation matrix is a very stable and flexible statistic for heavy-tailed big data and open a novel way of analysis of high-dimensional random matrices with self-normalized entries.

math.PR

Limiting distributions for eigenvalues of sample correlation matrices from heavy-tailed populations

Consider a $p$-dimensional population ${\mathbf x} \in\mathbb{R}^p$ with iid coordinates in the domain of attraction of a stable distribution with index $α\in (0,2)$. Since the variance of ${\mathbf x}$ is infinite, the sample covariance matrix ${\mathbf S}_n=n^{-1}\sum_{i=1}^n {{\mathbf x}_i}{\mathbf x}'_i$ based on a sample ${\mathbf x}_1,\ldots,{\mathbf x}_n$ from the population is not well behaved and it is of interest to use instead the sample correlation matrix ${\mathbf R}_n= \{\operatorname{diag}({\mathbf S}_n)\}^{-1/2}\, {\mathbf S}_n \{\operatorname{diag}({\mathbf S}_n)\}^{-1/2}$. This paper finds the limiting distributions of the eigenvalues of ${\mathbf R}_n$ when both the dimension $p$ and the sample size $n$ grow to infinity such that $p/n\to γ\in (0,\infty)$. The family of limiting distributions $\{H_{α,γ}\}$ is new and depends on the two parameters $α$ and $γ$. The moments of $H_{α,γ}$ are fully identified as sum of two contributions: the first from the classical Marčenko-Pastur law and a second due to heavy tails. Moreover, the family $\{H_{α,γ}\}$ has continuous extensions at the boundaries $α=2$ and $α=0$ leading to the Marčenko-Pastur law and a modified Poisson distribution, respectively. Our proofs use the method of moments, the path-shortening algorithm developed in [18] and some novel graph counting combinatorics. As a consequence, the moments of $H_{α,γ}$ are expressed in terms of combinatorial objects such as Stirling numbers of the second kind. A simulation study on these limiting distributions $H_{α,γ}$ is also provided for comparison with the Marčenko-Pastur law.

math.PR

Limiting spectral distribution for large sample correlation matrices

In this paper, we consider the empirical spectral distribution of the sample correlation matrix and investigate its asymptotic behavior under mild assumptions on the data's distribution, when dimension and sample size increase at the same rate. First, we give a characterization for the limiting spectral distribution to follow a Marchenko-Pastur law assuming that the underlying data matrix consists of i.i.d. entries. Subsequently, we provide the limiting spectral distribution of the sample correlation matrix when allowing for a dependence structure within the columns of the data matrix. In contrast to previous works, the fourth moment of the data may be infinite, resulting in a fundamental structural difference. More precisely, the standard argument of approximating the sample correlation matrix by its sample covariance companion breaks down and novel techniques for tackling the challenging dependency structure of the sample correlation matrix are introduced.

math.PR

Large Sample Covariance Matrices of Gaussian Observations with Uniform Correlation Decay

We derive the Marchenko-Pastur (MP) law for sample covariance matrices of the form $V_n=\frac{1}{n}XX^T$, where $X$ is a $p\times n$ data matrix and $p/n\to y\in(0,\infty)$ as $n,p \to \infty$. We assume the data in $X$ stems from a correlated joint normal distribution. In particular, the correlation acts both across rows and across columns of $X$, and we do not assume a specific correlation structure, such as separable dependencies. Instead, we assume that correlations converge uniformly to zero at a speed of $a_n/n$, where $a_n$ may grow mildly to infinity. We employ the method of moments tightly: We identify the exact condition on the growth of $a_n$ which will guarantee that the moments of the empirical spectral distributions (ESDs) converge to the MP moments. If the condition is not met, we can construct an ensemble for which all but finitely many moments of the ESDs diverge. We also investigate the operator norm of $V_n$ under a uniform correlation bound of $C/n^δ$, where $C,δ>0$ are fixed, and observe a phase transition at $δ=1$. In particular, convergence of the operator norm to the maximum of the support of the MP distribution can only be guaranteed if $δ>1$. The analysis leads to an example for which the MP law holds almost surely, but the operator norm remains stochastic in the limit, and we provide its exact limiting distribution.

math.PR

Large sample correlation matrices: a comparison theorem and its applications

In this paper, we show that the diagonal of a high-dimensional sample covariance matrix stemming from $n$ independent observations of a $p$-dimensional time series with finite fourth moments can be approximated in spectral norm by the diagonal of the population covariance matrix. We assume that $n,p\to \infty$ with $p/n$ tending to a constant which might be positive or zero. As applications, we provide an approximation of the sample correlation matrix ${\mathbf R}$ and derive a variety of results for its eigenvalues. We identify the limiting spectral distribution of ${\mathbf R}$ and construct an estimator for the population correlation matrix and its eigenvalues. Finally, the almost sure limits of the extreme eigenvalues of ${\mathbf R}$ in a generalized spiked correlation model are analyzed.

math.PR

On estimation of quadratic variation for multivariate pure jump semimartingales

In this paper we present the asymptotic analysis of the realised quadratic variation for multivariate symmetric $β$-stable Lévy processes, $β\in (0,2)$, and certain pure jump semimartingales. The main focus is on derivation of functional limit theorems for the realised quadratic variation and its spectrum. We will show that the limiting process is a matrix-valued $β$-stable Lévy process when the original process is symmetric $β$-stable, while the limit is conditionally $β$-stable in case of integrals with respect to symmetric $β$-stable motions. These asymptotic results are mostly related to the work [5], which investigates the univariate version of the problem. Furthermore, we will show the implications for estimation of eigenvalues and eigenvectors of the quadratic variation matrix, which is a useful result for the principle component analysis. Finally, we propose a consistent subsampling procedure in the Lévy setting to obtain confidence regions.

math.PR