SearcharxivSearch

arXiv subjects

Jiahui Xie

Publications and source records attributed to Jiahui Xie.

12 recordsLinked to original sources

Multiplier Bootstrap and Edge Phase Transitions of High-Dimensional Covariance Matrices

In this paper, we study the effects of employing multiplier bootstrap to analyze the asymptotic distributions of the largest eigenvalues of high-dimensional sample covariance matrices in both spiked and non-spiked models. Our findings demonstrate that the multiplier bootstrap establishes several phase transitions in the limiting edge distributions of both unconditional and conditional bootstrapped covariance matrices, provided the different classes of multipliers. In the nonspiked setting, unbounded multipliers lead to Frechet or Gumbel limits for the largest eigenvalue of the bootstrapped covariance matrix, both conditionally on the observed data and unconditionally. For bounded multipliers, the unconditional model exhibits transitions among Tracy-Widom, Gaussian, or Weibull limits, determined jointly by the aspect ratio p/n, the upper-endpoint behavior of the multipliers, and the population covariance matrix. The conditional model displays analogous Gaussian and Weibull regimes; in contrast, the conditional counterpart of the unconditional Tracy-Widom regime collapses to a point mass. In the spiked setting, under suitable signal-strength conditions, the leading eigenvalues of both the unconditional and conditional bootstrapped sample covariance matrices are asymptotically Gaussian for bounded as well as unbounded multipliers, under some mild assumptions. Our theoretical results also clarify the feasibility and adaptability of the multiplier bootstrap for spectral inference in high-dimensional sample covariance models. Numerical simulations confirm the accuracy of our results and the effectiveness of the proposed spectral inference procedures, which may be of independent interest.

math.ST

Bias-Corrected Multiplier Bootstrap Inference for Spectral Edges of Large Covariance Matrices

Inference for spectral edges of large covariance matrices is a fundamental problem in high-dimensional statistics. A major difficulty is that the largest non-spiked sample eigenvalues, which serve as natural estimators of the edge, fluctuate on the Tracy--Widom scale. Consequently, valid inference requires accurate centering by the deterministic spectral edge together with a precise scaling constant, both of which are often difficult to estimate in practice under general unknown population covariance structures. In this paper, we propose a bias-corrected multiplier bootstrap procedure for inference on the deterministic edge of the bulk spectrum. The key idea is to introduce a carefully calibrated multiplier perturbation that regularizes the edge fluctuation to a slightly larger scale at which Gaussian approximation becomes tractable. The resulting confidence interval is constructed directly from bootstrap eigenvalues, together with a data-driven recentering step that corrects the bootstrap-induced shift of the deterministic edge. On the theoretical side, we show that, after bias correction and rescaling, the largest few non-spiked bootstrap eigenvalues are asymptotically Gaussian conditionally on the data. Building on this result, we establish the asymptotic validity of the proposed confidence interval, whose length is only slightly larger than the Tracy--Widom scale, and prove vanishing coverage under alternatives in which additional spikes separate from the bulk at a local scale larger than $n^{-1/6}$. As a consequence, the same confidence interval yields a threshold-free estimator for the number of spikes, without requiring the spikes to be distinct or very large. Equivalently, the procedure yields a data-driven and theoretically justified cutoff for the scree plot.

stat.ME

The logarithmic law of sample correlation matrices

Let $\mathbf{R}$ be the sample correlation matrix constructed from $\mathbf{X}\in \mathbb{R}^{p\times n}$, whose entries are independent and identically distributed random variables with mean zero and tail probability condition $\lim_{x\rightarrow \infty}x^3\mathbb{P}(|\xi|>x)=0$. We derive the universal logarithmic law for $\log \det \mathbf{R}$, \begin{equation*} \frac{\log \det \mathbf{R}-(p-n+1/2)\log (1-\frac{p-1}{n})+p-\frac{p}{n}}{\sqrt{-2\log (1-\frac{p-1}{n})-2\frac{p}{n}}}\stackrel{d}{\rightarrow} {N}(0,1), \end{equation*} if $p\le n$ as $p,n\rightarrow \infty$. Moreover, under the near-singularity case $0\le n-p\le n^{1-w}$ for any $w\in (0,1)$, it is shown that the tail probability condition can be weakened to $\lim_{x\rightarrow \infty}x^3(\log x)^{-1/4+\mathfrak{c}}\mathbb{P}(|\xi|>x)<\infty$ for any constant $0<\mathfrak{c}<1/4$.

math.PR

On Convergence Rates of Spiked Eigenvalue Estimates: A General Study of Global and Local Laws in Sample Covariance Matrices

This paper investigates global and local laws for sample covariance matrices with general growth rates of dimensions. The sample size $N$ and population dimension $M$ can have the same order in logarithm, which implies that their ratio $M/N$ can approach zero, a constant, or infinity. These theories are utilized to determine the convergence rate of spiked eigenvalue estimates.

math.ST

The Spurious Factor Dilemma: Robust Inference in Heavy-Tailed Elliptical Factor Models

Standard methods for determining the number of factors often overestimate the true number when data exhibit heavy-tailed randomness, misinterpreting noise-induced outliers as genuine factors. This paper addresses this challenge within the framework of Elliptical Factor Models (EFM), which accommodate both heavy tails and potential non-linear dependencies common in real-world data. We demonstrate, both theoretically and empirically, that heavy-tailed noise generates spurious eigenvalues that mimic true factor signals. To distinguish these, we propose a novel methodology based on a fluctuation magnification algorithm. Under mild conditions, we show that, by magnifying perturbations, the eigenvalues associated with real factors exhibit significantly less fluctuation (stabilizing asymptotically) than spurious eigenvalues arising from heavy-tailed effects. We develop a formal testing procedure based on this principle and apply it to the problem of accurately selecting the number of common factors in heavy-tailed EFMs. Simulation studies and real data analysis confirm the effectiveness of our approach, particularly in scenarios with pronounced heavy-tailedness.

stat.ME

Representational Transfer Learning for Matrix Completion

We propose to transfer representational knowledge from multiple sources to a target noisy matrix completion task by aggregating singular subspaces information. Under our representational similarity framework, we first integrate linear representation information by solving a two-way principal component analysis problem based on a properly debiased matrix-valued dataset. After acquiring better column and row representation estimators from the sources, the original high-dimensional target matrix completion problem is then transformed into a low-dimensional linear regression, of which the statistical efficiency is guaranteed. A variety of extensional arguments, including post-transfer statistical inference and robustness against negative transfer, are also discussed alongside. Finally, extensive simulation results and a number of real data cases are reported to support our claims.

stat.ML

Necessary and sufficient condition for CLT of linear spectral statistics of sample correlation matrices

In this paper, we establish the central limit theorem (CLT) for the linear spectral statistics (LSS) of sample correlation matrix $R$, constructed from a $p\times n$ data matrix $X$ with independent and identically distributed (i.i.d.) entries having mean zero, variance one, and infinite fourth moments in the high-dimensional regime $n/p\rightarrow \phi\in \mathbb{R}_+\backslash \{1\}$. We derive a necessary and sufficient condition for the CLT. More precisely, under the assumption that the identical distribution $\xi$ of the entries in $X$ satisfies $\mathbb{P}(|\xi|>x)\sim l(x)x^{-\alpha}$ when $x\rightarrow \infty$ for $\alpha \in (2,4]$, where $l(x)$ is a slowly varying function, we conclude that: (i). When $\alpha\in(3,4]$, the universal asymptotic normality for the LSS of sample correlation matrix holds, with the same asymptotic mean and variance as in the finite fourth moment scenario; (ii) We identify a necessary and sufficient condition $\lim_{x\rightarrow\infty}x^3\mathbb{P}(|\xi|>x)=0$ for the universal CLT; (iii) We establish a local law for $\alpha \in (2, 4]$. Overall, our proof strategy follows the routine of the matrix resampling, intermediate local law, Green function comparison, and characteristic function estimation. In various parts of the proof, we are required to come up with new approaches and ideas to solve the challenges posed by the special structure of sample correlation matrix. Our results also demonstrate that the symmetry condition is unnecessary for the CLT of LSS for sample correlation matrix, but the tail index $\alpha$ plays a crucial role in determining the asymptotic behaviors of LSS for $\alpha \in (2, 3)$.

math.PR

Necessity of orthogonal basis vectors for the two-anyon problem in one-dimensional lattice

Few-body physics for anyons has been intensively studied within the anyon-Hubbard model, including the quantum walk and Bloch oscillations of two-anyon states. However, the known theoretical proposal and experimental simulations of two-anyon states in one-dimensional lattice have been carried out by expanding the wavefunction in terms of non-orthogonal basis vectors, which introduces extra non-physical degrees of freedom. In the present work, we deduce the finite difference equations for the two-anyon state in the one-dimensional lattice by solving the Schr\"odinger equation with orthogonal basis vectors. Such an orthogonal scheme gives all the orthogonal physical eigenstates for the time-independent two-anyon Schr\"odinger equation, while the conventional (non-orthogonal) method produces a lot of non-physical redundant eigen-solutions whose components violate the anyonic relations. The dynamical property of the two-anyon states in a sufficiently large lattice has been investigated and compared in both the orthogonal and conventional schemes, which proves to depend crucially on the initial states. When the initial states with two anyons on the same site or (next-)neighboring sites are suitably chosen to be in accordance with the anyonic coefficient relation, we observe exactly the same dynamical behavior in the two schemes, including the revival probability, the probability density function, and the two-body correlation, otherwise, the conventional scheme will produce erroneous results which not any more describe anyons. The period of the Bloch oscillation in the pseudo-fermionic limit is found to be twice that in the bosonic limit, while the oscillations disappear for statistical parameters in between. Our findings are vital for quantum simulations of few-body physics with anyons in the lattice.

cond-mat.quant-gas

Tracy-Widom distribution for the edge eigenvalues of elliptical model

In this paper, we study the largest eigenvalues of sample covariance matrices with elliptically distributed data. We consider the sample covariance matrix $Q=YY^*,$ where the data matrix $Y \in \mathbb{R}^{p \times n}$ contains i.i.d. $p$-dimensional observations $\mathbf{y}_i=ξ_iT\mathbf{u}_i,\;i=1,\dots,n.$ Here $\mathbf{u}_i$ is distributed on the unit sphere, $ξ_i \sim ξ$ is independent of $\mathbf{u}_i$ and $T^*T=Σ$ is some deterministic matrix. Under some mild regularity assumptions of $Σ,$ assuming $ξ^2$ has bounded support and certain proper behavior near its edge so that the limiting spectral distribution (LSD) of $Q$ has a square decay behavior near the spectral edge, we prove that the Tracy-Widom law holds for the largest eigenvalues of $Q$ when $p$ and $n$ are comparably large.

math.PR

Extreme eigenvalues of sample covariance matrices under generalized elliptical models with applications

We consider the extreme eigenvalues of the sample covariance matrix $Q=YY^*$ under the generalized elliptical model that $Y=Σ^{1/2}XD.$ Here $Σ$ is a bounded $p \times p$ positive definite deterministic matrix representing the population covariance structure, $X$ is a $p \times n$ random matrix containing either independent columns sampled from the unit sphere in $\mathbb{R}^p$ or i.i.d. centered entries with variance $n^{-1},$ and $D$ is a diagonal random matrix containing i.i.d. entries and independent of $X.$ Such a model finds important applications in statistics and machine learning. In this paper, assuming that $p$ and $n$ are comparably large, we prove that the extreme edge eigenvalues of $Q$ can have several types of distributions depending on $Σ$ and $D$ asymptotically. These distributions include: Gumbel, Fréchet, Weibull, Tracy-Widom, Gaussian and their mixtures. On the one hand, when the random variables in $D$ have unbounded support, the edge eigenvalues of $Q$ can have either Gumbel or Fréchet distribution depending on the tail decay property of $D.$ On the other hand, when the random variables in $D$ have bounded support, under some mild regularity assumptions on $Σ,$ the edge eigenvalues of $Q$ can exhibit Weibull, Tracy-Widom, Gaussian or their mixtures. Based on our theoretical results, we consider two important applications. First, we propose some statistics and procedure to detect and estimate the possible spikes for elliptically distributed data. Second, in the context of a factor model, by using the multiplier bootstrap procedure via selecting the weights in $D,$ we propose a new algorithm to infer and estimate the number of factors in the factor model. Numerical simulations also confirm the accuracy and powerfulness of our proposed methods and illustrate better performance compared to some existing methods in the literature.

stat.ME

Testing Kronecker Product Covariance Matrices for High-dimensional Matrix-Variate Data

Kronecker product covariance structure provides an efficient way to modeling the inter-correlations of matrix-variate data. In this paper, we propose testing statistics for Kronecker product covariance matrix based on linear spectral statistics of renormalized sample covariance matrices. Central limit theorem is proved for the linear spectral statistics with explicit formulas for mean and covariance functions, which fills the gap in the literature. We then theoretically justify that the proposed testing statistics have well-controlled sizes and strong powers. To facilitate practical usefulness, we further propose a bootstrap resampling algorithm to approximate the limiting distributions of associated linear spectral statistics. Consistency of the bootstrap procedure is guaranteed under mild conditions. A more general model which allows the existence of noises will also be discussed. In the simulations, the empirical sizes of the proposed testing procedure and its bootstrapped version are close to corresponding theoretical values, while the powers converge to one quickly as the dimension and sample size grow.

math.ST

Tracy-Widom limit for the largest eigenvalue of high-dimensional covariance matrices in elliptical distributions

Let $X$ be an $M\times N$ random matrix consisting of independent $M$-variate elliptically distributed column vectors $\mathbf{x}_{1},\dots,\mathbf{x}_{N}$ with general population covariance matrix $Σ$. In the literature, the quantity $XX^{*}$ is referred to as the sample covariance matrix after scaling, where $X^{*}$ is the transpose of $X$. In this article, we prove that the limiting behavior of the scaled largest eigenvalue of $XX^{*}$ is universal for a wide class of elliptical distributions, namely, the scaled largest eigenvalue converges weakly to the same limit regardless of the distributions that $\mathbf{x}_{1},\dots,\mathbf{x}_{N}$ follow as $M,N\to\infty$ with $M/N\toϕ_0>0$ if the weak fourth moment of the radius of $\mathbf{x}_{1}$ exists . In particular, via comparing the Green function with that of the sample covariance matrix of multivariate normally distributed data, we conclude that the limiting distribution of the scaled largest eigenvalue is the celebrated Tracy-Widom law.

math.ST