SearcharxivSearch

arXiv subjects

Tiefeng Jiang

Publications and source records attributed to Tiefeng Jiang.

At least 19 recordsLinked to original sources

Extreme principal minors of Wishart and deformed GOE matrices

We study the laws of large numbers for the largest eigenvalues among all principal minors of Wishart matrices and deformed GOE matrices. We propose a new method based on identifying the deterministic sets to which the random sets formed by suitably normalized principal minors converge in Hausdorff distance, thereby reducing the original extreme-value problems to finite-dimensional convex optimization problems. We demonstrate the effectiveness of this method in regimes not covered by the existing second-moment arguments in \cite{cai2021asymptotic,hu2023extreme}. For deformed GOE matrices with fixed minor size \(k\), we determine the limit for every diagonal variance \(a>0\) and identify a phase transition at \(a=2\). Above the transition, the limiting constant satisfies an explicit recursion with no close-form expression, and the optimizers exhibit a nested hierarchical structure, thereby resolving the case left open in \cite{cai2021asymptotic}. For Wishart matrices with general sub-Gaussian entries and fixed \(k\), we characterize the limit through an entropy-constrained deterministic convex set. When the entries are standard Gaussian, we solve the resulting optimization problem explicitly and obtain the exact value of the limiting constant.

math.PR

Detecting non-uniform patterns on high-dimensional hyperspheres

We propose a new probabilistic characterization of the uniform distribution on the hypersphere in terms of the distribution of pairwise inner products, extending the ideas of \citep{cuesta2009projection,cuesta2007sharp} in a data-driven manner. This characterization naturally leads to an Ingster-type distance for quantifying deviations from uniformity, whose asymptotic behavior can be analyzed systematically via Edgeworth-type expansions. Perhaps surprisingly, we show that this distance captures the minimax rates for testing uniformity simultaneously across several high-dimensional parametric models, even in the models where densities with respect to the uniform law do not exist. We then introduce a simple test for spherical uniformity based on this distance and study its detection rates and consistency against various classes of alternatives, both local and non-local. The proposed test is universally consistent in fixed dimensions, minimax-optimal over a variety of high-dimensional parametric models, and consistent against non-local high-dimensional alternatives. This is different from previously studied high-dimensional Sobolev tests and extreme-value-based tests, which are rate-suboptimal or inconsistent against one or more classes of alternatives. We also establish the local asymptotic distribution of the proposed test under the considered classes of alternatives, along with new information lower bounds.

math.ST

Maximum of sparsely equicorrelated Gaussian fields and applications

We investigate the extreme values of a sparse and equicorrelated Gaussian field on a triangle: the correlations on every vertical or horizontal line are all equal to a parameter $r \in [0,1/2]$ and are zero everywhere else. This problem is closely linked with various problems in high-dimensional statistics and extreme-value theory. We identify the threshold for $r$ at which the standard Gumbel law breaks down. Our result is based on a subtle application of the Chen-Stein method for Poisson approximation. As applications, we discuss the implication of our results on multiple testing and resolve several questions that were left open in \cite{heiny2024maximum}, \cite{tang2022asymptotic} and \cite{Jiang19}.

math.PR

Eigenvalues of Product of Ginibre Ensembles and Their Inverses and that of Truncated Haar Unitary Matrices and Their Inverses

Consider two types of products of independent random matrices, including products of Ginibre matrices and inverse Ginibre matrices and products of truncated Haar unitary matrices and inverse truncated Haar matrices. Each product matrix has $m$ multiplicands of $n$ by $n$ square matrices, and the empirical distribution based on the $n$ eigenvalues of the product matrix is called empirical spectral distribution of the matrix. In this paper, we investigate the limiting empirical spectral distribution of the product matrices when $n$ tends to infinity and $m$ changes with $n$. For properly scaled eigenvalues for two types of the product matrices, we obtain the necessary and sufficient conditions for the convergence of the empirical spectral distributions.

math.PR

Asymptotic analysis of high-dimensional uniformity tests under heavy-tailed alternatives

We study the high-dimensional uniformity testing problem, which involves testing whether the underlying distribution is the uniform distribution, given $n$ data points on the $p$-dimensional unit hypersphere. While this problem has been extensively studied in scenarios with fixed $p$, only three testing procedures are known in high-dimensional settings: the Rayleigh test \cite{Cutting-P-V}, the Bingham test \cite{Cutting-P-V2}, and the packing test \cite{Jiang13}. Most existing research focuses on the former two tests, and the consistency of the packing test remains open. We show that under certain classes of alternatives involving projections of heavy-tailed distributions, the Rayleigh test is asymptotically blind, and the Bingham test has asymptotic power equivalent to random guessing. In contrast, we show theoretically that the packing test is powerful against such alternatives, and empirically that its size suffers from severe distortion due to the slow convergence nature of extreme-value statistics. By exploiting the asymptotic independence of these three tests, we then propose a new test based on Fisher's combination technique that combines their strengths. The new test is shown to enjoy all the optimality properties of each individual test, and unlike the packing test, it maintains excellent type-I error control.

math.ST

Largest Eigenvalues of Principal Minors of Deformed Gaussian Orthogonal Ensembles and Wishart Matrices

Consider a high-dimensional Wishart matrix $\bd{W}=\bd{X}^T\bd{X}$ where the entries of $\bd{X}$ are i.i.d. random variables with mean zero, variance one, and a finite fourth moment $η$. Motivated by problems in signal processing and high-dimensional statistics, we study the maximum of the largest eigenvalues of any two-by-two principal minors of $\bd{W}$. Under certain restrictions on the sample size and the population dimension of $\bd{W}$, we obtain the limiting distribution of the maximum, which follows the Gumbel distribution when $η$ is between 0 and 3, and a new distribution when $η$ exceeds 3. To derive this result, we first address a simpler problem on a new object named a deformed Gaussian orthogonal ensemble (GOE). The Wishart case is then resolved using results from the deformed GOE and a high-dimensional central limit theorem. Our proof strategy combines the Stein-Poisson approximation method, conditioning, U-statistics, and the Hájek projection. This method may also be applicable to other extreme-value problems. Some open questions are posed.

math.PR

Fast Sampling and Inference via Preconditioned Langevin Dynamics

Sampling from distributions play a crucial role in aiding practitioners with statistical inference. However, in numerous situations, obtaining exact samples from complex distributions is infeasible. Consequently, researchers often turn to approximate sampling techniques to address this challenge. Fast approximate sampling from complicated distributions has gained much traction in the last few years with considerable progress in this field. Previous work has shown that for some problems a preconditioning can make the algorithm faster. In our research, we explore the Langevin Monte Carlo (LMC) algorithm and demonstrate its effectiveness in enabling inference from the obtained samples. Additionally, we establish a convergence rate for the LMC Markov chain in total variation. Lastly, we derive non-asymptotic bounds for approximate sampling from specific target distributions in the Wasserstein distance, particularly when the preconditioning is spatially invariant.

stat.CO

Asymptotic Distributions of Largest Pearson Correlation Coefficients under Dependent Structures

Given a random sample from a multivariate normal distribution whose covariance matrix is a Toeplitz matrix, we study the largest off-diagonal entry of the sample correlation matrix. Assuming the multivariate normal distribution has the covariance structure of an auto-regressive sequence, we establish a phase transition in the limiting distribution of the largest off-diagonal entry. We show that the limiting distributions are of Gumbel-type (with different parameters) depending on how large or small the parameter of the autoregressive sequence is. At the critical case, we obtain that the limiting distribution is the maximum of two independent random variables of Gumbel distributions. This phase transition establishes the exact threshold at which the auto-regressive covariance structure behaves differently than its counterpart with the covariance matrix equal to the identity. Assuming the covariance matrix is a general Toeplitz matrix, we obtain the limiting distribution of the largest entry under the ultra-high dimensional settings: it is a weighted sum of two independent random variables, one normal and the other following a Gumbel-type law. The counterpart of the non-Gaussian case is also discussed. As an application, we study a high-dimensional covariance testing problem.

math.ST

Non-Asymptotic Analysis of Online Multiplicative Stochastic Gradient Descent

Past research has indicated that the covariance of the Stochastic Gradient Descent (SGD) error done via minibatching plays a critical role in determining its regularization and escape from low potential points. Motivated by some new research in this area, we prove universality results by showing that noise classes that have the same mean and covariance structure of SGD via minibatching have similar properties. We mainly consider the Multiplicative Stochastic Gradient Descent (M-SGD) algorithm as introduced in previous work, which has a much more general noise class than the SGD algorithm done via minibatching. We establish non asymptotic bounds for the M-SGD algorithm in the Wasserstein distance. We also show that the M-SGD error is approximately a scaled Gaussian distribution with mean $0$ at any fixed point of the M-SGD algorithm.

stat.ML

Asymptotic Properties of Random Restricted Partitions

We study two types of probability measures on the set of integer partitions of $n$ with at most $m$ parts. The first one chooses the random partition with a chance related to its largest part only. We then obtain the limiting distributions of all of the parts together and that of the largest part as $n$ tends to infinity while $m$ is fixed or tends to infinity. In particular, if $m$ goes to infinity not too fast, the largest part satisfies the central limit theorem. The second measure is very general. It includes the Dirichlet distribution and the uniform distribution as special cases. We derive the asymptotic distributions of the parts jointly by taking limits of $n$ and $m$ in the same manner as that in the first probability measure.

math.PR

Contiguity under high dimensional Gaussianity with applications to covariance testing

Le Cam's third/contiguity lemma is a fundamental probabilistic tool to compute the limiting distribution of a given statistic $T_n$ under a non-null sequence of probability measures $\{Q_n\}$, provided its limiting distribution under a null sequence $\{P_n\}$ is available, and the log likelihood ratio $\{\log (dQ_n/dP_n)\}$ has a distributional limit. Despite its wide-spread applications to low-dimensional statistical problems, the stringent requirement of Le Cam's third/contiguity lemma on the distributional limit of the log likelihood ratio makes it challenging, or even impossible to use in many modern high-dimensional statistical problems. This paper provides a non-asymptotic analogue of Le Cam's third/contiguity lemma under high dimensional normal populations. Our contiguity method is particularly compatible with sufficiently regular statistics $T_n$: the regularity of $T_n$ effectively reduces both the problems of (i) obtaining a null (Gaussian) limit distribution and of (ii) verifying our new quantitative contiguity condition, to those of derivative calculations and moment bounding exercises. More important, our method bypasses the need to understand the precise behavior of the log likelihood ratio, and therefore possibly works even when it necessarily fails to stabilize -- a regime beyond the reach of classical contiguity methods. As a demonstration of the scope of our new contiguity method, we obtain asymptotically exact power formulae for a number of widely used high-dimensional covariance tests, including the likelihood ratio tests and trace tests, that hold uniformly over all possible alternative covariance under mild growth conditions on the dimension-to-sample ratio. These new results go much beyond the scope of previous available case-specific techniques, and exhibit new phenomenon regarding the behavior of these important class of covariance tests.

math.ST

Asymptotic Independence of the Sum and Maximum of Dependent Random Variables with Applications to High-Dimensional Tests

For a set of dependent random variables, without stationary or the strong mixing assumptions, we derive the asymptotic independence between their sums and maxima. Then we apply this result to high-dimensional testing problems, where we combine the sum-type and max-type tests and propose a novel test procedure for the one-sample mean test, the two-sample mean test and the regression coefficient test in high-dimensional setting. Based on the asymptotic independence between sums and maxima, the asymptotic distributions of test statistics are established. Simulation studies show that our proposed tests have good performance regardless of data being sparse or not. Examples on real data are also presented to demonstrate the advantages of our proposed methods.

stat.ME

Mean Test with Fewer Observation than Dimension and Ratio Unbiased Estimator for Correlation Matrix

Hotelling's T-squared test is a classical tool to test if the normal mean of a multivariate normal distribution is a specified one or the means of two multivariate normal means are equal. When the population dimension is higher than the sample size, the test is no longer applicable. Under this situation, in this paper we revisit the tests proposed by Srivastava and Du (2008), who revise the Hotelling's statistics by replacing Wishart matrices with their diagonal matrices. They show the revised statistics are asymptotically normal. We use the random matrix theory to examine their statistics again and find that their discovery is just part of the big picture. In fact, we prove that their statistics, decided by the Euclidean norm of the population correlation matrix, can go to normal, mixing chi-squared distributions and a convolution of both. Examples are provided to show the phase transition phenomenon between the normal and mixing chi-squared distributions. The second contribution of ours is a rigorous derivation of an asymptotic ratio-unbiased-estimator of the squared Euclidean norm of the correlation matrix.

math.ST

Statistical properties of eigenvalues of Laplace-Beltrami operators

We study the eigenvalues of a Laplace-Beltrami operator defined on the set of the symmetric polynomials, where the eigenvalues are expressed in terms of partitions of integers. By assigning partitions with the restricted uniform measure, the restricted Jack measure, the uniform measure or the Plancherel measure, we prove that the global distribution of the eigenvalues is asymptotically a new distribution $μ$, the Gamma distribution, the Gumbel distribution and the Tracy-Widom distribution, respectively. An explicit representation of $μ$ is obtained by a function of independent random variables. We also derive an independent result on random partitions itself: a law of large numbers for the restricted uniform measure. Two open problems are also asked.

math.PR

Max-sum tests for cross-sectional dependence of high-demensional panel data

We consider a testing problem for cross-sectional dependence for high-dimensional panel data, where the number of cross-sectional units is potentially much larger than the number of observations. The cross-sectional dependence is described through a linear regression model. We study three tests named the sum test, the max test and the max-sum test, where the latter two are new. The sum test is initially proposed by Breusch and Pagan (1980). We design the max and sum tests for sparse and non-sparse residuals in the linear regressions, respectively.And the max-sum test is devised to compromise both situations on the residuals. Indeed, our simulation shows that the max-sum test outperforms the previous two tests. This makes the max-sum test very useful in practice where sparsity or not for a set of data is usually vague. Towards the theoretical analysis of the three tests, we have settled two conjectures regarding the sum of squares of sample correlation coefficients asked by Pesaran (2004 and 2008). In addition, we establish the asymptotic theory for maxima of sample correlations coefficients appeared in the linear regression model for panel data, which is also the first successful attempt to our knowledge. To study the max-sum test, we create a novel method to show asymptotic independence between maxima and sums of dependent random variables. We expect the method itself is useful for other problems of this nature. Finally, an extensive simulation study as well as a case study are carried out. They demonstrate advantages of our proposed methods in terms of both empirical powers and robustness for residuals regardless of sparsity or not.

math.ST

Limiting behavior of largest entry of random tensor constructed by high-dimensional data

Let ${X}_{k}=(x_{k1}, \cdots, x_{kp})', k=1,\cdots,n$, be a random sample of size $n$ coming from a $p$-dimensional population. For a fixed integer $m\geq 2$, consider a hypercubic random tensor $\mathbf{T}$ of $m$-th order and rank $n$ with \begin{eqnarray*} \mathbf{T}= \sum_{k=1}^{n}\underbrace{{X}_{k}\otimes\cdots\otimes {X}_{k}}_{m~multiple}=\Big(\sum_{k=1}^{n} x_{ki_{1}}x_{ki_{2}}\cdots x_{ki_{m}}\Big)_{1\leq i_{1},\cdots, i_{m}\leq p}. \end{eqnarray*} Let $W_n$ be the largest off-diagonal entry of $\mathbf{T}$. We derive the asymptotic distribution of $W_n$ under a suitable normalization for two cases. They are the ultra-high dimension case with $p\to\infty$ and $\log p=o(n^β)$ and the high-dimension case with $p\to \infty$ and $p=O(n^α)$ where $α,β>0$. The normalizing constant of $W_n$ depends on $m$ and the limiting distribution of $W_n$ is a Gumbel-type distribution involved with parameter $m$.

math.PR

Likelihood Ratio Test in Multivariate Linear Regression: from Low to High Dimension

Multivariate linear regressions are widely used statistical tools in many applications to model the associations between multiple related responses and a set of predictors. To infer such associations, it is often of interest to test the structure of the regression coefficients matrix, and the likelihood ratio test (LRT) is one of the most popular approaches in practice. Despite its popularity, it is known that the classical $χ^2$ approximations for LRTs often fail in high-dimensional settings, where the dimensions of responses and predictors $(m,p)$ are allowed to grow with the sample size $n$. Though various corrected LRTs and other test statistics have been proposed in the literature, the fundamental question of when the classic LRT starts to fail is less studied, an answer to which would provide insights for practitioners, especially when analyzing data with $m/n$ and $p/n$ small but not negligible. Moreover, the power performance of the LRT in high-dimensional data analysis remains underexplored. To address these issues, the first part of this work gives the asymptotic boundary where the classical LRT fails and develops the corrected limiting distribution of the LRT for a general asymptotic regime. The second part of this work further studies the test power of the LRT in the high-dimensional setting. The result not only advances the current understanding of asymptotic behavior of the LRT under alternative hypothesis, but also motivates the development of a power-enhanced LRT. The third part of this work considers the setting with $p>n$, where the LRT is not well-defined. We propose a two-step testing procedure by first performing dimension reduction and then applying the proposed LRT. Theoretical properties are developed to ensure the validity of the proposed method. Numerical studies are also presented to demonstrate its good performance.

math.ST

Asymptotic Analysis for Extreme Eigenvalues of Principal Minors of Random Matrices

Consider a standard white Wishart matrix with parameters $n$ and $p$. Motivated by applications in high-dimensional statistics and signal processing, we perform asymptotic analysis on the maxima and minima of the eigenvalues of all the $m \times m$ principal minors, under the asymptotic regime that $n,p,m$ go to infinity. Asymptotic results concerning extreme eigenvalues of principal minors of real Wigner matrices are also obtained. In addition, we discuss an application of the theoretical results to the construction of compressed sensing matrices, which provides insights to compressed sensing in signal processing and high dimensional linear regression in statistics.

math.ST