SearcharxivSearch

arXiv subjects

Guangming Pan

Publications and source records attributed to Guangming Pan.

At least 19 recordsLinked to original sources

Spectral Analysis of Gram Matrices with Missing at Random Observations: Convergence, Central Limit Theorems, and Applications in Statistical Inference

Motivated by the statistical inference using the Gram matrix in the context of missing at random observations, this paper investigates the spectral properties of the random matrices $\mb S_n=\frac{1}{n}\mb Z\mb Z^*$, where $\mb Z=\mb D\circ(\boldsymbol{\Sigma^{1/2}}\mb X)$ represents a Hadamard random matrix with entries determined by independent Bernoulli variables $\mb D$. Operating within the high-dimensional framework, we establish the convergence of the empirical spectral distribution of $\mb S_n$ to a well-defined limiting distribution. In addition, we explore the impact of the missing mechanism on the second-order properties of the spectral distribution of the Gram matrix $\mathbf{S}_n$. We establish the central limit theorem for the linear spectral statistics of $\mathbf{S}_n$, shedding light on their fluctuations. Surprisingly, our analysis reveals that even in the ideal Gaussian distribution scenario, the fluctuations of statistics generated by eigenvalues are influenced by the eigenvectors of the population covariance matrix in the missing-at-random case. This discovery uncovers a remarkable phenomenon that starkly contrasts with the classical case. Subsequently, we demonstrate the practical application of our central limit theorem in hypothesis testing for the population covariance matrix.

math.ST

Multi-kernel spectral clustering: Entrywise eigenvector perturbation bounds and exact recovery

Kernel spectral clustering with a single bandwidth can be inadequate for data exhibiting multiple characteristic pairwise-distance scales, a problem particularly prevalent in the high-dimensional regime. We address this issue through a multi-kernel formulation that aggregates kernels with different bandwidths. The bandwidths are selected as prescribed empirical quantiles of the pairwise squared distances, thereby capturing the relevant distance scales without requiring prior population-scale information. We develop a rigorous theoretical analysis of the resulting method under a general high-dimensional, multi-scale mixture model with heterogeneous cluster centers and covariance geometries. We construct a blockwise constant, low-rank informative approximation to the empirical multi-kernel matrix and establish row-wise $\ell_{2,\infty}$ perturbation bounds for its leading spectral components, as well as for the associated normalized Laplacian matrix. These bounds yield observation-level control of the spectral embedding, which is more informative than conventional global eigenspace perturbation estimates. Under suitable eigen-gap and cluster-separation conditions, we show that approximate $K$-means applied to the multi-kernel spectral embedding achieves exact recovery with high probability.

stat.ML

Phase transition of Schott's statistic for high-dimensional heavy-tailed data

Consider Schott's statistic (Schott, 2005) defined as the squared Frobenius norm of the sample correlation matrix for data from $\alpha$-regularly varying populations. We investigate its asymptotic distribution in a general framework characterized by data dimension p, sample size n, and regularly varying coefficients $\alpha$. In particular, we identify a phase transition phenomenon in the asymptotic behavior. For light-tailed populations ($\alpha > 3$), we revisit the $\alpha$-free asymptotic distribution but relax the constraint on the ratio of $p/n$. For heavy-tailed populations ($\alpha < 3$), we derive a new asymptotic normal distribution whose variance explicitly depends on $\alpha$. We also propose a consistent estimator for the asymptotic variance such that the standardized Schott's test statistic remains applicable for unknown location parameters and all $\alpha > 0$.

math.ST

Phase Transition of Spectral Fluctuations in Large Gram Matrices with a Variance Profile: A Unified Framework for Sparse CLTs

We study the asymptotic spectral behavior of high-dimensional random Gram matrices with sparsity and a variance profile, motivated by applications in wireless communications. Specifically, we consider the Gram matrices $\mathbf S_n=\mathbf Y_n\mathbf Y_n^*$, where the entries of $\mathbf Y_n$ are independent, centered, heteroscedastic, and sparse through Bernoulli masking. The sparsity level is parameterized as $s=q^2/n$, where $q$ ranges from polynomial order up to order $n^{1/2}$. We investigate two asymptotic regimes: a moderate-sparsity regime with fixed $s\in(0,1]$, and a high-sparsity regime where $s\to0$. In both regimes, we establish the convergence of the empirical spectral distribution of $\mathbf S_n$ to a deterministic limit, and further derive central limit theorems for linear spectral statistics using resolvent techniques and martingale difference arguments. Our analysis reveals a phase transition in the fluctuation behavior across the two regimes. In the high-sparsity regime, the asymptotic fluctuations are entirely governed by fourth-moment effects, with sparsity-scaled contributions being suppressed. Moreover, the leading deterministic term and the variance of the linear spectral statistic scale at different rates in $q$, causing the standard centering to fail and necessitating an explicit correction to recover a valid CLT. The results apply to both Gaussian and non-Gaussian entries and are illustrated through applications to hypothesis testing and outage probability analysis in large-scale MIMO systems.

math.ST

High-Dimensional Precision Matrix Quadratic Forms: Estimation Framework for $p > n$

We propose a novel estimation framework for quadratic functionals of precision matrices in high-dimensional settings, particularly in regimes where the feature dimension $p$ exceeds the sample size $n$. Traditional moment-based estimators with bias correction remain consistent when $p n$, highlighting a fundamental distinction between the two regimes due to rank deficiency and high-dimensional complexity. Our approach resolves these issues by combining a spectral-moment representation with constrained optimization, resulting in consistent estimation under mild moment conditions. The proposed framework provides a unified approach for inference on a broad class of high-dimensional statistical measures. We illustrate its utility through two representative examples: the optimal Sharpe ratio in portfolio optimization and the multiple correlation coefficient in regression analysis. Simulation studies demonstrate that the proposed estimator effectively overcomes the fundamental $p>n$ barrier where conventional methods fail.

stat.ME

Simultaneous Detection and Localization of Mean and Covariance Changes in High Dimensions

Existing methods for high-dimensional changepoint detection and localization typically focus on changes in either the mean vector or the covariance matrix separately. This separation reduces detection power and localization accuracy when both parameters change simultaneously. We propose a simple yet powerful method that jointly monitors shifts in both the mean and covariance structures. Under mild conditions, the test statistics for detecting these shifts jointly converge in distribution to a bivariate standard normal distribution, revealing their asymptotic independence. This independence enables the combination of the individual p-values using Fisher's method, and the development of an adaptive p-value-based estimator for the changepoint. Theoretical analysis and extensive simulations demonstrate the superior performance of our method in terms of both detection power and localization accuracy.

math.ST

Decentralized Quantile Regression for Feature-Distributed Massive Datasets with Privacy Guarantees

In this paper, we introduce a novel decentralized surrogate gradient-based algorithm for quantile regression in a feature-distributed setting, where global features are dispersed across multiple machines within a decentralized network. The proposed algorithm, \texttt{DSG-cqr}, utilizes a convolution-type smoothing approach to address the non-smooth nature of the quantile loss function. \texttt{DSG-cqr} is fully decentralized, conjugate-free, easy to implement, and achieves linear convergence up to statistical precision. To ensure privacy, we adopt the Gaussian mechanism to provide $(\epsilon,\delta)$-differential privacy. To overcome the exact residual calculation problem, we estimate residuals using auxiliary variables and develop a confidence interval construction method based on Wald statistics. Theoretical properties are established, and the practical utility of the methods is also demonstrated through extensive simulations and a real-world data application.

stat.CO

Necessary and sufficient condition for CLT of linear spectral statistics of sample correlation matrices

In this paper, we establish the central limit theorem (CLT) for the linear spectral statistics (LSS) of sample correlation matrix $R$, constructed from a $p\times n$ data matrix $X$ with independent and identically distributed (i.i.d.) entries having mean zero, variance one, and infinite fourth moments in the high-dimensional regime $n/p\rightarrow \phi\in \mathbb{R}_+\backslash \{1\}$. We derive a necessary and sufficient condition for the CLT. More precisely, under the assumption that the identical distribution $\xi$ of the entries in $X$ satisfies $\mathbb{P}(|\xi|>x)\sim l(x)x^{-\alpha}$ when $x\rightarrow \infty$ for $\alpha \in (2,4]$, where $l(x)$ is a slowly varying function, we conclude that: (i). When $\alpha\in(3,4]$, the universal asymptotic normality for the LSS of sample correlation matrix holds, with the same asymptotic mean and variance as in the finite fourth moment scenario; (ii) We identify a necessary and sufficient condition $\lim_{x\rightarrow\infty}x^3\mathbb{P}(|\xi|>x)=0$ for the universal CLT; (iii) We establish a local law for $\alpha \in (2, 4]$. Overall, our proof strategy follows the routine of the matrix resampling, intermediate local law, Green function comparison, and characteristic function estimation. In various parts of the proof, we are required to come up with new approaches and ideas to solve the challenges posed by the special structure of sample correlation matrix. Our results also demonstrate that the symmetry condition is unnecessary for the CLT of LSS for sample correlation matrix, but the tail index $\alpha$ plays a crucial role in determining the asymptotic behaviors of LSS for $\alpha \in (2, 3)$.

math.PR

Penalized Principal Component Analysis for Large-dimension Factor Model with Group Pursuit

This paper investigates the intrinsic group structures within the framework of large-dimensional approximate factor models, which portrays homogeneous effects of the common factors on the individuals that fall into the same group. To this end, we propose a fusion Penalized Principal Component Analysis (PPCA) method and derive a closed-form solution for the $\ell_2$-norm optimization problem. We also show the asymptotic properties of our proposed PPCA estimates. With the PPCA estimates as an initialization, we identify the unknown group structure by a combination of the agglomerative hierarchical clustering algorithm and an information criterion. Then the factor loadings and factor scores are re-estimated conditional on the identified latent groups. Under some regularity conditions, we establish the consistency of the membership estimators as well as that of the group number estimator derived from the information criterion. Theoretically, we show that the post-clustering estimators for the factor loadings and factor scores with group pursuit achieve efficiency gains compared to the estimators by conventional PCA method. Thorough numerical studies validate the established theory and a real financial example illustrates the practical usefulness of the proposed method.

stat.ME

Eigenvector overlaps in large sample covariance matrices and nonlinear shrinkage estimators

Consider a data matrix $Y = [\mathbf{y}_1, \cdots, \mathbf{y}_N]$ of size $M \times N$, where the columns are independent observations from a random vector $\mathbf{y}$ with zero mean and population covariance $\Sigma$. Let $\mathbf{u}_i$ and $\mathbf{v}_j$ denote the left and right singular vectors of $Y$, respectively. This study investigates the eigenvector/singular vector overlaps $\langle {\mathbf{u}_i, D_1 \mathbf{u}_j} \rangle$, $\langle {\mathbf{v}_i, D_2 \mathbf{v}_j} \rangle$ and $\langle {\mathbf{u}_i, D_3 \mathbf{v}_j} \rangle$, where $D_k$ are general deterministic matrices with bounded operator norms. We establish the convergence in probability of these eigenvector overlaps toward their deterministic counterparts with explicit convergence rates, when the dimension $M$ scales proportionally with the sample size $N$. Building on these findings, we offer a more precise characterization of the loss for Ledoit and Wolf's nonlinear shrinkage estimators of the population covariance $\Sigma$.

math.ST

Asymptotic distribution of spiked eigenvalues in the large signal-plus-noise models

Consider large signal-plus-noise data matrices of the form $S + \Sigma^{1/2} X$, where $S$ is a low-rank deterministic signal matrix and the noise covariance matrix $\Sigma$ can be anisotropic. We establish the asymptotic joint distribution of its spiked singular values when the dimensionality and sample size are comparably large and the signals are supercritical under general assumptions concerning the structure of $(S, \Sigma)$ and the distribution of the random noise $X$. It turns out that the asymptotic distributions exhibit nonuniversality in the sense of dependence on the distributions of the entries of $X$, which contrasts with what has previously been established for the spiked sample eigenvalues in the context of spiked population models. Such a result yields the asymptotic distribution of the sample spiked eigenvalues associated with mixture models. We also explore the application of these findings in detecting mean heterogeneity of data matrices.

math.ST

Asymptotic limits of spiked eigenvalues and eigenvectors of signal-plus-noise matrices with weak signals and heteroskedastic noise

This paper is to study a signal-plus-noise model in high dimensional settings when the dimension and the sample size are comparable. Specifically, we assume that the noise has a general covariance matrix that allows for heteroskedasticity, and that the deterministic signal has the same magnitude as the noise and can have a rank that tends to infinity. We develop the asymptotic limits of the left and right spiked singular vectors of the signal-plusnoise data matrix and the limits of the spiked eigenvalues of the corresponding Gram matrix. As an application, we propose a new criterion to estimate the number of clusters in clustering problems.

math.ST

CLT for random quadratic forms based on sample means and sample covariance matrices

In this paper, we use the dimensional reduction technique to study the central limit theory (CLT) random quadratic forms based on sample means and sample covariance matrices. Specifically, we use a matrix denoted by $U_{p\times q}$, to map $q$-dimensional sample vectors to a $p$ dimensional subspace, where $q\geq p$ or $q\gg p$. Under the condition of $p/n\rightarrow 0$ as $(p,n)\rightarrow \infty$, we obtain the CLT of random quadratic forms for the sample means and sample covariance matrices.

math.ST

A new model for preferential attachment scheme with time-varying parameters

We propose an extension of the preferential attachment scheme by allowing the connecting probability to depend on time t. We estimate the parameters involved in the model by minimizing the expected squared difference between the number of vertices of degree one and its conditional expectation. The asymptotic properties of the estimators are also investigated when the parameters are time-varying by establishing the central limit theorem (CLT) of the number of vertices of degree one. We propose a new statistic to test whether the parameters have change points. We also offer some methods to estimate the number of change points and detect the locations of change points. Simulations are conducted to illustrate the performances of the above results.

stat.ME

Spectral Machine Learning for Pancreatic Mass Imaging Classification

We present a novel spectral machine learning (SML) method in screening for pancreatic mass using CT imaging. Our algorithm is trained with approximately 30,000 images from 250 patients (50 patients with normal pancreas and 200 patients with abnormal pancreas findings) based on public data sources. A test accuracy of 94.6 percents was achieved in the out-of-sample diagnosis classification based on a total of approximately 15,000 images from 113 patients, whereby 26 out of 32 patients with normal pancreas and all 81 patients with abnormal pancreas findings were correctly diagnosed. SML is able to automatically choose fundamental images (on average 5 or 9 images for each patient) in the diagnosis classification and achieve the above mentioned accuracy. The computational time is 75 seconds for diagnosing 113 patients in a laptop with standard CPU running environment. Factors that influenced high performance of a well-designed integration of spectral learning and machine learning included: 1) use of eigenvectors corresponding to several of the largest eigenvalues of sample covariance matrix (spike eigenvectors) to choose input attributes in classification training, taking into account only the fundamental information of the raw images with less noise; 2) removal of irrelevant pixels based on mean-level spectral test to lower the challenges of memory capacity and enhance computational efficiency while maintaining superior classification accuracy; 3) adoption of state-of-the-art machine learning classification, gradient boosting and random forest. Our methodology showcases practical utility and improved accuracy of image diagnosis in pancreatic mass screening in the era of AI.

cs.CV

Large-Dimensional Random Matrix Theory and Its Applications in Deep Learning and Wireless Communications

Large-dimensional random matrix theory, RMT for short, which originates from the research field of quantum physics, has shown tremendous capability in providing deep insights into large dimensional systems. With the fact that we have entered an unprecedented era full of massive amounts of data and large complex systems, RMT is expected to play more important roles in the analysis and design of modern systems. In this paper, we review the key results of RMT and its applications in two emerging fields: wireless communications and deep learning. In wireless communications, we show that RMT can be exploited to design the spectrum sensing algorithms for cognitive radio systems and to perform the design and asymptotic analysis for large communication systems. In deep learning, RMT can be utilized to analyze the Hessian, input-output Jacobian and data covariance matrix of the deep neural networks, thereby to understand and improve the convergence and the learning speed of the neural networks. Finally, we highlight some challenges and opportunities in applying RMT to the practical large dimensional systems.

math.SP

Factor Modelling for Clustering High-dimensional Time Series

We propose a new unsupervised learning method for clustering a large number of time series based on a latent factor structure. Each cluster is characterized by its own cluster-specific factors in addition to some common factors which impact on all the time series concerned. Our setting also offers the flexibility that some time series may not belong to any clusters. The consistency with explicit convergence rates is established for the estimation of the common factors, the cluster-specific factors, the latent clusters. Numerical illustration with both simulated data as well as a real data example is also reported. As a spin-off, the proposed new approach also advances significantly the statistical inference for the factor model of Lam and Yao (2012).

math.ST

Tracy-Widom law for the extreme eigenvalues of large signal-plus-noise matrices

Let $\bY =\bR+\bX$ be an $M\times N$ matrix, where $\bR$ is a rectangular diagonal matrix and $\bX$ consists of $i.i.d.$ entries. This is a signal-plus-noise type model. Its signal matrix could be full rank, which is rarely studied in literature compared with the low rank cases. This paper is to study the extreme eigenvalues of $\bY\bY^*$. We show that under the high dimensional setting ($M/N\rightarrow c\in(0,1]$) and some regularity conditions on $\bR$ the rescaled extreme eigenvalue converges in distribution to Tracy-Widom distribution ($TW_1$).

math.ST