SearcharxivSearch

arXiv subjects

Jianfeng Yao

Publications and source records attributed to Jianfeng Yao.

At least 19 recordsLinked to original sources

Quadratic form of heavy-tailed self-normalized random vector with applications in $α$-heavy Marčenko--Pastur law

Let $\mathbf{x}$ be a random vector with $n$ i.i.d.\ real-valued components in the domain attraction of an $α$-stable law with $α\in(0,2)$, and let $\mathbf{y}=\mathbf{x}/\|\mathbf{x}\|_2$ be the associated self-normalized vector on the unit sphere. For a (possibly random) Hermitian matrix $\mathbf{A}_n=\big(a_{ij}^{(n)}\big)$ independent of $\mathbf{y}$, we study the asymptotic law of the quadratic form $\mathbf{y}^\top \mathbf{A}_n \mathbf{y}$. Building on the sharp separation between diagonal and off-diagonal contributions in this heavy-tailed setting, we show that under a mild assumption on the Frobenius norm of the off-diagonal part of $\mathbf{A}_n$ the limiting law is solely governed by the empirical distribution of the diagonal entries and the index $α$. More precisely, if $n^{-1}\sum_{i=1}^n δ_{a^{(n)}_{ii}}$ converges weakly almost surely to a deterministic $ν$, then $Q_n$ converges in distribution to a non-degenerate law $μ_{ν,α}$ characterized through its Stieltjes transform. The law $μ_{ν,α}$ is shown to be atom-free (provided that $ν$ is non-degenerate) with an explicit density and tractable tail behavior. As an application in random matrix theory, we derive an implicit resolvent-based representation of the $α$-heavy Marčenko--Pastur law $H_{α,γ}$ for heavy-tailed sample correlation matrices and prove that $H_{α,γ}$ has no atoms except possibly at the origin. For comparison with the light-tailed setting, we also provide a Hanson--Wright-type concentration inequality for $\mathbf{y}^\top \mathbf{A}_n \mathbf{y}$ when the components of $\mathbf{x}$ are sub-Gaussian.

math.PR

Deviation Tests for a High-dimensional Mean

This paper investigates testing for deviation of a high-dimensional mean vector $\boldsymbolμ$. In contrast to the standard one-sample significance test of the form: $H_0^\texttt{e} : \boldsymbolμ = \boldsymbolμ_0$ versus $H_1^\texttt{e} : \boldsymbolμ \neq \boldsymbolμ_0$, we focus on testing the deviation $H_0 : \|\boldsymbolμ - \boldsymbolμ_0\|_2 \ge d_0$ versus $H_1 : \|\boldsymbolμ - \boldsymbolμ_0\|_2 < d_0$ for a prespecified length $d_0 > 0$. Constructing a valid test statistic for this problem is technically nontrivial. By applying the concept of positive and negative feedback processes from control theory, we propose a test statistic based on a two-armed bandit (TAB) process. The deviation test is also extended to the two-sample setting. Simulation experiments confirm a good performance of the tests in finite samples. Finally, a real data analysis demonstrates the practical significance of the proposed deviation tests.

stat.ME

Extreme principal minors of Wishart and deformed GOE matrices

We study the laws of large numbers for the largest eigenvalues among all principal minors of Wishart matrices and deformed GOE matrices. We propose a new method based on identifying the deterministic sets to which the random sets formed by suitably normalized principal minors converge in Hausdorff distance, thereby reducing the original extreme-value problems to finite-dimensional convex optimization problems. We demonstrate the effectiveness of this method in regimes not covered by the existing second-moment arguments in \cite{cai2021asymptotic,hu2023extreme}. For deformed GOE matrices with fixed minor size \(k\), we determine the limit for every diagonal variance \(a>0\) and identify a phase transition at \(a=2\). Above the transition, the limiting constant satisfies an explicit recursion with no close-form expression, and the optimizers exhibit a nested hierarchical structure, thereby resolving the case left open in \cite{cai2021asymptotic}. For Wishart matrices with general sub-Gaussian entries and fixed \(k\), we characterize the limit through an entropy-constrained deterministic convex set. When the entries are standard Gaussian, we solve the resulting optimization problem explicitly and obtain the exact value of the limiting constant.

math.PR

Alignment and matching tests for high-dimensional tensor signals via tensor contraction

We consider two hypothesis testing problems for low-rank and high-dimensional tensor signals, namely the tensor signal alignment and tensor signal matching problems. These problems are challenging due to the high dimension of tensors and the lack of suitable test statistics. By exploiting a recent tensor contraction method, we propose and validate relevant test statistics using eigenvalues of a data matrix resulting from the tensor contraction. The matrix entries exhibit long-range dependence, which makes the analysis of the matrix challenging, involved, and distinct from standard random matrix theory. Our approach provides a novel framework for addressing hypothesis testing problems in the context of high-dimensional tensor signals.

stat.ME

Spectral analysis of large dimensional Chatterjee's rank correlation matrix

This paper studies the spectral behavior of large dimensional Chatterjee's rank correlation matrix when observations are independent draws from a high-dimensional random vector with independent continuous components. Limits for the empirical spectral distributions of its two symmetrized versions are established in the proportional high-dimensional regime, one of them being the semicircle law, thereby giving a first example of a correlation matrix with a non-Marchenko--Pastur spectral limit, in contrast to the Pearson, Kendall, and Spearman cases. We further establish central limit theorems for linear spectral statistics of the symmetrized matrices. As an important application of this theory, we develop Chatterjee's rank correlation-based tests for the complete independence among the components.

math.ST

Mean-Shift PCA by Knockoff Mean

Removing noise is difficult, but adding noise is easy. In this work, we show how to eliminate mean-shift noisy components from PCA by deliberately introducing knockoff mean-shift perturbation. Standard PCA is highly sensitive to shifts in the sample mean: a small fraction of samples from a shifted distribution can cause large deviations in the leading principal components. In high-dimensional regimes, existing Robust PCA approaches cannot handle the mean-shift contamination structure inherent in the mixture model. Using tools from Random Matrix Theory, we prove that the mean-shift spikes are spectrally separable from the stable eigenvalues of the original covariance. Furthermore, the original eigenspace remains asymptotically invariant to the contamination, independent of the mixture weight. Exploiting this spectral stability, we propose a simple, two-stage PCA algorithm by adding knockoff mean that identifies and removes the mean-shift component using only standard PCA operations.

stat.ML

Limiting spectral distributions of large consistent rank correlation matrices

We study random matrices whose entries are obtained by applying consistent rank correlations, such as Hoeffding's $D$, pairwise to a high-dimensional random vector with mutually independent components. Prior work has shown that, in the proportional high-dimensional regime, the empirical spectral distributions of large Kendall's tau and Spearman's rho matrices converge weakly almost surely to the Marchenko--Pastur law. By contrast, we prove that for consistent rank correlations such as Hoeffding's $D$, the limiting spectral distribution is given by the semicircle law. Our result thus generalizes a recent work of Dong, Han, and Yao (2025), who considered the special case of Chatterjee's rank correlation and established the first semicircle law for a large correlation matrix in the proportional regime.

math.PR

Many-sample tests for the dimensionality hypothesis for large covariance matrices among groups

In this paper, we consider procedures for testing hypotheses on the dimension of the linear span generated by a growing number of $p\times p$ covariance matrices from independent $q$ populations. Under a proper limiting scheme where all the parameters, $q$, $p$, and the sample sizes from the $q$ populations, are allowed to increase to infinity, we derive the asymptotic normality of the proposed test statistics. The proposed test procedures show satisfactory performance in finite samples under both the null and the alternative. We also apply the proposed many-sample dimensionality test to investigate a matrix-valued gene dataset from the Mouse Aging Project and gain some new knowledge about its covariance structures.

math.ST

Towards Quantifying the Hessian Structure of Neural Networks

Empirical studies reported that the Hessian matrix of neural networks (NNs) exhibits a near-block-diagonal structure, yet its theoretical foundation remains unclear. In this work, we reveal that the reported Hessian structure comes from a mixture of two forces: a ``static force'' rooted in the architecture design, and a ''dynamic force'' arisen from training. We then provide a rigorous theoretical analysis of ''static force'' at random initialization. We study linear models and 1-hidden-layer networks for classification tasks with $C$ classes. By leveraging random matrix theory, we compare the limit distributions of the diagonal and off-diagonal Hessian blocks and find that the block-diagonal structure arises as $C$ becomes large. Our findings reveal that $C$ is one primary driver of the near-block-diagonal structure. These results may shed new light on the Hessian structure of large language models (LLMs), which typically operate with a large $C$ exceeding $10^4$.

cs.LG

A robust and p-hacking-proof significance test under variance uncertainty

P-hacking poses challenges to traditional hypothesis testing. In this paper, we propose a robust method for the one-sample significance test that can protect against p-hacking from sample manipulation. Precisely, assuming a sequential arrival of the data whose variance can be time-varying and for which only lower and upper bounds are assumed to exist with possibly unknown values, we use the modern theory of sublinear expectation to build a testing procedure which is robust under such variance uncertainty, and can protect the significance level against potential data manipulation by an experimenter. It is shown that our new method can effectively control the type I error while preserving a satisfactory power, yet a traditional rejection criterion performs poorly under such variance uncertainty. Our theoretical results are well confirmed by a detailed simulation study.

math.ST

CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation

Multimodal emotion recognition in conversation (MER) aims to accurately identify emotions in conversational utterances by integrating multimodal information. Previous methods usually treat multimodal information as equal quality and employ symmetric architectures to conduct multimodal fusion. However, in reality, the quality of different modalities usually varies considerably, and utilizing a symmetric architecture is difficult to accurately recognize conversational emotions when dealing with uneven modal information. Furthermore, fusing multi-modality information in a single granularity may fail to adequately integrate modal information, exacerbating the inaccuracy in emotion recognition. In this paper, we propose a novel Cross-Modality Augmented Transformer with Hierarchical Variational Distillation, called CMATH, which consists of two major components, i.e., Multimodal Interaction Fusion and Hierarchical Variational Distillation. The former is comprised of two submodules, including Modality Reconstruction and Cross-Modality Augmented Transformer (CMA-Transformer), where Modality Reconstruction focuses on obtaining high-quality compressed representation of each modality, and CMA-Transformer adopts an asymmetric fusion strategy which treats one modality as the central modality and takes others as auxiliary modalities. The latter first designs a variational fusion network to fuse the fine-grained representations learned by CMA- Transformer into a coarse-grained representations. Then, it introduces a hierarchical distillation framework to maintain the consistency between modality representations with different granularities. Experiments on the IEMOCAP and MELD datasets demonstrate that our proposed model outperforms previous state-of-the-art baselines. Implementation codes can be available at https://github.com/ cjw-MER/CMATH.

cs.MM

The First-stage F Test with Many Weak Instruments

A widely adopted approach for detecting weak instruments is to use the first-stage $F$ statistic. While this method was developed with a fixed number of instruments, its performance with many instruments remains insufficiently explored. We show that the first-stage $F$ test exhibits distorted sizes for detecting many weak instruments, regardless of the choice of pretested estimators or Wald tests. These distortions occur due to the inadequate approximation using classical noncentral Chi-squared distributions. As a byproduct of our main result, we present an alternative approach to pre-test many weak instruments with the corrected first-stage $F$ statistic. An empirical illustration with Angrist and Keueger (1991)'s returns to education data confirms its usefulness.

econ.EM

Many-sample tests for the equality and the proportionality hypotheses between large covariance matrices

This paper proposes procedures for testing the equality hypothesis and the proportionality hypothesis involving a large number of $q$ covariance matrices of dimension $p\times p$. Under a limiting scheme where $p$, $q$ and the sample sizes from the $q$ populations grow to infinity in a proper manner, the proposed test statistics are shown to be asymptotically normal. Simulation results show that finite sample properties of the test procedures are satisfactory under both the null and alternatives. As an application, we derive a test procedure for the Kronecker product covariance specification for transposable data. Empirical analysis of datasets from the Mouse Aging Project and the 1000 Genomes Project (phase 3) is also conducted.

math.ST

Robust estimation for number of factors in high dimensional factor modeling via Spearman correlation matrix

Determining the number of factors in high-dimensional factor modeling is essential but challenging, especially when the data are heavy-tailed. In this paper, we introduce a new estimator based on the spectral properties of Spearman sample correlation matrix under the high-dimensional setting, where both dimension and sample size tend to infinity proportionally. Our estimator is robust against heavy tails in either the common factors or idiosyncratic errors. The consistency of our estimator is established under mild conditions. Numerical experiments demonstrate the superiority of our estimator compared to existing methods.

stat.ME

Multiple Descent in the Multiple Random Feature Model

Recent works have demonstrated a double descent phenomenon in over-parameterized learning. Although this phenomenon has been investigated by recent works, it has not been fully understood in theory. In this paper, we investigate the multiple descent phenomenon in a class of multi-component prediction models. We first consider a ''double random feature model'' (DRFM) concatenating two types of random features, and study the excess risk achieved by the DRFM in ridge regression. We calculate the precise limit of the excess risk under the high dimensional framework where the training sample size, the dimension of data, and the dimension of random features tend to infinity proportionally. Based on the calculation, we further theoretically demonstrate that the risk curves of DRFMs can exhibit triple descent. We then provide a thorough experimental study to verify our theory. At last, we extend our study to the ''multiple random feature model'' (MRFM), and show that MRFMs ensembling $K$ types of random features may exhibit $(K+1)$-fold descent. Our analysis points out that risk curves with a specific number of descent generally exist in learning multi-component prediction models.

math.ST

Unified and robust Lagrange multiplier type tests for cross-sectional independence in large panel data models

This paper revisits the Lagrange multiplier type test for the null hypothesis of no cross-sectional dependence in large panel data models. We propose a unified test procedure and its power enhancement version, which show robustness for a wide class of panel model contexts. Specifically, the two procedures are applicable to both heterogeneous and fixed effects panel data models with the presence of weakly exogenous as well as lagged dependent regressors, allowing for a general form of nonnormal error distribution. With the tools from Random Matrix Theory, the asymptotic validity of the test procedures is established under the simultaneous limit scheme where the number of time periods and the number of cross-sectional units go to infinity proportionally. The derived theories are accompanied by detailed Monte Carlo experiments, which confirm the robustness of the two tests and also suggest the validity of the power enhancement technique.

econ.EM

A specification test for the strength of instrumental variables

This paper develops a new specification test for the instrument weakness when the number of instruments $K_n$ is large with a magnitude comparable to the sample size $n$. The test relies on the fact that the difference between the two-stage least squares (2SLS) estimator and the ordinary least squares (OLS) estimator asymptotically disappears when there are many weak instruments, but otherwise converges to a non-zero limit. We establish the limiting distribution of the difference within the above two specifications, and introduce a delete-$d$ Jackknife procedure to consistently estimate the asymptotic variance/covariance of the difference. Monte Carlo experiments demonstrate the good performance of the test procedure for both cases of single and multiple endogenous variables. Additionally, we re-examine the analysis of returns to education data in Angrist and Keueger (1991) using our proposed test. Both the simulation results and empirical analysis indicate the reliability of the test.

econ.EM

Ratio-consistent estimation for long range dependent Toeplitz covariance with application to matrix data whitening

We consider a data matrix $X:=C_N^{1/2}ZR_M^{1/2}$ from a multivariate stationary process with a separable covariance function, where $C_N$ is a $N\times N$ positive semi-definite matrix, $Z$ a $N\times M$ random matrix of uncorrelated standardized white noise, and $R_M$ a $M\times M$ Toeplitz matrix. Under the assumption of long range dependence (LRD), we re-examine the consistency of two toeplitzifized estimators $\hat R_M$ (unbiased) and $\hat R_M^b$ (biased) for $R_M$, which are known to be norm consistent with $R_M$ when the process is short range dependent (SRD). However in the LRD case, some simulations suggest that the norm consistency does not hold in general for both estimators. Instead, a weaker {\it ratio consistency} is established for the unbiased estimator $\hat R_M$, and a further weaker {\it ratio LSD consistency} is established for the biased estimator $\hat R_M^b$. The main result leads to a consistent whitening procedure on the original data matrix $X$, which is further applied to two real world questions, one is a signal detection problem, and the other is PCA on the space covariance $C_N$ to achieve a noise reduction and data compression.

math.PR