SearcharxivSearch

arXiv subjects

Yukun He

Publications and source records attributed to Yukun He.

18 recordsLinked to original sources

The oriented Kesten--McKay law for random regular digraphs

We consider the adjacency matrix of a random directed $d$-regular graph on $N$ vertices. For fixed $d\geq 2$, we prove that the empirical eigenvalue density converges in probability to the oriented Kesten--McKay law as $N\to \infty$. The key technical input is the small-ball probability estimate for the smallest singular value. The proof combines a fixed-rank transposition argument with finite-field anticoncentration for shifted inverse compressions. We also prove a polynomial hard-edge estimate, which allows us to deduce the global law from the vanishing small-ball probability.

math.PR

Free energy and spectral edge of the SYK model

We introduce a microscopic framework that gives the first rigorous derivation of the Schwinger--Dyson thermodynamics of the Sachdev--Ye--Kitaev model. For fixed, even interaction order $q\geq 4$, we prove that the annealed and quenched normalized pressures converge to the Schwinger--Dyson pressure $p_{{\rm SD},q}(\beta)$, at any fixed inverse temperature $\beta>0$. The proof is derived from the finite-$N$ Gibbs state, and it contains three new ingredients: a single-site cavity expansion that keeps the bulk Gibbs state intact, a finite-dimensional locality estimate yielding label-uniform conditional factorization of the Euclidean cavity fields, and an exact Majorana-bath representation of the leading diagrams. We further determine the asymptotic location of the largest eigenvalue, which answers a question posed by Feng--Tian--Wei. Consequently, the sample free-energy density converges almost surely to the zero-temperature value along every sequence $\beta\equiv \beta_N\to\infty$.

math-ph

Gaussian Waves and Edge Eigenvectors of Random Regular Graphs

Backhausz and Szegedy (2019) demonstrated that the almost eigenvectors of random regular graphs converge to Gaussian waves with variance $0\leq \sigma^2\leq 1$. In this paper, we present an alternative proof of this result for the edge eigenvectors of random regular graphs, establishing that the variance must be $\sigma^2=1$. Furthermore, we show that the eigenvalues and eigenvectors are asymptotically independent. Our approach introduces a simple framework linking the weak convergence of the imaginary part of the Green's function to the convergence of eigenvectors, which may be of independent interest.

math.PR

Extremal eigenvectors of sparse random matrices

We consider a class of sparse random matrices, which includes the adjacency matrix of Erd\H{o}s-R\'enyi graph ${\bf G}(N,p)$. For $N^{-1+o(1)}\leq p\leq 1/2$, we show that the non-trivial edge eigenvectors are asymptotically jointly normal. The main ingredient of the proof is an algorithm that directly computes the joint eigenvector distributions, without comparisons with GOE. The method is applicable in general. As an illustration, we also use it to prove the normal fluctuation in quantum ergodicity at the edge for Wigner matrices. Another ingredient of the proof is the isotropic local law for sparse matrices, which at the same time improves several existing results.

math.PR

Edge universality of sparse Erd\H{o}s-R\'enyi digraphs

Let $\mathcal A$ be the adjacency matrix of the Erd\H{o}s-R\'{e}nyi directed graph $\mathscr G(N,p)$. We denote the eigenvalues of $\mathcal A$ by $\lambda_1^{\cal A},...,\lambda^{\cal A}_N$, and $|\lambda_1^{\cal A}|=\max_i|\lambda_i^{\cal A}|$. For $N^{-1+o(1)}\leq p\leq 1/2$, we show that \[ \max_{i=2,3,...,N} \bigg|\frac{\lambda_i^{\mathcal A}}{\sqrt{Np(1-p)}}\bigg| =1+O(N^{-1/2+o(1)}) \] with very high probability. In addition, we prove that near the unit circle, the local eigenvalue statistics of ${\mathcal A}/\sqrt{Np(1-p)}$ coincide with those of the real Ginibre ensemble. As a by-product, we also show that all non-trivial eigenvectors of $\mathcal A$ are completely delocalized. For Hermitian random matrices, it is known that the edge statistics are sensitive to the sparsity: in the very sparse regime, one needs to remove many noise random variables (which affect both the mean and the fluctuation) to recover the Tracy-Widom distribution. Our results imply that, compared to their analogues in the Hermitian case, the edge statistics of non-Hermitian sparse random matrices are more robust.

math.PR

Spectral gap and edge universality of dense random regular graphs

Let $\mathcal A$ be the adjacency matrix of a random $d$-regular graph on $N$ vertices, and we denote its eigenvalues by $\lambda_1\geq \lambda_2\cdots \geq \lambda_{N}$. For $N^{2/3}\ll d\leq N/2$, we prove optimal rigidity estimates of the extreme eigenvalues of $\mathcal A$, which in particular imply that \[ \max\{|\lambda_N|,\lambda_2\} <2\sqrt{d-1} \] with overwhelming probability. In the same regime of $d$, we also show that \[ N^{2/3}\bigg(\frac{\lambda_2+d/N}{\sqrt{d(N-d)/N}}-2\bigg) \overset{d}{\longrightarrow} \mathrm{TW}_1\,, \]where $\mathrm{TW}_1$ is the Tracy-Widom distribution for GOE; analogues results also hold for other non-trivial extreme eigenvalues.

math.PR

Towards IID representation learning and its application on biomedical data

Due to the heterogeneity of real-world data, the widely accepted independent and identically distributed (IID) assumption has been criticized in recent studies on causality. In this paper, we argue that instead of being a questionable assumption, IID is a fundamental task-relevant property that needs to be learned. Consider $k$ independent random vectors $\mathsf{X}^{i = 1, \ldots, k}$, we elaborate on how a variety of different causal questions can be reformulated to learning a task-relevant function $ϕ$ that induces IID among $\mathsf{Z}^i := ϕ\circ \mathsf{X}^i$, which we term IID representation learning. For proof of concept, we examine the IID representation learning on Out-of-Distribution (OOD) generalization tasks. Concretely, by utilizing the representation obtained via the learned function that induces IID, we conduct prediction of molecular characteristics (molecular prediction) on two biomedical datasets with real-world distribution shifts introduced by a) preanalytical variation and b) sampling protocol. To enable reproducibility and for comparison to the state-of-the-art (SOTA) methods, this is done by following the OOD benchmarking guidelines recommended from WILDS. Compared to the SOTA baselines supported in WILDS, the results confirm the superior performance of IID representation learning on OOD tasks. The code is publicly accessible via https://github.com/CTPLab/IID_representation_learning.

cs.LG

Fluctuations of extreme eigenvalues of sparse Erdős-Rényi graphs

We consider a class of sparse random matrices which includes the adjacency matrix of the Erdős-Rényi graph $\mathcal{G}(N,p)$. We show that if $N^{\varepsilon} \leq Np \leq N^{1/3-\varepsilon}$ then all nontrivial eigenvalues away from 0 have asymptotically Gaussian fluctuations. These fluctuations are governed by a single random variable, which has the interpretation of the total degree of the graph. This extends the result [19] on the fluctuations of the extreme eigenvalues from $Np \geq N^{2/9 + \varepsilon}$ down to the optimal scale $Np \geq N^{\varepsilon}$. The main technical achievement of our proof is a rigidity bound of accuracy $N^{-1/2-\varepsilon} \, (Np)^{-1/2}$ for the extreme eigenvalues, which avoids the $(Np)^{-1}$-expansions from [9,19,24]. Our result is the last missing piece, added to [8, 12, 19, 24], of a complete description of the eigenvalue fluctuations of sparse random matrices for $Np \geq N^{\varepsilon}$.

math.PR

Quantitative CLT for linear eigenvalue statistics of Wigner matrices

In this article, we establish a near-optimal convergence rate for the CLT of linear eigenvalue statistics of Wigner matrices, in Kolmogorov-Smirnov distance. For all test functions $f\in C^5(\mathbb R)$, we show that the convergence rate is either $N^{-1/2+\varepsilon}$ or $N^{-1+\varepsilon}$, depending on the first Chebyshev coefficient of $f$ and the third moment of the diagonal matrix entries. The condition that distinguishes these two rates is necessary and sufficient. For a general class of test functions, we further identify matching lower bounds for the convergence rates. In addition, we identify an explicit, non-universal contribution in the linear eigenvalue statistics, which is responsible for the slow rate $N^{-1/2+\varepsilon}$ for non-Gaussian ensembles. By removing this non-universal part, we show that the shifted linear eigenvalue statistics have the unified convergence rate $N^{-1+\varepsilon}$ for all test functions.

math.PR

On Cramér-von Mises statistic for the spectral distribution of random matrices

Let $F_N$ and $F$ be the empirical and limiting spectral distributions of an $N\times N$ Wigner matrix. The Cramér-von Mises (CvM) statistic is a classical goodness-of-fit statistic that characterizes the distance between $F_N$ and $F$ in $\ell^2$-norm. In this paper, we consider a mesoscopic approximation of the CvM statistic for Wigner matrices, and derive its limiting distribution. In the appendix, we also give the limiting distribution of the CvM statistic (without approximation) for the toy model CUE.

math.PR

Bulk eigenvalue fluctuations of sparse random matrices

We consider a class of sparse random matrices, which includes the adjacency matrix of Erdős-Rényi graphs $\mathcal G(N,p)$ for $p \in [N^{\varepsilon-1},N^{-\varepsilon}]$. We identify the joint limiting distributions of the eigenvalues away from 0 and the spectral edges. Our result indicates that unlike Wigner matrices, the eigenvalues of sparse matrices satisfy central limit theorems with normalization $N\sqrt{p}$. In addition, the eigenvalues fluctuate simultaneously: the correlation of two eigenvalues of the same/different sign is asymptotically 1/-1. We also prove CLTs for the eigenvalue counting function and trace of the resolvent at mesoscopic scales.

math.PR

Mesoscopic eigenvalue density correlations of Wigner matrices

We investigate to what extent the microscopic Wigner-Gaudin-Mehta-Dyson (WGMD) (or sine kernel) statistics of random matrix theory remain valid on mesoscopic scales. To that end, we compute the connected two-point spectral correlation function of a Wigner matrix at two mesoscopically separated points. In the mesoscopic regime, density correlations are much weaker than in the microscopic regime. Our result is an explicit formula for the two-point function. This formula implies that the WGMD statistics are valid to leading order on all mesoscopic scales, that in the real symmetric case there are subleading corrections matching precisely the WGMD statistics, while in the complex Hermitian case these subleading corrections are absent. We also uncover non-universal subleading correlations, which dominate over the universal ones beyond a certain intermediate mesoscopic scale. The proof is based on a hierarchy of Schwinger-Dyson equations for a sufficiently large class of polynomials in the entries of the Green function. The hierarchy is indexed by a tree, whose depth is controlled using stopping rules. A key ingredient in the derivation of the stopping rules is a new estimate on the density of states, which we prove to have bounded derivatives of all order on all mesoscopic scales.

math.PR

Diffusion Profile for Random Band Matrices: a Short Proof

Let $H$ be a Hermitian random matrix whose entries $H_{xy}$ are independent, centred random variables with variances $S_{xy} = \mathbb E|H_{xy}|^2$, where $x, y \in (\mathbb Z/L\mathbb Z)^d$ and $d \geq 1$. The variance $S_{xy}$ is negligible if $|x - y|$ is bigger than the band width $W$. For $ d = 1$ we prove that if $L \ll W^{1 + \frac{2}{7}}$ then the eigenvectors of $H$ are delocalized and that an averaged version of $|G_{xy}(z)|^2$ exhibits a diffusive behaviour, where $ G(z) = (H-z)^{-1}$ is the resolvent of $ H$. This improves the previous assumption $L \ll W^{1 + \frac{1}{4}}$ by Erdős et al. (2013). In higher dimensions $d \geq 2$, we obtain similar results that improve the corresponding by Erdős et al. Our results hold for general variance profiles $S_{xy}$ and distributions of the entries $H_{xy}$. The proof is considerably simpler and shorter than that by Erdős et al. It relies on a detailed Fourier space analysis combined with isotropic estimates for the fluctuating error terms. It avoids the intricate fluctuation averaging machinery used by Erdős and collaborators.

math-ph

Mesoscopic linear statistics of Wigner matrices of mixed symmetry class

We prove a central limit theorem for the mesoscopic linear statistics of $N\times N$ Wigner matrices $H$ satisfying $\mathbb{E}|H_{ij}|^2=1/N$ and $\mathbb{E} H_{ij}^2= σ/N$, where $σ\in [-1,1]$. We show that on all mesoscopic scales $η$ ($1/N \ll η\ll 1$), the linear statistics of $H$ have a sharp transition at $1-σ\sim η$. As an application, we identify the mesoscopic linear statistics of Dyson's Brownian motion $H_t$ started from a real symmetric Wigner matrix $H_0$ at any nonnegative time $t \in [0,\infty]$. In particular, we obtain the transition from the central limit theorem for GOE to the one for GUE at time $t \sim η$.

math.PR

Local law and complete eigenvector delocalization for supercritical Erdős-Rényi graphs

We prove a local law for the adjacency matrix of the Erdős-Rényi graph $G(N, p)$ in the supercritical regime $ pN \geq C\log N$ where $G(N,p)$ has with high probability no isolated vertices. In the same regime, we also prove the complete delocalization of the eigenvectors. Both results are false in the complementary subcritical regime. Our result improves the corresponding results from [11] by extending them all the way down to the critical scale $pN = O(\log N)$. A key ingredient of our proof is a new family of multilinear large deviation estimates for sparse random vectors, which carefully balance mixed $\ell^2$ and $\ell^\infty$ norms of the coefficients with combinatorial factors, allowing us to prove strong enough concentration down to the critical scale $pN = O(\log N)$. These estimates are of independent interest and we expect them to be more generally useful in the analysis of very sparse random matrices.

math.PR

Efficient Two-Step Adversarial Defense for Deep Neural Networks

In recent years, deep neural networks have demonstrated outstanding performance in many machine learning tasks. However, researchers have discovered that these state-of-the-art models are vulnerable to adversarial examples: legitimate examples added by small perturbations which are unnoticeable to human eyes. Adversarial training, which augments the training data with adversarial examples during the training process, is a well known defense to improve the robustness of the model against adversarial attacks. However, this robustness is only effective to the same attack method used for adversarial training. Madry et al.(2017) suggest that effectiveness of iterative multi-step adversarial attacks and particularly that projected gradient descent (PGD) may be considered the universal first order adversary and applying the adversarial training with PGD implies resistance against many other first order attacks. However, the computational cost of the adversarial training with PGD and other multi-step adversarial examples is much higher than that of the adversarial training with other simpler attack techniques. In this paper, we show how strong adversarial examples can be generated only at a cost similar to that of two runs of the fast gradient sign method (FGSM), allowing defense against adversarial attacks with a robustness level comparable to that of the adversarial training with multi-step adversarial examples. We empirically demonstrate the effectiveness of the proposed two-step defense approach against different attack methods and its improvements over existing defense strategies.

cs.LG

Isotropic self-consistent equations for mean-field random matrices

We present a simple and versatile method for deriving (an)isotropic local laws for general random matrices constructed from independent random variables. Our method is applicable to mean-field random matrices, where all independent variables have comparable variances. It is entirely insensitive to the expectation of the matrix. In this paper we focus on the probabilistic part of the proof -- the derivation of the self-consistent equations. As a concrete application, we settle in complete generality the local law for Wigner matrices with arbitrary expectation.

math.PR

Mesoscopic eigenvalue statistics of Wigner matrices

We prove that the linear statistics of the eigenvalues of a Wigner matrix converge to a universal Gaussian process on all mesoscopic spectral scales, i.e. scales larger than the typical eigenvalue spacing and smaller than the global extent of the spectrum.

math.PR