SearcharxivSearch

arXiv subjects

Lucas Benigni

Publications and source records attributed to Lucas Benigni.

18 recordsLinked to original sources

Local law and delocalization for the Sachdev-Ye-Kitaev model

We establish a mesoscopic local law and quantitative eigenvector-delocalization estimates for the Sachdev--Ye--Kitaev model (SYK) Hamiltonian of $N$ interacting Majorana fermions. For even $q\ll N^{\frac{1}{2}}$, we prove that the normalized Stieltjes transform converges uniformly on bounded energy intervals to that of the standard Gaussian law, down to scales of order $qN^{-\frac{1}{2}}$, up to logarithmic factors. The result holds both on the full Hilbert space and in each fermion-parity sector. We derive mesoscopic eigenvalue counting and spectral form factor estimates, as well as an averaged inverse participation ratio bound. For fixed $q$, we further obtain high-probability $\ell^\infty$-delocalization bounds in any deterministic orthonormal basis for individual bulk eigenvectors. These are the first mesoscopic laws and eigenvector delocalization estimates for the SYK Hamiltonian, which we obtain using the resolvent method.

math.PR

Quantitative eigenvector universality for generalized Wigner matrices

We present a novel approach to eigenvector universality for generalized Wigner matrices. Our main consequences are asymptotic normality of joint eigenvector projections everywhere in the spectrum as well as a quantitative lower bound on the largest entry of an eigenvector. In the case of smooth entries, we are able to obtain joint normality of an explicit growing number of eigenvector projections, and we are also able to obtain an explicit rate of convergence in Kolmogorov distance. This is based on a new analysis of the Dyson vector flow which does not rely on the eigenvector moment flow.

math.PR

Phases of Muon: When Muon Eclipses SignSGD

Recently, Muon and related spectral optimizers have demonstrated strong empirical performance as scalable stochastic methods, often outperforming Adam. Yet their behaviour remains poorly understood. We analyze stochastic spectral optimizers, including Muon, on a high-dimensional matrix-valued least squares problem. We derive explicit deterministic dynamics that provide a tractable framework for studying learning behaviour with a focus on (stochastic) SignSVD, which Muon approximates, and (stochastic) SignSGD, the latter serving as a proxy for Adam. Our analysis shows that for large batch size, SignSVD performs a square-root preconditioning with respect to the data covariance spectrum, while for small batch size smaller eigenmodes behave like SGD, slowing down convergence. We contrast with SignSGD which for generic covariance performs no preconditioning and has no transition, leading to different optimal learning rates and convergence characteristics. The two methods match up to a constant factor with isotropic data, but behave differently with anisotropic data. An analysis of a power law covariance model with data exponent $\alpha$ and target exponent $\beta$ shows there are three phases in the $(\alpha,\beta)$ plane: one where SignSGD is uniformly favored, one where SignSVD is uniformly favored, and a third where the two methods exhibit a trade-off in performance.

math.OC

The delocalization of eigenvectors of real elliptic matrices

We investigate delocalization phenomena for eigenvectors of real random matrices that are invariant by orthogonal transformations. A specific phenomenon with these ensembles is that an eigenvector is typically more localized when its eigenvalue is closer to the real axis while for unitarily invariant ensembles, all eigenvectors are delocalized at the same level. More precisely, we measure the delocalization level of a vector $x\in \mathbb{C}^N$ using the Inverse Participation Ratio $\mathrm{IPR}(x) = N|x|_4^4 / |x|_2^4 \geqslant 1$. A higher IPR means a more localized vector. Using the exact distribution of the Schur decomposition of some paradigmatic rotation-invariant matrix models, we prove that conditionally on having an eigenvalue $\lambda$ with $|\mathfrak{Im}(\lambda)| = y / \sqrt{N}$, the IPR of the associated eigenvector converges in distribution towards a random variable $\ell_y$ with an explicit density depending only on $y$. We then prove that $\ell_y \to 3$ when $y \to 0$ and $\ell_y \to 2$ when $y\to +\infty$, coherently with the observed phenomenon. This result is explicitly proved for higher-order IPRs and for the real Elliptic Ginibre ensemble at every non-symmetry parameter $\tau \in [0,1[$, including the classical real Ginibre ensemble ($\tau=0$).

math.PR

Convergence of local eigenvector processes of generalized Wigner matrices

We prove convergence of eigenvector processes of the form $(\sqrt{N}\langle \mathbf{u}_k,A_t\mathbf{u}_k\rangle)_{t\in[0,1]}$ where $\mathbf{u}_k$ is a bulk eigenvector of generalized Wigner matrices and $(A_t)$ a family of symmetric matrices with bounded norm and H\"{o}lder regularity. We give explicit examples of limiting processes and prove that a large class of Gaussian process with H\"{o}lder-continuous covariance function can be obtained as such a limit using its Karhunen--Lo\`eve expansion. The proof is based on the multi-dimensional convergence proved Benigni and Cipolloni (2024) and a tightness criterion proved using H\"{o}lder regularity of the observables.

math.PR

Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling

We compute the asymptotic eigenvalue distribution of the neural tangent kernel of a two-layer neural network under a specific scaling of dimension. Namely, if $X\in\mathbb{R}^{n\times d}$ is an i.i.d random matrix, $W\in\mathbb{R}^{d\times p}$ is an i.i.d $\mathcal{N}(0,1)$ matrix and $D\in\mathbb{R}^{p\times p}$ is a diagonal matrix with i.i.d bounded entries, we consider the matrix \[ \mathrm{NTK} = \frac{1}{d}XX^\top \odot \frac{1}{p} \sigma'\left( \frac{1}{\sqrt{d}}XW \right)D^2 \sigma'\left( \frac{1}{\sqrt{d}}XW \right)^\top \] where $\sigma'$ is a pseudo-Lipschitz function applied entrywise and under the scaling $\frac{n}{dp}\to \gamma_1$ and $\frac{p}{d}\to \gamma_2$. We describe the asymptotic distribution as the free multiplicative convolution of the Marchenko--Pastur distribution with a deterministic distribution depending on $\sigma$ and $D$.

math.PR

Eigenvalue distribution of the Hadamard product of sample covariance matrices in a quadratic regime

In this note, we prove that if $X\in\mathbb{R}^{n\times d}$ and $Y\in\mathbb{R}^{n\times p}$ are two independent matrices with i.i.d entries then the empirical spectral distribution of $\frac{1}{d}XX^\top \odot \frac{1}{p}YY^\top$, where $\odot$ denotes the Hadamard product, converges to the Marchenko--Pastur distribution of shape $\gamma$ in the quadratic regime of dimension $\frac{n}{dp}\to \gamma$ and $\frac{p}{d}\to a$.

math.PR

Fluctuations in local quantum unique ergodicity for generalized Wigner matrices

We study the eigenvector mass distribution for generalized Wigner matrices on a set of coordinates $I$, where $N^\varepsilon \le | I | \le N^{1- \varepsilon}$, and prove it converges to a Gaussian at every energy level, including the edge, as $N\rightarrow \infty$. The key technical input is a four-point decorrelation estimate for eigenvectors of matrices with a large Gaussian component. Its proof is an application of the maximum principle to a new set of moment observables satisfying parabolic evolution equations. Additionally, we prove high-probability Quantum Unique Ergodicity and Quantum Weak Mixing bounds for all eigenvectors and all deterministic sets of entries using a novel bootstrap argument.

math.PR

Fluctuations in Quantum Unique Ergodicity at the Spectral Edge

We study the eigenvector mass distribution of an $N\times N$ Wigner matrix on a set of coordinates $I$ satisfying $| I | \ge c N$ for some constant $c >0$. For eigenvectors corresponding to eigenvalues at the spectral edge, we show that the sum of the mass on these coordinates converges to a Gaussian in the $N \rightarrow \infty$ limit, after a suitable rescaling and centering. The proof proceeds by a two moment matching argument. We directly compare edge eigenvector observables of an arbitrary Wigner matrix to those of a Gaussian matrix, which may be computed explicitly.

math.PR

Largest Eigenvalues of the Conjugate Kernel of Single-Layered Neural Networks

This paper is concerned with the asymptotic distribution of the largest eigenvalues for some nonlinear random matrix ensemble stemming from the study of neural networks. More precisely we consider $M= \frac{1}{m} YY^\top$ with $Y=f(WX)$ where $W$ and $X$ are random rectangular matrices with i.i.d. centered entries. This models the data covariance matrix or the Conjugate Kernel of a single layered random Feed-Forward Neural Network. The function $f$ is applied entrywise and can be seen as the activation function of the neural network. We show that the largest eigenvalue has the same limit (in probability) as that of some well-known linear random matrix ensembles. In particular, we relate the asymptotic limit of the largest eigenvalue for the nonlinear model to that of an information-plus-noise random matrix, establishing a possible phase transition depending on the function $f$ and the distribution of $W$ and $X$. This may be of interest for applications to machine learning.

math.PR

Determinantal structures for Bessel fields

A Bessel field $\mathcal{B}=\{\mathcal{B}(α,t), α\in\mathbb{N}_0, t\in\mathbb{R}\}$ is a two-variable random field such that for every $(α,t)$, $\mathcal{B}(α,t)$ has the law of a Bessel point process with index $α$. The Bessel fields arise as hard edge scaling limits of the Laguerre field, a natural extension of the classical Laguerre unitary ensemble. It is recently proved in [LW21] that for fixed $α$, $\{\mathcal{B}(α,t), t\in\mathbb{R}\}$ is a squared Bessel Gibbsian line ensemble. In this paper, we discover rich integrable structures for the Bessel fields: along a time-like or a space-like path, $\mathcal{B}$ is a determinantal point process with an explicit correlation kernel; for fixed $t$, $\{\mathcal{B}(α,t),α\in\mathbb{N}_0\}$ is an exponential Gibbsian line ensemble.

math.PR

Optimal Delocalization for Generalized Wigner Matrices

We study the eigenvectors of generalized Wigner matrices with subexponential entries and prove that they delocalize at the optimal rate with overwhelming probability. We also prove high probability delocalization bounds with sharp constants. Our proof uses an analysis of the eigenvector moment flow introduced by Bourgade and Yau (2017) to bound logarithmic moments of eigenvector entries for random matrices with small Gaussian components. We then extend this control to all generalized Wigner matrices by comparison arguments based on a framework of regularized eigenvectors, level repulsion, and the observable employed by Landon, Lopatto, and Marcinek (2018) to compare extremal eigenvalue statistics. Additionally, we prove level repulsion and eigenvalue overcrowding estimates for the entire spectrum, which may be of independent interest.

math.PR

Fermionic eigenvector moment flow

We exhibit new functions of the eigenvectors of the Dyson Brownian motion which follow an equation similar to the Bourgade-Yau eigenvector moment flow. These observables can be seen as a Fermionic counterpart to the original (Bosonic) ones. By analyzing both Fermionic and Bosonic observables, we obtain new correlations between eigenvectors. The fluctuations $\sum_{α\in I}u_k(α)^2-{\vert I\vert}/{N}$ decorrelate for distinct eigenvectors as the dimension $N$ grows and an optimal estimate on the partial inner product $\sum_{α\in I}u_k(α)u_\ell(α)$ between two eigenvectors is given. These static results obtained by integrable dynamics are stated for generalized Wigner matrices and should apply to wide classes of mean field models.

math.PR

Eigenvectors distribution and quantum unique ergodicity for deformed Wigner matrices

We analyze the distribution of eigenvectors for mesoscopic, mean-field perturbations of diagonal matrices in the bulk of the spectrum. Our results apply to a generalized $N\times N$ Rosenzweig-Porter model. We prove that the eigenvectors entries are asymptotically Gaussian with a specific variance, localizing them onto a small, explicit part of the spectrum. For a well spread initial spectrum, this variance profile universally follows a heavy-tailed Cauchy distribution. In the case of smooth entries, we also obtain a strong form of quantum unique ergodicity as an overwhelming probability bound on the eigenvectors probability mass. The proof relies on a priori local laws for this model and the eigenvector moment flow.

math.PR

Eigenvalue distribution of nonlinear models of random matrices

This paper is concerned with the asymptotic empirical eigenvalue distribution of a non linear random matrix ensemble. More precisely we consider $M= \frac{1}{m} YY^*$ with $Y=f(WX)$ where $W$ and $X$ are random rectangular matrices with i.i.d. centered entries. The function $f$ is applied pointwise and can be seen as an activation function in (random) neural networks. We compute the asymptotic empirical distribution of this ensemble in the case where $W$ and $X$ have sub-Gaussian tails and $f$ is real analytic. This extends a previous result where the case of Gaussian matrices $W$ and $X$ is considered. We also investigate the same questions in the multi-layer case, regarding neural network applications.

math.PR