Searcharxiv⌕ Search

arXiv subjects

Jiaoyang Huang

Publications and source records attributed to Jiaoyang Huang.

At least 19 recordsLinked to original sources

Fast Convergence for High-Order ODE Solvers in Diffusion Probabilistic Models

Diffusion probabilistic models generate samples by learning to reverse a noise-injection process that transforms data into noise. A key development is the reformulation of the reverse sampling process as a deterministic probability flow ordinary differential equation (ODE), which allows for efficient sampling using high-order numerical solvers. Unlike traditional time integrator analysis, the accuracy of this sampling procedure depends not only on numerical integration errors but also on the approximation quality and regularity of the learned score function, as well as their interaction. In this work, we present a rigorous convergence analysis of deterministic samplers derived from probability flow ODEs for general forward processes with arbitrary variance schedules. Specifically, we develop and analyze $p$-th order (exponential) Runge-Kutta schemes, under the practical assumption that the first and second derivatives of the learned score function are bounded. We prove that the total variation distance between the generated and target distributions can be bounded as \begin{align*} O\bigl(d^{\frac{7}{4}}\varepsilon_{\text{score}}^{\frac{1}{2}} +d(dH_{\max})^p\bigr), \end{align*} where $\varepsilon^2_{\text{score}}$ denotes the $L^2$ error in the score function approximation, $d$ is the data dimension, and $H_{\max}$ represents the maximum solver step size. Numerical experiments on benchmark datasets further confirm that the derivatives of the learned score function are bounded in practice.

cs.LG↗

Loop Equations Characterize Random Matrix Statistics

We prove that the universal local point processes of random matrix theory are characterized by their loop equation hierarchies. More precisely, for every rational $β>0$, the $\mathrm{Sine}_β$ point process is the unique solution of the bulk loop equation hierarchy, and the $\mathrm{Airy}_β$ point process is the unique solution of the edge loop equation hierarchy. These uniqueness results provide a direct route to universality: it suffices to verify the corresponding approximate loop equations for the ensemble. In many models, these equations follow from local laws and integration by parts.

math.PR↗

Height fluctuation for Lozenge Tilings of Polygons

We establish Gaussian free field fluctuations for uniformly random lozenge tilings of simply connected polygonal domains with $3d$ sides whose directions cycle through the three lattice directions. More precisely, assuming that the liquid region is connected and that the boundary data do not force the height at any interior point, we prove that the fluctuations of centered height function converge to the Gaussian free field in the liquid region, confirming a prediction of Kenyon and Okounkov from 2007. We introduce a tiling action function that encodes the geometry of the limit shape through its critical points. The action function has a complex conjugate pair of critical points in the liquid region, repeated real critical points on the arctic boundary, and distinct real critical points in the frozen region. Using this tiling action function, we construct an approximation to the inverse Kasteleyn matrix in terms of explicit single-contour and double-contour integrals and prove that the approximation is uniform throughout the polygonal domain. The convergence to the Gaussian free field then follows from standard kernel computations.

math.PR↗

The oriented Kesten--McKay law for random regular digraphs

We consider the adjacency matrix of a random directed $d$-regular graph on $N$ vertices. For fixed $d\geq 2$, we prove that the empirical eigenvalue density converges in probability to the oriented Kesten--McKay law as $N\to \infty$. The key technical input is the small-ball probability estimate for the smallest singular value. The proof combines a fixed-rank transposition argument with finite-field anticoncentration for shifted inverse compressions. We also prove a polynomial hard-edge estimate, which allows us to deduce the global law from the vanishing small-ball probability.

math.PR↗

Support of Dyson Brownian Motion

We consider beta-Dyson Brownian motion, with $β>= 1$, started from a deterministic configuration with uniformly bounded support. Let $μ_t$ be the semicircular free-convolution flow issued from the initial empirical measure, and set $S_t = supp(μ_t)$. For every fixed $T, ε> 0$, with probability at least $1 - C \exp(-(log n)^2)$, every particle remains within an epsilon-neighborhood of $S_t$ for all $0 <= t <= T$. The result holds from time zero, requires no regularity assumption at the initial spectral edges, and applies to multi-cut supports with macroscopic interior gaps. A key ingredient is a deterministic local resolvent exclusion principle: an $o((n η)^(-1))$ comparison of Stieltjes transforms on a complex disc above a real point separated from the reference support excludes eigenvalues from the corresponding real interval. This gives a model-independent mechanism for converting local resolvent estimates into spectral confinement.

math.PR↗

Non-symmetric vector dyson equations

We study the vector Dyson equation $$-\frac{1}{m(z)}=z\mathbf{1}+\mathbf{a}+Sm(z),$$ with parameter $z$ in the complex upper half-plane $\mathbb{C}_+$, where $\mathbf{a}\in\mathbb R^d$ and $S$ is a nonnegative matrix, not necessarily symmetric. This equation has a unique vector solution $m(z)\in\mathbb{C}_+^d$, for which we establish a complete measure decomposition and prove regularity. We then develop a graph-theoretic approach to the singularity and stability problem for non-symmetric matrices $S$. The graph structure of $S$ identifies the possible degeneracies of the stability operator as $z$ approaches the real axis. In particular, for non-backtracking matrices, we prove square-root growth at regular edges, cubic-root growth at regular cusps, and complete stability estimates. We also obtain the corresponding estimates for symmetric matrices in the periodic setting.

math.CV↗

A convergence framework for Airy$_β$ line ensemble via pole evolution

The Airy$_β$ line ensemble is an infinite sequence of random curves. It is a natural extension of the Tracy-Widom$_β$ distributions, and is expected to be the universal edge scaling limit of a range of models in random matrix theory and statistical mechanics. In this work, we provide a framework of proving convergence to the Airy$_β$ line ensemble, via a characterization through the pole evolution of meromorphic functions satisfying certain stochastic differential equations. Our framework is then applied to prove the universality of the Airy$_β$ line ensemble as the edge limit of various continuous time processes, including Dyson Brownian motions with general $β$ and potentials, Laguerre processes and Jacobi processes.

math.PR↗

Fisher-Rao Gradient Flow: Geodesic Convexity and Functional Inequalities

The dynamics of probability density functions have been extensively studied in computational science and engineering to understand physical phenomena and facilitate algorithmic design. Of particular interest are dynamics formulated as gradient flows of energy functionals under the Wasserstein metric. The development of functional inequalities, such as the log-Sobolev inequality, plays a pivotal role in analyzing the convergence of these dynamics. This paper aims to extend the success of functional inequality techniques to dynamics that are gradient flows under the Fisher-Rao metric, with various $f$-divergences serving as energy functionals. Such dynamics take the form of nonlocal differential equations, for which existing analyses critically rely on explicit solution formulas in special cases. We provide a comprehensive study of functional inequalities and the relevant geodesic convexity for Fisher-Rao gradient flows under minimal assumptions. A notable feature of our functional inequalities is their independence from the log-concavity or log-Sobolev constants of the target distribution. Consequently, the convergence rate of the dynamics (assuming well-posedness) remains uniform across general target distributions.

math.AP↗

Fluctuations for non-Hermitian dynamics

We prove that under the Brownian evolution on large non-Hermitian matrices the log-determinant converges in distribution to a 2+1 dimensional Gaussian field in the Edwards-Wilkinson regularity class, namely it is logarithmically correlated for the parabolic distance. This dynamically extends a seminal result by Rider and Virág about convergence to the Gaussian free field. The convergence holds out of equilibrium for centered, i.i.d. matrix entries as an initial condition. A remarkable aspect of the limiting field is its non-Markovianity, due to long range correlations of the eigenvector overlaps, for which we identify the exact space-time polynomial decay. In the proof, we obtain a quantitative, optimal relaxation at the hard edge, for a broad extension of the Dyson Brownian motion, with a driving noise arbitrarily correlated in space.

math.PR↗

Extremal eigenvectors of sparse random matrices

We consider a class of sparse random matrices, which includes the adjacency matrix of Erdős-Rényi graph ${\bf G}(N,p)$. For $N^{-1+o(1)}\leq p\leq 1/2$, we show that the non-trivial edge eigenvectors are asymptotically jointly normal. The main ingredient of the proof is an algorithm that directly computes the joint eigenvector distributions, without comparisons with GOE. The method is applicable in general. As an illustration, we also use it to prove the normal fluctuation in quantum ergodicity at the edge for Wigner matrices. Another ingredient of the proof is the isotropic local law for sparse matrices, which at the same time improves several existing results.

math.PR↗

Lecture Notes on Edge Universality for Random Regular Graphs

The purpose of this note is to explain the structure, general strategy, and main ideas of the proof in the work of Huang, McKenzie, and Yau (2024) on the Ramanujan property and edge universality of random regular graphs. The core of the argument is the derivation of self-consistent equations and a microscopic version of the loop equations for random $d$-regular graphs. We first recall the local law for random $d$-regular graphs, and then illustrate the main ideas behind the derivation of the self-consistent equations and the first loop equation.

math.PR↗

Local geometry of high-dimensional mixture models: Effective spectral theory and dynamical transitions

We study the local geometry of empirical risks in high dimensions via the spectral theory of their Hessian and information matrices. We focus on settings where the data, $(Y_\ell)_{\ell =1}^n \in \mathbb{R}^d$, are i.i.d. draws of a $k$-Gaussian mixture model, and the loss depends on the projection of the data into a fixed number of vectors, namely $\mathbf{x}^\top Y$, where $\mathbf{x}\in \mathbb{R}^{d\times C}$ are the parameters, and $C$ need not equal $k$. This setting captures a broad class of problems such as classification by one and two-layer networks and regression on multi-index models. We provide exact formulas for the limits of the empirical spectral distribution and outlier eigenvalues and eigenvectors of such matrices in the proportional asymptotics limit, where the number of samples and dimension $n,d\to\infty$ and $n/d=ϕ\in (0,\infty)$. These limits depend on the parameters $\mathbf{x}$ only through the summary statistic of the $(C+k)\times (C+k)$ Gram matrix of the parameters and class means, $\mathbf{G} = (\mathbf{x},\boldsymbolμ)^\top(\mathbf{x},\boldsymbolμ)$. It is known that under general conditions, when $\mathbf{x}$ is trained by online stochastic gradient descent, the evolution of these same summary statistics along training converges to the solution of an autonomous system of ODEs, called the effective dynamics. This enables us to connect the training dynamics to the spectral theory of these matrices generated with test data. We demonstrate our general results by analyzing the effective spectrum along the effective dynamics in the case of multi-class logistic regression. In this setting, the empirical Hessian and information matrices have substantially different spectra, each with their own static and even dynamical spectral transitions.

math.ST↗

Strong Characterization for the Airy Line Ensemble

In this paper we show that a Brownian Gibbsian line ensemble whose top curve approximates a parabola must be given by the parabolic Airy line ensemble. More specifically, we prove that if $\boldsymbol{\mathcal{L}} = (\mathcal{L}_1, \mathcal{L}_2, \ldots )$ is a line ensemble satisfying the Brownian Gibbs property, such that for any $\varepsilon > 0$ there exists a constant $\mathfrak{K} (\varepsilon) > 0$ with $$\mathbb{P} \Big[ \big| \mathcal{L}_1 (t) + 2^{-1/2} t^2 \big| \le \varepsilon t^2 + \mathfrak{K} (\varepsilon) \Big] \ge 1 - \varepsilon, \qquad \text{for all $t \in \mathbb{R}$},$$ then $\boldsymbol{\mathcal{L}}$ is the parabolic Airy line ensemble, up to an independent affine shift. Specializing this result to the case when $\boldsymbol{\mathcal{L}} (t) + 2^{-1/2} t^2$ is translation-invariant confirms a prediction of Okounkov and Sheffield from 2006 and Corwin-Hammond from 2014.

math.PR↗

Fluctuation of the Largest Eigenvalue of a Kernel Matrix with application in Graphon-based Random Graphs

In this article, we explore the spectral properties of general random kernel matrices $[K(U_i,U_j)]_{1\leq i\neq j\leq n}$ from a Lipschitz kernel $K$ with $n$ independent random variables $U_1,U_2,\ldots, U_n$ distributed uniformly over $[0,1]$. In particular we identify a dichotomy in the extreme eigenvalue of the kernel matrix, where, if the kernel $K$ is degenerate, the largest eigenvalue of the kernel matrix (after proper normalization) converges weakly to a weighted sum of independent chi-squared random variables. In contrast, for non-degenerate kernels, it converges to a normal distribution extending and reinforcing earlier results from Koltchinskii and Giné (2000). Further, we apply this result to show a dichotomy in the asymptotic behavior of extreme eigenvalues of $W$-random graphs, which are pivotal in modeling complex networks and analyzing large-scale graph behavior. These graphs are generated using a kernel $W$, termed as graphon, by connecting vertices $i$ and $j$ with probability $W(U_i, U_j)$. Our results show that for a Lipschitz graphon $W$, if the degree function is constant, the fluctuation of the largest eigenvalue (after proper normalization) converges to the weighted sum of independent chi-squared random variables and an independent normal distribution. Otherwise, it converges to a normal distribution.

math.PR↗

Spectral alignment of stochastic gradient descent for high-dimensional classification tasks

We rigorously study the relation between the training dynamics via stochastic gradient descent (SGD) and the spectra of empirical Hessian and gradient matrices. We prove that in two canonical classification tasks for multi-class high-dimensional mixtures and either 1 or 2-layer neural networks, both the SGD trajectory and emergent outlier eigenspaces of the Hessian and gradient matrices align with a common low-dimensional subspace. Moreover, in multi-layer settings this alignment occurs per layer, with the final layer's outlier eigenspace evolving over the course of training, and exhibiting rank deficiency when the SGD converges to sub-optimal classifiers. This establishes some of the rich predictions that have arisen from extensive numerical studies in the last decade about the spectra of Hessian and information matrices over the course of training in overparametrized networks.

cs.LG↗

Convergence Analysis of Probability Flow ODE for Score-based Generative Models

Score-based generative models have emerged as a powerful approach for sampling high-dimensional probability distributions. Despite their effectiveness, their theoretical underpinnings remain relatively underdeveloped. In this work, we study the convergence properties of deterministic samplers based on probability flow ODEs from both theoretical and numerical perspectives. Assuming access to $L^2$-accurate estimates of the score function, we prove the total variation between the target and the generated data distributions can be bounded above by $\mathcal{O}(d^{3/4}δ^{1/2})$ in the continuous time level, where $d$ denotes the data dimension and $δ$ represents the $L^2$-score matching error. For practical implementations using a $p$-th order Runge-Kutta integrator with step size $h$, we establish error bounds of $\mathcal{O}(d^{3/4}δ^{1/2} + d\cdot(dh)^p)$ at the discrete level. Finally, we present numerical studies on problems up to 128 dimensions to verify our theory.

cs.LG↗

On self-training of summary data with genetic applications

Prediction model training is often hindered by limited access to individual-level data due to privacy concerns and logistical challenges, particularly in biomedical research. Resampling-based self-training presents a promising approach for building prediction models using only summary-level data. These methods leverage summary statistics to sample pseudo datasets for model training and parameter optimization, allowing for model development without individual-level data. Although increasingly used in precision medicine, the general behaviors of self-training remain unexplored. In this paper, we leverage a random matrix theory framework to establish the statistical properties of self-training algorithms for high-dimensional sparsity-free summary data. We demonstrate that, within a class of linear estimators, resampling-based self-training achieves the same asymptotic predictive accuracy as conventional training methods that require individual-level datasets. These results suggest that self-training with only summary data incurs no additional cost in prediction accuracy, while offering significant practical convenience. Our analysis provides several valuable insights and counterintuitive findings. For example, while pseudo-training and validation datasets are inherently dependent, their interdependence unexpectedly cancels out when calculating prediction accuracy measures, preventing overfitting in self-training algorithms. Furthermore, we extend our analysis to show that the self-training framework maintains this no-cost advantage when combining multiple methods or when jointly training on data from different distributions. We numerically validate our findings through simulations and real data analyses using the UK Biobank. Our study highlights the potential of resampling-based self-training to advance genetic risk prediction and other fields that make summary data publicly available.

stat.ME↗

Gaussian Waves and Edge Eigenvectors of Random Regular Graphs

Backhausz and Szegedy (2019) demonstrated that the almost eigenvectors of random regular graphs converge to Gaussian waves with variance $0\leq σ^2\leq 1$. In this paper, we present an alternative proof of this result for the edge eigenvectors of random regular graphs, establishing that the variance must be $σ^2=1$. Furthermore, we show that the eigenvalues and eigenvectors are asymptotically independent. Our approach introduces a simple framework linking the weak convergence of the imaginary part of the Green's function to the convergence of eigenvectors, which may be of independent interest.

math.PR↗