SearcharxivSearch

arXiv subjects

Hong Chang Ji

Publications and source records attributed to Hong Chang Ji.

13 recordsLinked to original sources

Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning

It is folklore that reusing training data more than once can improve the statistical efficiency of gradient-based learning. While this phenomenon has been extensively studied in linear regression, the benefit of multi-pass gradient descent (GD, which reuses all the data) over one-pass stochastic gradient descent (online SGD, which uses each data point only once) is not well-understood in nonlinear and non-convex settings, except for a loss modification mechanism achieved by the first two passes on the data. In this work, we consider learning a $d$-dimensional single-index model with a quadratic activation, for which it is known that one-pass SGD requires $n\gtrsim d\log d$ samples to achieve weak recovery. We first show that this $\log d$ factor in the sample complexity persists for full-batch spherical GD on the correlation loss; however, by simply truncating the activation, full-batch GD exhibits a favorable optimization landscape at $n \simeq d$ samples, thereby outperforming one-pass SGD (with the same activation) in statistical efficiency. We complement this result with a trajectory analysis of full-batch GD on the squared loss from small initialization, showing that $n \gtrsim d$ samples and $T \gtrsim\log d$ gradient steps suffice to achieve strong (exact) recovery.

stat.ML

Optimal Estimation in Orthogonally Invariant Generalized Linear Models: Spectral Initialization and Approximate Message Passing

We consider the problem of parameter estimation from a generalized linear model with a random design matrix that is orthogonally invariant in law. Such a model allows the design have an arbitrary distribution of singular values and only assumes that its singular vectors are generic. It is a vast generalization of the i.i.d. Gaussian design typically considered in the theoretical literature, and is motivated by the fact that real data often have a complex correlation structure so that methods relying on i.i.d. assumptions can be highly suboptimal. Building on the paradigm of spectrally-initialized iterative optimization, this paper proposes optimal spectral estimators and combines them with an approximate message passing (AMP) algorithm, establishing rigorous performance guarantees for these two algorithmic steps. Both the spectral initialization and the subsequent AMP meet existing conjectures on the fundamental limits to estimation -- the former on the optimal sample complexity for efficient weak recovery, and the latter on the optimal errors. Numerical experiments suggest the effectiveness of our methods and accuracy of our theory beyond orthogonally invariant data.

math.ST

On the spectral edge of non-Hermitian random matrices

For general non-Hermitian random matrices $X$ and deterministic deformation matrices $A$, we prove that the local eigenvalue statistics of $A+X$ close to the typical edge points of its spectrum are universal. Furthermore, we show that under natural assumptions on $A$ the spectrum of $A+X$ does not have outliers at a distance larger than the natural fluctuation scale of the eigenvalues. As a consequence, the number of eigenvalues in each component of $\mathrm{Spec}(A+X)$ is deterministic.

math.PR

Non-Hermitian spectral universality at critical points

For general large non-Hermitian random matrices $X$ and deterministic normal deformations $A$, we prove that the local eigenvalue statistics of $A+X$ close to the critical edge points of its spectrum are universal. This concludes the proof of the third and last remaining typical universality class for non-Hermitian random matrices, after bulk and sharp edge universalities have been established in recent years.

math.PR

Spectral Estimators for Structured Generalized Linear Models via Approximate Message Passing

We consider the problem of parameter estimation in a high-dimensional generalized linear model. Spectral methods obtained via the principal eigenvector of a suitable data-dependent matrix provide a simple yet surprisingly effective solution. However, despite their wide use, a rigorous performance characterization, as well as a principled way to preprocess the data, are available only for unstructured (i.i.d.\ Gaussian and Haar orthogonal) designs. In contrast, real-world data matrices are highly structured and exhibit non-trivial correlations. To address the problem, we consider correlated Gaussian designs capturing the anisotropic nature of the features via a covariance matrix $Σ$. Our main result is a precise asymptotic characterization of the performance of spectral estimators. This allows us to identify the optimal preprocessing that minimizes the number of samples needed for parameter estimation. Surprisingly, such preprocessing is universal across a broad set of designs, which partly addresses a conjecture on optimal spectral estimators for rotationally invariant models. Our principled approach vastly improves upon previous heuristic methods, including for designs common in computational imaging and genetics. The proposed methodology, based on approximate message passing, is broadly applicable and opens the way to the precise characterization of spiked matrices and of the corresponding spectral methods in a variety of settings.

math.ST

Density of Brown measure of free circular Brownian motion

We consider the Brown measure of the free circular Brownian motion, $\boldsymbol{a}+\sqrt{t}\boldsymbol{x}$, with an arbitrary initial condition $\boldsymbol{a}$, i.e. $\boldsymbol{a}$ is a general non-normal operator and $\boldsymbol{x}$ is a circular element $*$-free from $\boldsymbol{a}$. We prove that, under a mild assumption on $\boldsymbol{a}$, the density of the Brown measure has one of the following two types of behavior around each point on the boundary of its support -- either (i) sharp cut, i.e. a jump discontinuity along the boundary, or (ii) quadratic decay at certain critical points on the boundary. Our result is in direct analogy with the previously known phenomenon for the spectral density of free semicircular Brownian motion, whose singularities are either a square-root edge or a cubic cusp. We also provide several examples and counterexamples, one of which shows that our assumption on $\boldsymbol{a}$ is necessary.

math.PR

Wegner estimate and upper bound on the eigenvalue condition number of non-Hermitian random matrices

We consider $N\times N$ non-Hermitian random matrices of the form $X+A$, where $A$ is a general deterministic matrix and $\sqrt{N}X$ consists of independent entries with zero mean, unit variance, and bounded densities. For this ensemble, we prove (i) a Wegner estimate, i.e. that the local density of eigenvalues is bounded by $N^{1+o(1)}$ and (ii) that the expected condition number of any bulk eigenvalue is bounded by $N^{1+o(1)}$; both results are optimal up to the factor $N^{o(1)}$. The latter result complements the very recent matching lower bound obtained in [15] (arXiv:2301.03549) and improves the $N$-dependence of the upper bounds in [5,6,32] (arXiv:1906.11819, arXiv:2005.08930, arXiv:2005.08908). Our main ingredient, a near-optimal lower tail estimate for the small singular values of $X+A-z$, is of independent interest.

math.PR

Tracy-Widom limit for free sum of random matrices

We consider fluctuations of the largest eigenvalues of the random matrix model $A+UBU^{*}$ where $A$ and $B$ are $N \times N$ deterministic Hermitian (or symmetric) matrices and $U$ is a Haar-distributed unitary (or orthogonal) matrix. We prove that the largest eigenvalue weakly converges to the Tracy-Widom distribution, under mild assumptions on $A$ and $B$ to guarantee that the density of states of the model decays as square root around the upper edge. Our proof is based on the comparison of the Green function along the Dyson Brownian motion starting from the matrix $A + UBU^{*}$ and ending at time $N^{-1/3+χ}$. As a byproduct of our proof, we also prove an optimal local law for the Dyson Brownian motion up to the constant time scale.

math.PR

Spiked multiplicative random matrices and principal components

In this paper, we study the eigenvalues and eigenvectors of the spiked invariant multiplicative models when the randomness is from Haar matrices. We establish the limits of the outlier eigenvalues $\widehatλ_i$ and the generalized components ($\langle \mathbf{v}, \widehat{\mathbf{u}}_i \rangle$ for any deterministic vector $\mathbb{v}$) of the outlier eigenvectors $\widehat{\mathbf{u}}_i$ with optimal convergence rates. Moreover, we prove that the non-outlier eigenvalues stick with those of the unspiked matrices and the non-outlier eigenvectors are delocalized. The results also hold near the so-called BBP transition and for degenerate spikes. On one hand, our results can be regarded as a refinement of the counterparts of [12] under additional regularity conditions. On the other hand, they can be viewed as an analog of [34] by replacing the random matrix with i.i.d. entries with Haar random matrix.

math.PR

Local laws for multiplication of random matrices

Consider the random matrix model $A^{1/2} UBU^* A^{1/2},$ where $A$ and $B$ are two $N \times N$ deterministic matrices and $U$ is either an $N \times N$ Haar unitary or orthogonal random matrix. It is well-known that on the macroscopic scale, the limiting empirical spectral distribution (ESD) of the above model is given by the free multiplicative convolution of the limiting ESDs of $A$ and $B,$ denoted as $μ_α\boxtimes μ_β,$ where $μ_α$ and $μ_β$ are the limiting ESDs of $A$ and $B,$ respectively. In this paper, we study the asymptotic microscopic behavior of the edge eigenvalues and eigenvectors statistics. We prove that both the density of $μ_A \boxtimes μ_B,$ where $μ_A$ and $μ_B$ are the ESDs of $A$ and $B,$ respectively and the associated subordination functions have a regular behavior near the edges. Moreover, we establish the local laws near the edges on the optimal scale. In particular, we prove that the entries of the resolvent are close to some functionals depending only on the eigenvalues of $A, B$ and the subordination functions with optimal convergence rates. Our proofs and calculations are based on the techniques developed for the additive model $A+UBU^*$ in [3,5,6,8], and our results can be regarded as the counterparts of [8] for the multiplicative model.

math.PR

Functional CLT for non-Hermitian random matrices

For large dimensional non-Hermitian random matrices $X$ with real or complex independent, identically distributed, centered entries, we consider the fluctuations of $f(X)$ as a matrix where $f$ is an analytic function around the spectrum of $X$. We prove that for a generic bounded square matrix $A$, the quantity $\mathrm{Tr}f(X)A$ exhibits Gaussian fluctuations as the matrix size grows to infinity, which consists of two independent modes corresponding to the tracial and traceless parts of $A$. We find a new formula for the variance of the traceless part that involves the Frobenius norm of $A$ and the $L^{2}$-norm of $f$ on the boundary of the limiting spectrum.

math.PR

Regularity properties of free multiplicative convolution on the positive line

Given two nondegenerate Borel probability measures $μ$ and $ν$ on $\mathbb{R}_{+}=[0,\infty)$, we prove that their free multiplicative convolution $μ\boxtimesν$ has zero singular continuous part and its absolutely continuous part has a density bounded by $x^{-1}$. When $μ$ and $ν$ are compactly supported Jacobi measures on $(0,\infty)$ having power law behavior with exponents in $(-1,1)$, we prove that $μ\boxtimesν$ is another Jacobi measure whose density has square root decay at the edges of its support.

math.PR

Gaussian fluctuations for linear spectral statistics of deformed Wigner matrices

We consider large-dimensional Hermitian or symmetric random matrices of the form $W=M+\vartheta V$ where $M$ is a Wigner matrix and $V$ is a real diagonal matrix whose entries are independent of $M$. For a large class of diagonal matrices $V$, we prove that the fluctuations of linear spectral statistics of $W$ for $C^{2}_{c}$ test function can be decomposed into that of $M$ and of $V$, and that each of those weakly converges to a Gaussian distribution. We also calculate the formulae for the means and variances of the limiting distributions.

math.PR