SearcharxivSearch

arXiv subjects

Elizabeth Collins-Woodfin

Publications and source records attributed to Elizabeth Collins-Woodfin.

10 recordsLinked to original sources

Exact Dynamics of Multi-class Stochastic Gradient Descent

We develop a framework for analyzing the learning dynamics of high-dimensional problems trained using one-pass stochastic gradient descent (SGD) with data from multiple anisotropic classes. Our main theorem provides exact expressions for quantities of interest, including the risk and the overlap with the true signal, in terms of a deterministic system of ODEs, valid in the high-dimensional limit. The theorem holds for a broad class of optimization problems and extends to settings where the number of classes grows with dimension. To illustrate its utility, we investigate in detail the effect of the data's anisotropic structure on the problems of binary logistic regression and least-squares (LS) loss. We study the LS in a linear multiclass setup and derive a learning-rate threshold that depends on the average eigenvalue of the covariance matrices. In the binary logistic regression, we study three cases: isotropic covariances, data covariance matrices with a large fraction of zero eigenvalues (denoted as the zero-one model), and covariance matrices with power-law spectra. We show that a structural phase transition occurs. In particular, for the zero-one model and the power-law model with sufficiently large power, SGD aligns more closely with values of the class mean that are projected onto the ``clean directions'' (i.e., directions of smaller variance). This is supported by analytical studies and numerical simulations, which show the exact asymptotic behavior of the loss in the high-dimensional limit. The effects of data anisotropy that we demonstrate are likely to hold beyond these examples and illustrate one application of the broader theorem that we prove.

stat.ML

Free energy fluctuations in SK and related spin glass models: A literature survey

Over the past 50 years, spin glass models have generated a broad range of literature in mathematics, physics, and computer science. There has been much progress in characterizing and proving the limiting free energy of various models, stemming from the original formulas of Parisi. Comparatively less is known about the more detailed topic of free energy fluctuations. This paper concerns a family of models in which there has been considerable progress on fluctuations, namely the Sherrington-Kirkpatrick (SK) and spherical Sherrington-Kirkpatrick (SSK) models, along with their multi-species analogs. We present a survey of the literature on free energy fluctuations in these 2-spin models, discussing results from different temperature regimes, with and without an external field, including results on phase transitions.

math.PR

Order of fluctuations of the free energy in the positive semi-definite MSK model at critical temperature

In this note, we consider the multi-species Sherrington-Kirkpatrick spin glass model at its conjectured critical temperature, and we show that, when the variance profile matrix $Δ^2$ is positive semi-definite, the variance of the free energy is $O(\log^2N)$. Furthermore, when one approaches this temperature threshold from the low temperature side at a rate of $O(N^{-α})$ with $α>0$, the variance is $O(\log^2N+N^{1-α})$. This result is a direct extension of the work of Chen and Lam (2019) who proved an analogous result for the SK model, and our proof methods are adapted from theirs.

math.PR

The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms

We develop a framework for analyzing the training and learning rate dynamics on a large class of high-dimensional optimization problems, which we call the high line, trained using one-pass stochastic gradient descent (SGD) with adaptive learning rates. We give exact expressions for the risk and learning rate curves in terms of a deterministic solution to a system of ODEs. We then investigate in detail two adaptive learning rates -- an idealized exact line search and AdaGrad-Norm -- on the least squares problem. When the data covariance matrix has strictly positive eigenvalues, this idealized exact line search strategy can exhibit arbitrarily slower convergence when compared to the optimal fixed learning rate with SGD. Moreover we exactly characterize the limiting learning rate (as time goes to infinity) for line search in the setting where the data covariance has only two distinct eigenvalues. For noiseless targets, we further demonstrate that the AdaGrad-Norm learning rate converges to a deterministic constant inversely proportional to the average eigenvalue of the data covariance matrix, and identify a phase transition when the covariance density of eigenvalues follows a power law distribution. We provide our code for evaluation at https://github.com/amackenzie1/highline2024.

math.OC

Free energy of the bipartite spherical SK model at critical temperature

The spherical Sherrington-Kirkpatrick (SSK) model and its bipartite analog both exhibit the phenomenon that their free energy fluctuations are asymptotically Gaussian at high temperature but asymptotically Tracy-Widom at low temperature. This was proved in two papers by Baik and Lee, for all non-critical temperatures. The case of critical temperature was recently computed for the SSK model in two separate papers, one by Landon and the other by Johnstone, Klochkov, Onatski, Pavlyshyn. In the current paper, we derive the critical temperature result for the bipartite SSK model. In particular, we find that the free energy fluctuations exhibit a transition when the temperature is in a window of size $n^{-1/3}\sqrt{\log n}$ around the critical temperature, the same window for the SSK model. Within this transitional window, the asymptotic fluctuations of the free energy are the sum of independent Gaussian and Tracy-Widom random variables.

math.PR

Hitting the High-Dimensional Notes: An ODE for SGD learning dynamics on GLMs and multi-index models

We analyze the dynamics of streaming stochastic gradient descent (SGD) in the high-dimensional limit when applied to generalized linear models and multi-index models (e.g. logistic regression, phase retrieval) with general data-covariance. In particular, we demonstrate a deterministic equivalent of SGD in the form of a system of ordinary differential equations that describes a wide class of statistics, such as the risk and other measures of sub-optimality. This equivalence holds with overwhelming probability when the model parameter count grows proportionally to the number of data. This framework allows us to obtain learning rate thresholds for stability of SGD as well as convergence guarantees. In addition to the deterministic equivalent, we introduce an SDE with a simplified diffusion coefficient (homogenized SGD) which allows us to analyze the dynamics of general statistics of SGD iterates. Finally, we illustrate this theory on some standard examples and show numerical simulations which give an excellent match to the theory.

math.OC

An edge CLT for the log determinant of Laguerre beta ensembles

We obtain a CLT for $\log|\det(M_n-s_n)|$ where $M_n$ is a scaled Laguerre $β$ ensemble and $s_n=d_++σ_n n^{-2/3}$ with $d_+$ denoting the upper edge of the limiting spectrum of $M_n$ and $σ_n$ a slowly growing function ($\log\log^2 n\llσ_n\ll\log^2 n$). In the special cases of LUE and LOE, we prove that the CLT also holds for $σ_n$ of constant order. A similar result was proved for Wigner matrices by Johnstone, Klochkov, Onatski, and Pavlyshyn. Obtaining this type of CLT of Laguerre matrices is of interest for statistical testing of critically spiked sample covariance matrices as well as free energy of bipartite spherical spin glasses at critical temperature.

math.PR

High-dimensional limit of one-pass SGD on least squares

We give a description of the high-dimensional limit of one-pass single-batch stochastic gradient descent (SGD) on a least squares problem. This limit is taken with non-vanishing step-size, and with proportionally related number of samples to problem-dimensionality. The limit is described in terms of a stochastic differential equation in high dimensions, which is shown to approximate the state evolution of SGD. As a corollary, the statistical risk is shown to be approximated by the solution of a convolution-type Volterra equation with vanishing errors as dimensionality tends to infinity. The sense of convergence is the weakest that shows that statistical risks of the two processes coincide. This is distinguished from existing analyses by the type of high-dimensional limit given as well as generality of the covariance structure of the samples.

math.PR

Spherical spin glass model with external field

We analyze the free energy and the overlaps in the 2-spin spherical Sherrington Kirkpatrick spin glass model with an external field for the purpose of understanding the transition between this model and the one without an external field. We compute the limiting values and fluctuations of the free energy as well as three types of overlaps in the setting where the strength of the external field goes to zero as the dimension of the spin variable grows. In particular, we consider overlaps with the external field, the ground state, and a replica. Our methods involve a contour integral representation of the partition function along with random matrix techniques. We also provide computations for the matching between different scaling regimes. Finally, we discuss the implications of our results for susceptibility and for the geometry of the Gibbs measure. Some of the findings of this paper are confirmed rigorously by Landon and Sosoe in their recent paper which came out independently and simultaneously.

cond-mat.dis-nn

Overlap of a spherical spin glass model with microscopic external field

We examine the behavior of the 2-spin spherical Sherrington-Kirkpatrick model with an external field by analyzing the overlap of a spin with the external field. Previous research has noted that, at low temperature, this overlap exhibits dramatically different behavior in the presence of an external field as compared to the model with no external field. The transition between those two settings was examined in a recent physics paper by Baik, Collins-Woodfin, Le Doussal, and Wu as well as a recent math paper by Landon and Sosoe. Both papers focus on the setting in which the external field strength, $h$, approaches zero as the dimension, $N$, approaches infinity. In particular, the paper of Baik et al studies the overlap with a microscopic external field ($h\sim N^{-1/2}$) but without a rigorous proof. This paper aims to give a proof of that result. The proof involves representing the generating function of the overlap as a ratio of contour integrals and then analyzing the asymptotics of those contour integrals using results from random matrix theory.

math.PR