SearcharxivSearch

arXiv · 2306.02108

Random matrix theory and the loss surfaces of neural networks

Abstract

Neural network models are one of the most successful approaches to machine learning, enjoying an enormous amount of development and research over recent years and finding concrete real-world applications in almost any conceivable area of science, engineering and modern life in general. The theoretical understanding of neural networks trails significantly behind their practical success and the engineering heuristics that have grown up around them. Random matrix theory provides a rich framework of tools with which aspects of neural network phenomenology can be explored theoretically. In this thesis, we establish significant extensions of prior work using random matrix theory to understand and describe the loss surfaces of large neural networks, particularly generalising to different architectures. Informed by the historical applications of random matrix theory in physics and elsewhere, we establish the presence of local random matrix universality in real neural networks and then utilise this as a modeling assumption to derive powerful and novel results about the Hessians of neural network loss surfaces and their spectra. In addition to these major contributions, we make use of random matrix models for neural network loss surfaces to shed light on modern neural network training approaches and even to derive a novel and effective variant of a popular optimisation algorithm. Overall, this thesis provides important contributions to cement the place of random matrix theory in the theoretical study of modern neural networks, reveals some of the limits of existing approaches and begins the study of an entirely new role for random matrix theory in the theory of deep learning with important experimental discoveries and novel theoretical results based on local random matrix universality.

Explore related subjects

Keep this discovery

BibTeXRIS

Nicholas P Baskerville. 2023-06-03. Random matrix theory and the loss surfaces of neural networks. https://arxiv.org/abs/2306.02108

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Why we should condition denoising diffusion generative models on windows of past observations

Data assimilation (DA) is, traditionally, a cycling process that relies on time-dependent priors to propagate information from past observations to future cycles. Using denoising diffusion generative modeling for DA is challenging because standard approaches use a fixed training data set, which in turn leads to a static prior that ignores information from past observations. Because past observations are ignored, DA systems with static priors lead to larger posterior errors than cycling DA systems. Incorporating time-dependent priors into generative models, however, requires expensive and frequent retraining. Motivated by linear systems theory - where the dependence of a prediction of a Kalman filter on past observations decays exponentially - we condition diffusion models on short windows of past observations. Specifically, we describe training procedures for two frameworks: a diffusion DA system predicting the current state given a set of past observations, and a diffusion ``direct observation prediction'' (DOP) system, predicting future observations given a set of past observations. Using a canonical linear system, we show that both systems can achieve the minimal posterior error characteristic of a fully-cycled DA/DOP system, without re-training, provided the time windows are long enough. The linear setup ensures analytical tractability, avoids confounding neural network training errors, and confirms that conditioning on windows of past observations is required for efficient and accurate diffusion-based DA or DOP.

math-ph

The kinematic structures and the inertial geometry of a moving charge

We ask how much of the geometry a charged particle moves in is fixed by its motion, and how much a particle must bring. We ask of a symplectic structure only that it relate velocity to momentum as Hamilton's equations do, and we ask it of every energy at once. In particular, we show that the structures meeting that demand are the canonical one and its twists by a closed two-form of the base. A field provides the two-form, a particle the multiplier before it, which we identify constitutively with its charge. Thus, a single energy governs a family of structures, and each particle takes the one its charge fixes. We then ask what a particle must bring to be given a momentum, and we show that the degree of that map settles the degree at which a field enters Newton's Second Law. An antisymmetric bilinear form returns no Lorentz force, whilst a Randers metric returns one --- a length whose difference from a Riemannian one is linear in the velocity. Moreover, we find that metric already within the twisted structure, as its primitive over a level of the free energy, and its law of transport to be nonlinear, no affine connection being known to serve. Under an indefinite signature the length parts from the dynamics, and the extremals turn from shortest to longest. On the round sphere a monopole flux leaves no such metric, whilst the transport remains and prequantisation, given a unit of action, restricts the charge to a lattice. In this manner, we conclude that each charge-to-mass ratio receives a geometry of its own, so that by a functionalist criterion none of them is the spacetime of a charged particle.

math-ph