Searcharxiv⌕ Search

arXiv subjects

Boris Hanin

Publications and source records attributed to Boris Hanin.

At least 55 records · Page 3Linked to original sources

Local Universality for Zeros and Critical Points of Monochromatic Random Waves

This paper concerns the asymptotic behavior of zeros and critical points for monochromatic random waves $ϕ_λ$ of frequency $λ$ on a compact, smooth, Riemannian manifold $(M,g)$ as $λ\rightarrow \infty$. We prove that the measure of integration over the zero set of $ϕ_λ$ restricted to balls of radius $\approx λ^{-1}$ converges in distribution to the measure of integration over the zero set of a frequency $1$ random wave on $\mathbb R^n$, where $n$ is the dimension of $M$. We also prove convergence of finite moments for the counting measure of the critical points of ϕλ, again restricted to balls of radius $\approx λ^{-1}$, to the corresponding moments for frequency $1$ random waves. We then patch together these local results to obtain new global variance estimates on the volume of the zero set and numbers of critical points of $ϕ_λ$ on all of $M.$ Our local results hold under conditions about the structure of geodesics on $M$ that are generic in the space of all metrics on $M$, while our global results hold whenever $(M,g)$ has no conjugate points (e.g is negatively curved).

math.PR↗

Deep ReLU Networks Have Surprisingly Few Activation Patterns

The success of deep networks has been attributed in part to their expressivity: per parameter, deep networks can approximate a richer class of functions than shallow networks. In ReLU networks, the number of activation patterns is one measure of expressivity; and the maximum number of patterns grows exponentially with the depth. However, recent work has showed that the practical expressivity of deep networks - the functions they can learn rather than express - is often far from the theoretical maximum. In this paper, we show that the average number of activation patterns for ReLU networks at initialization is bounded by the total number of neurons raised to the input dimension. We show empirically that this bound, which is independent of the depth, is tight both at initialization and during training, even on memorization tasks that should maximize the number of activation patterns. Our work suggests that realizing the full expressivity of deep networks may not be possible in practice, at least with current methods.

stat.ML↗

Finite Depth and Width Corrections to the Neural Tangent Kernel

We prove the precise scaling, at finite depth and width, for the mean and variance of the neural tangent kernel (NTK) in a randomly initialized ReLU network. The standard deviation is exponential in the ratio of network depth to width. Thus, even in the limit of infinite overparameterization, the NTK is not deterministic if depth and width simultaneously tend to infinity. Moreover, we prove that for such deep and wide networks, the NTK has a non-trivial evolution during training by showing that the mean of its first SGD update is also exponential in the ratio of network depth to width. This is sharp contrast to the regime where depth is fixed and network width is very large. Our results suggest that, unlike relatively shallow and wide networks, deep and wide ReLU networks are capable of learning data-dependent features even in the so-called lazy training regime.

cs.LG↗

Complexity of Linear Regions in Deep Networks

It is well-known that the expressivity of a neural network depends on its architecture, with deeper networks expressing more complex functions. In the case of networks that compute piecewise linear functions, such as those with ReLU activation, the number of distinct linear regions is a natural measure of expressivity. It is possible to construct networks with merely a single region, or for which the number of linear regions grows exponentially with depth; it is not clear where within this range most networks fall in practice, either before or after training. In this paper, we provide a mathematical framework to count the number of linear regions of a piecewise linear network and measure the volume of the boundaries between these regions. In particular, we prove that for networks at initialization, the average number of regions along any one-dimensional subspace grows linearly in the total number of neurons, far below the exponential upper bound. We also find that the average distance to the nearest region boundary at initialization scales like the inverse of the number of neurons. Our theory suggests that, even after training, the number of linear regions is far below exponential, an intuition that matches our empirical observations. We conclude that the practical expressivity of neural networks is likely far below that of the theoretical maximum, and that this gap can be quantified.

stat.ML↗

Interface Asymptotics of Wigner-Weyl Distributions for the Harmonic Oscillator

We prove several types of scaling results for Wigner distributions of spectral projections of the isotropic Harmonic oscillator on $\mathbb R^d$. In prior work, we studied Wigner distributions $W_{\hbar, E_N(\hbar)}(x, ξ)$ of individual eigenspace projections. In this continuation, we study Weyl sums of such Wigner distributions as the eigenvalue $E_N(\hbar)$ ranges over spectral intervals $[E - δ(\hbar), E + δ(\hbar)]$ of various widths $δ(\hbar)$ and as $(x, ξ) \in T^*\mathbb R^d$ ranges over tubes of various widths around the classical energy surface $Σ_E \subset T^*\mathbb R^d$. The main results pertain to interface Airy scaling asymptotics around $Σ_E$, which divides phase space into an allowed and a forbidden region. The first result pertains to $δ(\hbar) = \hbar$ widths and generalizes our earlier results on Wigner distributions of individual eigenspace projections. Our second result pertains to $δ(\hbar) = \hbar^{2/3}$ spectral widths and Airy asymptotics of the Wigner distributions in $\hbar^{2/3}$-tubes around $Σ_E$. Our third result pertains to bulk spectral intervals of fixed width and the behavior of the Wigner distributions inside the energy surface, outside the energy surface and in a thin neighborhood of the energy surface.

math-ph↗

Interface Asymptotics of Eigenspace Wigner distributions for the Harmonic Oscillator

Eigenspaces of the quantum isotropic Harmonic Oscillator $\hat{H}_{\hbar} : = - \frac{\hbar^2}{2} Δ+ \frac{||x||^2}{2}$ on $\mathbb{R}^d$ have extremally high multiplicites and the eigenspace projections $Π_{\hbar, E_N(\hbar)} $ have special asymptotic properties. This article gives a detailed study of their Wigner distributions $W_{\hbar, E_N(\hbar)}(x, ξ)$. Heuristically, if $E_N(\hbar) = E$, $W_{\hbar, E_N(\hbar)}(x, ξ)$ is the `quantization' of the energy surface $Σ_E$, and should be like the delta-function $δ_{Σ_E}$ on $Σ_E$; rigorously, $W_{\hbar, E_N(\hbar)}(x, ξ)$ tends in a weak* sense to $δ_{Σ_E}$. But its pointwise asymptotics and scaling asymptotics have more structure. The main results give Bessel asymptotics of $W_{\hbar, E_N(\hbar)}(x, ξ)$ in the interior $H(x, ξ) < E$ of $Σ_E$; interface Airy scaling asymptotics in tubes of radius $\hbar^{2/3}$ around $Σ_E$, with $(x, ξ)$ either in the interior or exterior of the energy ball; and exponential decay rates in the exterior of the energy surface.

math-ph↗

Products of Many Large Random Matrices and Gradients in Deep Neural Networks

We study products of random matrices in the regime where the number of terms and the size of the matrices simultaneously tend to infinity. Our main theorem is that the logarithm of the $\ell_2$ norm of such a product applied to any fixed vector is asymptotically Gaussian. The fluctuations we find can be thought of as a finite temperature correction to the limit in which first the size and then the number of matrices tend to infinity. Depending on the scaling limit considered, the mean and variance of the limiting Gaussian depend only on either the first two or the first four moments of the measure from which matrix entries are drawn. We also obtain explicit error bounds on the moments of the norm and the Kolmogorov-Smirnov distance to a Gaussian. Finally, we apply our result to obtain precise information about the stability of gradients in randomly initialized deep neural networks with ReLU activations. This provides a quantitative measure of the extent to which the exploding and vanishing gradient problem occurs in a fully connected neural network with ReLU activations and a given architecture.

math.PR↗

How to Start Training: The Effect of Initialization and Architecture

We identify and study two common failure modes for early training in deep ReLU nets. For each we give a rigorous proof of when it occurs and how to avoid it, for fully connected and residual architectures. The first failure mode, exploding/vanishing mean activation length, can be avoided by initializing weights from a symmetric distribution with variance 2/fan-in and, for ResNets, by correctly weighting the residual modules. We prove that the second failure mode, exponentially large variance of activation length, never occurs in residual nets once the first failure mode is avoided. In contrast, for fully connected nets, we prove that this failure mode can happen and is avoided by keeping constant the sum of the reciprocals of layer widths. We demonstrate empirically the effectiveness of our theoretical results in predicting when networks are able to start training. In particular, we note that many popular initializations fail our criteria, whereas correct initialization and architecture allows much deeper networks to be trained.

stat.ML↗

Which Neural Net Architectures Give Rise To Exploding and Vanishing Gradients?

We give a rigorous analysis of the statistical behavior of gradients in a randomly initialized fully connected network N with ReLU activations. Our results show that the empirical variance of the squares of the entries in the input-output Jacobian of N is exponential in a simple architecture-dependent constant beta, given by the sum of the reciprocals of the hidden layer widths. When beta is large, the gradients computed by N at initialization vary wildly. Our approach complements the mean field theory analysis of random networks. From this point of view, we rigorously compute finite width corrections to the statistics of gradients at the edge of chaos.

stat.ML↗

The lemniscate tree of a random polynomial

To each generic complex polynomial $p(z)$ there is associated a labeled binary tree (here referred to as a "lemniscate tree") that encodes the topological type of the graph of $|p(z)|$. The branching structure of the lemniscate tree is determined by the configuration (i.e., arrangement in the plane) of the singular components of those level sets $|p(z)|=t$ passing through a critical point. In this paper, we address the question "How many branches appear in a typical lemniscate tree?" We answer this question first for a lemniscate tree sampled uniformly from the combinatorial class and second for the lemniscate tree arising from a random polynomial generated by i.i.d. zeros. From a more general perspective, these results take a first step toward a probabilistic treatment (within a specialized setting) of Arnold's program of enumerating algebraic Morse functions.

math.PR↗

Approximating Continuous Functions by ReLU Nets of Minimal Width

This article concerns the expressive power of depth in deep feed-forward neural nets with ReLU activations. Specifically, we answer the following question: for a fixed $d_{in}\geq 1,$ what is the minimal width $w$ so that neural nets with ReLU activations, input dimension $d_{in}$, hidden layer widths at most $w,$ and arbitrary depth can approximate any continuous, real-valued function of $d_{in}$ variables arbitrarily well? It turns out that this minimal width is exactly equal to $d_{in}+1.$ That is, if all the hidden layer widths are bounded by $d_{in}$, then even in the infinite depth limit, ReLU nets can only express a very limited class of functions, and, on the other hand, any continuous function on the $d_{in}$-dimensional unit cube can be approximated to arbitrary precision by ReLU nets in which all hidden layers have width exactly $d_{in}+1.$ Our construction in fact shows that any continuous function $f:[0,1]^{d_{in}}\to\mathbb R^{d_{out}}$ can be approximated by a net of width $d_{in}+d_{out}$. We obtain quantitative depth estimates for such an approximation in terms of the modulus of continuity of $f$.

stat.ML↗

Universal Function Approximation by Deep Neural Nets with Bounded Width and ReLU Activations

This article concerns the expressive power of depth in neural nets with ReLU activations and bounded width. We are particularly interested in the following questions: what is the minimal width $w_{\text{min}}(d)$ so that ReLU nets of width $w_{\text{min}}(d)$ (and arbitrary depth) can approximate any continuous function on the unit cube $[0,1]^d$ aribitrarily well? For ReLU nets near this minimal width, what can one say about the depth necessary to approximate a given function? Our approach to this paper is based on the observation that, due to the convexity of the ReLU activation, ReLU nets are particularly well-suited for representing convex functions. In particular, we prove that ReLU nets with width $d+1$ can approximate any continuous convex function of $d$ variables arbitrarily well. These results then give quantitative depth estimates for the rate of approximation of any continuous scalar function on the $d$-dimensional cube $[0,1]^d$ by ReLU nets with width $d+3.$

stat.ML↗

Level Spacings and Nodal Sets at Infinity for Radial Perturbations of the Harmonic Oscillator

We study properties of the nodal sets of high frequency eigenfunctions and quasimodes for radial perturbations of the Harmonic Oscillator. In particular, we consider nodal sets on spheres of large radius (in the classically forbidden region) for quasimodes with energies lying in intervals around a fixed energy $E$. For well chosen intervals we show that these nodal sets exhibit quantitatively different behavior compared to those of the unperturbed Harmonic Oscillator. These energy intervals are defined via a careful analysis of the eigenvalue spacings for the perturbed operator, based on analytic perturbation theory and linearization formulas for Laguerre polynomials.

math-ph↗

Scaling of Harmonic Oscillator Eigenfunctions and Their Nodal Sets Around the Caustic

We study the scaling asymptotics of the eigenspace projection kernels $Π_{\hbar, E}(x,y)$ of the isotropic Harmonic Oscillator $- \hbar ^2 Δ+ |x|^2$ of eigenvalue $E = \hbar(N + \frac{d}{2})$ in the semi-classical limit $\hbar \to 0$. The principal result is an explicit formula for the scaling asymptotics of $Π_{\hbar, E}(x,y)$ for $x,y$ in a $\hbar^{2/3}$ neighborhood of the caustic $\mathcal C_E$ as $\hbar \to 0.$ The scaling asymptotics are applied to the distribution of nodal sets of Gaussian random eigenfunctions around the caustic as $\hbar \to 0$. In previous work we proved that the density of zeros of Gaussian random eigenfunctions of $\hat{H}_{\hbar}$ have different orders in the Planck constant $\hbar$ in the allowed and forbidden regions: In the allowed region the density is of order $\hbar^{-1}$ while it is $\hbar^{-1/2}$ in the forbidden region. Our main result on nodal sets is that the density of zeros is of order $\hbar^{-\frac{2}{3}}$ in an $\hbar^{\frac{2}{3}}$-tube around the caustic. This tube radius is the `critical radius'. For annuli of larger inner and outer radii $\hbar^α$ with $0< α< \frac{2}{3}$ we obtain density results which interpolate between this critical radius result and our prior ones in the allowed and forbidden region. We also show that the Hausdorff $(d-2)$-dimensional measure of the intersection of the nodal set with the caustic is of order $\hbar^{- \frac{2}{3}}$.

math-ph↗

Nodal Sets of Smooth Functions with Finite Vanishing Order and p-Sweepouts

We show that on a compact Riemmanian manifold $(M,g)$, nodal sets of linear combinations of any $p+1$ smooth functions form an admissible $p-$sweepout provided these linear combinations have uniformly bounded vanishing order. This applies in particular to finite linear combinations of Laplace eigenfunctions. As a result, we obtain a new proof of the Gromov, Guth, Marques--Neves upper bounds on the min-max $p$-widths of $M.$ We also prove that close to a point at which a smooth function on $\mathbb{R}^{n+1}$ vanishes to order $k$, its nodal set is contained in the union of $k$ $W^{1,p}$ graphs for some $p > 1$. This implies that the nodal set is locally countably $n$-rectifiable and has locally finite $\mathcal{H}^n$ measure, facts which also follow from a previous result of Bär. Finally, we prove the continuity of the Hausdorff measure of nodal sets under heat flow.

math.AP↗

C-infinity Scaling Asymptotics for the Spectral Function of the Laplacian

This article concerns new off-diagonal estimates on the remainder and its derivatives in the pointwise Weyl law on a compact n-dimensional Riemannian manifold. As an application, we prove that near any non self-focal point, the scaling limit of the spectral projector of the Laplacian onto frequency windows of constant size is a normalized Bessel function depending only on n.

math.AP↗

Pairing of Zeros and Critical Points for Random Polynomials

Let p_N be a random degree N polynomial in one complex variable whose zeros are chosen independently from a fixed probability measure mu on the Riemann sphere S^2. This article proves that if we condition p_N to have a zero at some fixed point xi in , then, with high probability, there will be a critical point w_xi a distance 1/N away from xi. This 1/N distance is much smaller than the one over root N typical spacing between nearest neighbors for N i.i.d. points on S^2. Moreover, with the same high probability, the argument of w_xi relative to xi is a deterministic function of mu plus fluctuations on the order of 1/N.

math.PR↗

Scaling Limit for the Kernel of the Spectral Projector and Remainder Estimates in the Pointwise Weyl Law

Let (M, g) be a compact smooth Riemannian manifold. We obtain new off-diagonal estimates as λ tend to infinity for the remainder in the pointwise Weyl Law for the kernel of the spectral projector of the Laplacian onto functions with frequency at most λ. A corollary is that, when rescaled around a non self-focal point, the kernel of the spectral projector onto the frequency interval (λ, λ+ 1] has a universal scaling limit as λ goes to infinity (depending only on the dimension of M). Our results also imply that if M has no conjugate points, then immersions of M into Euclidean space by an orthonormal basis of eigenfunctions with frequencies in (λ, λ+ 1] are embeddings for all λ sufficiently large.

math.SP↗