SearcharxivSearch

arXiv · 1810.11741

Deep Limits of Residual Neural Networks

Abstract

Neural networks have been very successful in many applications; we often, however, lack a theoretical understanding of what the neural networks are actually learning. This problem emerges when trying to generalise to new data sets. The contribution of this paper is to show that, for the residual neural network model, the deep layer limit coincides with a parameter estimation problem for a nonlinear ordinary differential equation. In particular, whilst it is known that the residual neural network model is a discretisation of an ordinary differential equation, we show convergence in a variational sense. This implies that optimal parameters converge in the deep layer limit. This is a stronger statement than saying for a fixed parameter the residual neural network model converges (the latter does not in general imply the former). Our variational analysis provides a discrete-to-continuum $Γ$-convergence result for the objective function of the residual neural network training step to a variational problem constrained by a system of ordinary differential equations; this rigorously connects the discrete setting to a continuum problem.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Matthew Thorpe, Yves van Gennip. 2022-11-21. Deep Limits of Residual Neural Networks. https://arxiv.org/abs/1810.11741

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

On the log-concavity of the composite Bessel function $x^{\alpha}J_{\nu }\left( \beta x^{\gamma}\right) $

For a twice differentiable function $f:\left( a,b\right) \rightarrow \mathbb{R}$ define $v\left( f\right) =f^{\prime}f^{\prime}-f^{\prime\prime }f.$ It is well known that the positivity of $v\left( f\right) $ implies that the function $\left\vert f\right\vert $ is strictly log-concave on each subinterval which does not contain zeros of $f.$ In this paper we provide criteria for the positivity of $v\left( F\right) $ for the composite Bessel function $F\left( x\right) =J_{\alpha,\beta,\gamma,\nu}\left( x\right) :=x^{\alpha}J_{\nu}\left( \beta x^{\gamma}\right) $ for positive numbers $\beta$ and $\gamma$ and real numbers $\alpha$ and $\nu.$

math.CA

Riesz capacity ratios with negative exponents

We investigate sharp inequalities for ratios of Riesz capacities with negative exponents by combining computational experiments with rigorous analysis. For finite subsets of the line, we prove positivity of equilibrium masses when $-1<p<0$, enabling numerical tests of conjectured extremal ratios. In the plane, comparisons of the disk with regular polygon vertex sets reveal a cascade of transitions among the tested competitors and suggest a precise conjecture for the equilibrium measure of odd polygons, for which we give a partial proof. Numerical intersections of equality curves show that the regions where these sets outperform the disk are not simply nested. Similar numerical intersections occur in three dimensions between the regular-simplex equality curve and those of explicit five-point and six-point configurations. Motivated by the dimensional dependence of these comparisons, we prove that for each fixed $p<-2<q<0$, the regular simplex has a larger capacity ratio than the ball in all sufficiently large dimensions. Accompanying Python and Mathematica code supports reproduction and further testing of the conjectures.

math.CA

Shorter proof of dimension-free $L^p$ estimates for maximal Riesz transforms

We provide a shorter and more direct proof of $L^p$ estimates for maximal Riesz transforms (of an arbitrary order) in terms of the corresponding Riesz transforms, with a constant independent of the dimension of the Euclidean space $\mathbb R^d$. This result was originally proved by Mateu, Orobitg, P\'erez and Verdera with a constant depending on the dimension, and improved to a dimension-free inequality by Kucharski, Wr\'obel and Zienkiewicz.

math.CA