SearcharxivSearch

arXiv subjects

Kathryn Lindsey

Publications and source records attributed to Kathryn Lindsey.

18 recordsLinked to original sources

Regularization Implies balancedness in the deep linear network

We use geometric invariant theory (GIT) to study the deep linear network (DLN). The Kempf-Ness theorem is used to establish that the $L^2$ regularizer is minimized on the balanced manifold. We introduce related balancing flows using the Riemannian geometry of fibers. The balancing flow defined by the $L^2$ regularizer is shown to converge to the balanced manifold at a uniform exponential rate. The balancing flow defined by the squared moment map is computed explicitly and shown to converge globally. This framework allows us to decompose the training dynamics into two distinct gradient flows: a regularizing flow on fibers and a learning flow on the balanced manifold. It also provides a common mathematical framework for balancedness in deep learning and linear systems theory. We use this framework to interpret balancedness in terms of fast-slow systems, model reduction and Bayesian principles.

cs.LG

Functional dimension of feedforward ReLU neural networks

It is well-known that the parameterized family of functions representable by fully-connected feedforward neural networks with ReLU activation function is precisely the class of piecewise linear functions with finitely many pieces. It is less well-known that for every fixed architecture of ReLU neural network, the parameter space admits positive-dimensional spaces of symmetries, and hence the local functional dimension near any given parameter is lower than the parametric dimension. In this work we carefully define the notion of functional dimension, show that it is inhomogeneous across the parameter space of ReLU neural network functions, and continue an investigation - initiated in [14] and [5] - into when the functional dimension achieves its theoretical maximum. We also study the quotient space and fibers of the realization map from parameter space to function space, supplying examples of fibers that are disconnected, fibers upon which functional dimension is non-constant, and fibers upon which the symmetry group acts non-transitively.

math.MG

On Functional Dimension and Persistent Pseudodimension

For any fixed feedforward ReLU neural network architecture, it is well-known that many different parameter settings can determine the same function. It is less well-known that the degree of this redundancy is inhomogeneous across parameter space. In this work, we discuss two locally applicable complexity measures for ReLU network classes and what we know about the relationship between them: (1) the local functional dimension [14, 18], and (2) a local version of VC dimension that we call persistent pseudodimension. The former is easy to compute on finite batches of points; the latter should give local bounds on the generalization gap, which would inform an understanding of the mechanics of the double descent phenomenon [7].

cs.LG

Master Teapots and Entropy Algorithms for the Mandelbrot Set

We construct an analogue of W. Thurston's "Master teapot" for each principal vein in the Mandelbrot set, and generalize geometric properties known for the corresponding object for real maps. In particular, we show that eigenvalues outside the unit circle move continuously, while we show "persistence" for roots inside the unit circle. As an application, this shows that the outside part of the corresponding "Thurston set" is path connected. In order to do this, we define a version of kneading theory for principal veins, and we prove the equivalence of several algorithms that compute the core entropy.

math.DS

Local and global topological complexity measures OF ReLU neural network functions

We apply a generalized piecewise-linear (PL) version of Morse theory due to Grunert-Kuhnel-Rote to define and study new local and global notions of topological complexity for fully-connected feedforward ReLU neural network functions, F: R^n -> R. Along the way, we show how to construct, for each such F, a canonical polytopal complex K(F) and a deformation retract of the domain onto K(F), yielding a convenient compact model for performing calculations. We also give a construction showing that local complexity can be arbitrarily high.

math.AT

Hidden symmetries of ReLU networks

The parameter space for any fixed architecture of feedforward ReLU neural networks serves as a proxy during training for the associated class of functions - but how faithful is this representation? It is known that many different parameter settings can determine the same function. Moreover, the degree of this redundancy is inhomogeneous: for some networks, the only symmetries are permutation of neurons in a layer and positive scaling of parameters at a neuron, while other networks admit additional hidden symmetries. In this work, we prove that, for any network architecture where no layer is narrower than the input, there exist parameter settings with no hidden symmetries. We also describe a number of mechanisms through which hidden symmetries can arise, and empirically approximate the functional dimension of different network architectures at initialization. These experiments indicate that the probability that a network has no hidden symmetries decreases towards 0 as depth increases, while increasing towards 1 as width and input dimension increase.

cs.LG

On the deck groups of iterates of bicritical rational maps

Given a rational map $f:\widehat{\mathbb C}\to\widehat{\mathbb C}$ on the Riemann sphere, we define $\mathrm{Deck}(f)$ to be the group of Möbius transformations $μ$ satisfying $f \circ μ= f$. In this note, we consider the groups $\mathrm{Deck}(f^k)$, where $f$ is a \emph{bicritical} rational map (that is, a rational map with exactly two critical points) and $f^k$ denotes the $k$th iterate of $f$. In particular, we give a complete description of which groups (up to isomorphism) arise as the groups $\mathrm{Deck}(f^k)$ for bicritical rational maps $f$.

math.DS

Bicritical rational maps with a common iterate

Let $f$ be a degree $d$ bicritical rational map with critical point set $\mathcal{C}_f$ and critical value set $\mathcal{V}_f$. Using the group $\textrm{Deck}(f^k)$ of deck transformations of $f^k$, we show that if $g$ is a bicritical rational map which shares an iterate with $f$ then $\mathcal{C}_f = \mathcal{C}_g$ and $\mathcal{V}_f = \mathcal{V}_g$. Using this, we show that if two bicritical rational maps of even degree $d$ share an iterate then they share a second iterate, and both maps belong to the symmetry locus of degree $d$ bicritical rational maps.

math.DS

On transversality of bent hyperplane arrangements and the topological expressiveness of ReLU neural networks

Let F:R^n -> R be a feedforward ReLU neural network. It is well-known that for any choice of parameters, F is continuous and piecewise (affine) linear. We lay some foundations for a systematic investigation of how the architecture of F impacts the geometry and topology of its possible decision regions for binary classification tasks. Following the classical progression for smooth functions in differential topology, we first define the notion of a generic, transversal ReLU neural network and show that almost all ReLU networks are generic and transversal. We then define a partially-oriented linear 1-complex in the domain of F and identify properties of this complex that yield an obstruction to the existence of bounded connected components of a decision region. We use this obstruction to prove that a decision region of a generic, transversal ReLU network F: R^n -> R with a single hidden layer of dimension (n + 1) can have no more than one bounded connected component.

math.CO

A characterization of Thurston's Master Teapot

We prove an explicit characterization of the points in Thurston's Master Teapot. This description can be implemented algorithmically to test whether a point in $\mathbb{C} \times \mathbb{R}$ belongs to the complement of the Master Teapot. As an application, we show that the intersection of the Master Teapot with the unit cylinder is not symmetrical under reflection through the plane that is the product of the imaginary axis of $\mathbb{C}$ and $\mathbb{R}$.

math.DS

The Shape of Thurston's Master Teapot

We establish basic geometric and topological properties of Thurston's Master Teapot and the Thurston set for superattracting unimodal self-maps of intervals. In particular, the Master Teapot is connected, contains the unit cylinder, and its intersection with a set $\mathbb{D} \times \{c\}$ grows monotonically with $c$. We show that the Thurston set described above is not equal to the Thurston set for postcritically finite tent maps, and we provide an arithmetic explanation for why certain gaps appear in plots of finite approximations of the Thurston set.

math.DS

Convex shapes and harmonic caps

Any planar shape $P\subset \mathbb{C}$ can be embedded isometrically as part of the boundary surface $S$ of a convex subset of $\mathbb{R}^3$ such that $\partial P$ supports the positive curvature of $S$. The complement $Q = S \setminus P$ is the associated {\em cap}. We study the cap construction when the curvature is harmonic measure on the boundary of $(\hat{\mathbb{C}}\setminus P, \infty)$. Of particular interest is the case when $P$ is a filled polynomial Julia set and the curvature is proportional to the measure of maximal entropy.

math.DS

Flat surface models of ergodic systems

We propose a general framework for constructing and describing infinite type flat surfaces of finite area. Using this method, we characterize the range of dynamical behaviors possible for the vertical translation flows on such flat surfaces. We prove a sufficient condition for ergodicity of this flow and apply the condition to several examples. We present specific examples of infinite type flat surfaces on which the translation flow exhibits dynamical phenomena not realizable by translation flows on finite type flat surfaces.

math.DS

Horocycle flow orbits and lattice surface characterizations

The orbit closure of any translation surface under the horocycle flow in almost any direction equals its $SL_2(\mathbb{R})$ orbit closure. This result gives rise to new characterizations of lattice surfaces in terms of the hororcycle flow.

math.DS

Counting invariant components of hyperelliptic translation surfaces

The flow in a fixed direction on a translation surface S determines a decomposition of S into closed invariant sets, each of which is either periodic or minimal. We study this decomposition for translation surfaces in the hyperelliptic connected components $\mathcal{H}^{hyp}(2g-2)$ and $\mathcal{H}^{hyp}(g-1,g-1)$ of the corresponding strata of the moduli space of translation surfaces. Specifically, we characterize the pairs of nonnegative integers (p,m) for which there exists a translation surface in $\mathcal{H}^{hyp}(2g-2)$ or $\mathcal{H}^{hyp}(g-1,g-1)$ with precisely p periodic components and m minimal components. This extends results by Naveh ([Naveh08]), who obtained tight upper bounds on the numbers of minimal components and invariant components a translation surface in any given stratum may have. Analogous results for the other connected components of moduli space are forthcoming.

math.DS

On ergodic transformations that are both weakly mixing and uniformly rigid

We examine some of the properties of uniformly rigid transformations, and analyze the compatibility of uniform rigidity and (measurable) weak mixing along with some of their asymptotic convergence properties. We show that on Cantor space, there does not exist a finite measure-preserving, totally ergodic, uniformly rigid transformation. We briefly discuss general group actions and show that (measurable) weak mixing and uniform rigidity can coexist in a more general setting.

math.DS

Measurable Sensitivity

We introduce the notion of measurable sensitivity, a measure-theoretic version of the condition of sensitive dependence on initial conditions. It is a consequence of light mixing, implies a transformation has only finitely many eigenvalues, and does not exist in the infinite measure-preserving case. Unlike the traditional notion of sensitive dependence, measurable sensitivity carries up to measure-theoretic isomorphism, thus ignoring the behavior of the function on null sets and eliminating dependence on the choice of metric.

math.DS