SearcharxivSearch

arXiv subjects

Hannes Matt

Publications and source records attributed to Hannes Matt.

4 recordsLinked to original sources

Maximally Spread Out Measures and Implications for Phase Transitions in Approximation Theory

$\newcommand{\X}{\mathbb{X}}\newcommand{\SC}{\mathcal{C}}$ We establish the existence of a "maximally spread out" Borel probability measure on a totally bounded subset $\SC$ of a (quasi)-Banach space $\X$ under two mild conditions: (i) a growth condition on the covering numbers $N(\SC, \epsilon)$ of $\SC$, and (ii) a technical topological condition that is in particular satisfied whenever $\SC\subset\X$ is closed, bounded, and convex. More formally, condition (i) requires that the so-called lower power-exponential Minkowski dimension of $\SC$, i.e., \[s_\ast:=\liminf_{\epsilon\downarrow 0}\frac{\log\log N(\SC,\epsilon)}{\log(1/\epsilon)}\] satisfies $s_\ast>0$. Under these conditions, we construct a Borel probability measure $\mu$ on $\X$ that is critical for $\SC$, or maximally spread out, meaning that the associated outer measure $\mu^\ast$ satisfies $\mu^\ast(\X\setminus\SC)= 0$ and furthermore satisfies for every $0<s<s_\ast$ the small-ball condition \[\mu^\ast(B(x,r))\le\exp\bigl(-c(s)\cdot(1/r)^s\bigr)\quad\text{ for all }x\in\X\text{ and }0<r<r_0(s).\] The existence of such a critical measure in particular implies that the so-called power-exponential Hausdorff dimension of $\SC$ introduced in [J.~Topol.~Anal.~4(2):203--235, 2012] coincides with the lower power-exponential Minkowski dimension. Previous work [Found.~Comput.~Math.~23(1):329--392, 2023] shows that such a critical measure gives rise to a phase transition regarding lossy compression and approximation by quantized neural networks of elements of $\SC$, provided that the $\liminf$ in the definition of $s_\ast$ exists as an actual limit. There, critical measures were constructed for unit balls of certain Besov and Sobolev spaces considered as subsets of $L^2$. In contrast, our construction is completely general. In particular, our results apply to function spaces of dominating mixed smoothness.

math.FA

Linear regression with overparameterized linear neural networks: Tight upper and lower bounds for implicit $\ell^1$-regularization

Modern machine learning models are often trained in a setting where the number of parameters exceeds the number of training samples. To understand the implicit bias of gradient descent in such overparameterized models, prior work has studied diagonal linear neural networks in the regression setting. These studies have shown that, when initialized with small weights, gradient descent tends to favor solutions with minimal $\ell^1$-norm - an effect known as implicit regularization. In this paper, we investigate implicit regularization in diagonal linear neural networks of depth $D\ge 2$ for overparameterized linear regression problems. We focus on analyzing the approximation error between the limit point of gradient flow trajectories and the solution to the $\ell^1$-minimization problem. By deriving tight upper and lower bounds on the approximation error, we precisely characterize how the approximation error depends on the scale of initialization $\alpha$. Our results reveal a qualitative difference between depths: for $D \ge 3$, the error decreases linearly with $\alpha$, whereas for $D=2$, it decreases at rate $\alpha^{1-\varrho}$, where the parameter $\varrho \in [0,1)$ can be explicitly characterized. Interestingly, this parameter is closely linked to so-called null space property constants studied in the sparse recovery literature. We demonstrate the asymptotic tightness of our bounds through explicit examples. Numerical experiments corroborate our theoretical findings and suggest that deeper networks, i.e., $D \ge 3$, may lead to better generalization, particularly for realistic initialization scales.

stat.ML

Universal approximation with complex-valued deep narrow neural networks

We study the universality of complex-valued neural networks with bounded widths and arbitrary depths. Under mild assumptions, we give a full description of those activation functions $\varrho:\mathbb{C}\to \mathbb{C}$ that have the property that their associated networks are universal, i.e., are capable of approximating continuous functions to arbitrary accuracy on compact domains. Precisely, we show that deep narrow complex-valued networks are universal if and only if their activation function is neither holomorphic, nor antiholomorphic, nor $\mathbb{R}$-affine. This is a much larger class of functions than in the dual setting of arbitrary width and fixed depth. Unlike in the real case, the sufficient width differs significantly depending on the considered activation function. We show that a width of $2n+2m+5$ is always sufficient and that in general a width of $max\{2n,2m\}$ is necessary. We prove, however, that a width of $n+m+3$ suffices for a rich subclass of the admissible activation functions. Here, $n$ and $m$ denote the input and output dimensions of the considered networks. Moreover, for the case of smooth and non-polyharmonic activation functions, we provide a quantitative approximation bound in terms of the depth of the considered networks.

math.FA

Banach gradient flows for various families of knot energies

We establish long-time existence of Banach gradient flows for generalised integral Menger curvatures and tangent-point energies, and for O'Hara's self-repulsive potentials $E^{α,p}$. In order to do so, we employ the theory of curves of maximal slope in slightly smaller spaces compactly embedding into the respective energy spaces associated to these functionals, and add a term involving the logarithmic strain, which controls the parametrisations of the flowing (knotted) loops. As a prerequisite, we prove in addition that O'Hara's knot energies $E^{α,p}$ are continuously differentiable.

math.CA