SearcharxivSearch

arXiv subjects

Felix Voigtlaender

Publications and source records attributed to Felix Voigtlaender.

At least 19 recordsLinked to original sources

Near-optimal estimates for the $\ell^p$-Lipschitz constants of deep random ReLU neural networks

This paper studies the $\ell^p$-Lipschitz constants of ReLU neural networks $Φ: \mathbb{R}^d \to \mathbb{R}$ with random parameters for $p \in [1,\infty]$. The distribution of the weights follows a variant of the He initialization. In the case of zero-bias networks, we derive high probability upper and lower bounds for wide networks that differ at most by a factor that is logarithmic in the network's depth. Remarkably, the behavior of the $\ell^p$-Lipschitz constant varies significantly between the regimes $ p \in [1,2) $ and $ p \in [2,\infty] $. For $p \in [2,\infty]$, the $\ell^p$-Lipschitz constant behaves similarly to $\Vert g\Vert_{p'}$, where $g \in \mathbb{R}^d$ is a $d$-dimensional standard Gaussian vector and $1/p + 1/p' = 1$. In contrast, for $p \in [1,2)$, the $\ell^p$-Lipschitz constant aligns more closely to $\Vert g \Vert_{2}$. We extend our analysis to networks with possibly non-zero biases drawn from arbitrary symmetric distributions. In this case, we obtain high probability upper and lower bounds that differ at most by a factor that is logarithmic in the network's width and linear in its depth.

stat.ML

Wavelet orthonormal bases for nonexpansive dilations

We construct wavelet orthonormal bases for nonexpansive dilations and integer translations that do not admit a wavelet set. In addition, these orthonormal bases do not satisfy the Calderón sum formula. In particular, we show the existence of a wavelet basis for a dilation matrix with determinant one.

math.CA

Maximally Spread Out Measures and Implications for Phase Transitions in Approximation Theory

$\newcommand{\X}{\mathbb{X}}\newcommand{\SC}{\mathcal{C}}$ We establish the existence of a "maximally spread out" Borel probability measure on a totally bounded subset $\SC$ of a (quasi)-Banach space $\X$ under two mild conditions: (i) a growth condition on the covering numbers $N(\SC, ε)$ of $\SC$, and (ii) a technical topological condition that is in particular satisfied whenever $\SC\subset\X$ is closed, bounded, and convex. More formally, condition (i) requires that the so-called lower power-exponential Minkowski dimension of $\SC$, i.e., \[s_\ast:=\liminf_{ε\downarrow 0}\frac{\log\log N(\SC,ε)}{\log(1/ε)}\] satisfies $s_\ast>0$. Under these conditions, we construct a Borel probability measure $μ$ on $\X$ that is critical for $\SC$, or maximally spread out, meaning that the associated outer measure $μ^\ast$ satisfies $μ^\ast(\X\setminus\SC)= 0$ and furthermore satisfies for every $0<s<s_\ast$ the small-ball condition \[μ^\ast(B(x,r))\le\exp\bigl(-c(s)\cdot(1/r)^s\bigr)\quad\text{ for all }x\in\X\text{ and }0<r<r_0(s).\] The existence of such a critical measure in particular implies that the so-called power-exponential Hausdorff dimension of $\SC$ introduced in [J.~Topol.~Anal.~4(2):203--235, 2012] coincides with the lower power-exponential Minkowski dimension. Previous work [Found.~Comput.~Math.~23(1):329--392, 2023] shows that such a critical measure gives rise to a phase transition regarding lossy compression and approximation by quantized neural networks of elements of $\SC$, provided that the $\liminf$ in the definition of $s_\ast$ exists as an actual limit. There, critical measures were constructed for unit balls of certain Besov and Sobolev spaces considered as subsets of $L^2$. In contrast, our construction is completely general. In particular, our results apply to function spaces of dominating mixed smoothness.

math.FA

Linear dependence of time-frequency shifts of a Schwartz function

We show that a finite number of time-frequency shifts of a Schwartz function can be linearly dependent. This disproves the so-called HRT conjecture of Heil, Ramanathan, and Topiwala. In particular, we provide an example consisting of 12 time-frequency shifts.

math.FA

Discrete Triebel-Lizorkin spaces and expansive matrices

We provide a characterization of two expansive dilation matrices yielding equal discrete anisotropic Triebel-Lizorkin spaces. For two such matrices $A$ and $B$, it is shown that $\dot{\mathbf{f}}^α_{p,q}(A) = \dot{\mathbf{f}}^α_{p,q}(B)$ for all $α\in \mathbb{R}$ and $p, q \in (0, \infty]$ if and only if the set $\{A^j B^{-j} : j \in \mathbb{Z}\}$ is finite, or in the trivial case when $p = q$ and $|\det(A)|^{α+ 1/2 - 1/p} = |\det(B)|^{α+ 1/2 - 1/p}$. This provides an extension of a result by Triebel for diagonal dilations to arbitrary expansive matrices. The obtained classification of dilations is different from corresponding results for anisotropic Triebel-Lizorkin function spaces.

math.CA

Upper and lower bounds for the Lipschitz constant of random neural networks

Empirical studies have widely demonstrated that neural networks are highly sensitive to small, adversarial perturbations of the input. The worst-case robustness against these so-called adversarial examples can be quantified by the Lipschitz constant of the neural network. In this paper, we study upper and lower bounds for the Lipschitz constant of random ReLU neural networks. Specifically, we assume that the weights and biases follow a generalization of the He initialization, where general symmetric distributions for the biases are permitted. For deep networks of fixed depth and sufficiently large width, our established upper bound is larger than the lower bound by a factor that is logarithmic in the width. In contrast, for shallow neural networks we characterize the Lipschitz constant up to an absolute numerical constant that is independent of all parameters.

stat.ML

On best approximation by multivariate ridge functions with applications to generalized translation networks

In this paper, we prove sharp upper and lower bounds for the approximation of Sobolev functions by sums of multivariate ridge functions, i.e., for approximation by functions of the form $\mathbb{R}^d \ni x \mapsto \sum_{k=1}^n \varrho_k(A_k x) \in \mathbb{R}$ with $\varrho_k : \mathbb{R}^\ell \to \mathbb{R}$ and $A_k \in \mathbb{R}^{\ell \times d}$. We show that the order of approximation asymptotically behaves as $n^{-r/(d-\ell)}$, where $r$ is the regularity (order of differentiability) of the Sobolev functions to be approximated. Our lower bound even holds when approximating $L^\infty$-Sobolev functions of regularity $r$ with error measured in $L^1$, while our upper bound applies to the approximation of $L^p$-Sobolev functions in $L^p$ for any $1 \leq p \leq \infty$. These bounds generalize well-known results regarding the approximation properties of univariate ridge functions to the multivariate case. We use our results to obtain sharp asymptotic bounds for the approximation of Sobolev functions using generalized translation networks and complex-valued neural networks.

math.FA

Besov regularity of random wavelet series

We study the Besov regularity of wavelet series on $\mathbb{R}^d$ with randomly chosen coefficients. More precisely, each coefficient is a product of a random factor and a parameterized deterministic factor (decaying with the scale $j$ and the norm of the shift $m$). Compared to the literature, we impose relatively mild conditions on the moments of the random variables in order to characterize the almost sure convergence of the wavelet series in Besov spaces $B^s_{p,q}(\mathbb{R}^d)$ and the finiteness of the moments as well as of the moment generating function of the Besov norm. In most cases, we achieve a complete characterization, i.e., the derived conditions are both necessary and sufficient.

math.PR

On wavelet coorbit spaces associated to different dilation groups

This paper develops methods based on coarse geometry for the comparison of wavelet coorbit spaces defined by different dilation groups, with emphasis on establishing a unified approach to both irreducible and reducible quasi-regular representations. We show that the use of reducible representations is essential to include a variety of examples, such as anisotropic Besov spaces defined by general expansive matrices, in a common framework. The obtained criteria yield, among others, a simple characterization of subgroups of a dilation group yielding the same coorbit spaces. They also allow to clarify which anisotropic Besov spaces have an alternative description as coorbit spaces associated to irreducible quasi-regular representations.

math.FA

Smoothness spaces for warped time-frequency representations -- Decomposition spaces and embedding relations

In a recent paper, we have shown that warped time-frequency representations provide a rich framework for the construction and study of smoothness spaces matched to very general phase space geometries obtained by diffeomorphic deformations of $\mathbb{R}^d$. Here, we study these spaces, obtained through the application of general coorbit theory, using the framework of decomposition spaces. This allows us to derive embedding relations between coorbit spaces associated to different warping functions, and relate them to established, important smothness spaces. In particular, we show that we obtain $α$-modulation spaces and spaces of dominating mixed smoothness as special cases and, in contrast, that this is only possible for Besov spaces if $d=1$.

math.FA

Classification of anisotropic local Hardy spaces and inhomogeneous Triebel-Lizorkin spaces

This paper provides a characterization of when two expansive matrices yield the same anisotropic local Hardy and inhomogeneous Triebel-Lizorkin spaces. The characterization is in terms of the coarse equivalence of certain quasi-norms associated to the matrices. For nondiagonal matrices, these conditions are strictly weaker than those classifying the coincidence of the corresponding homogeneous function spaces. The obtained results complete the classification of anisotropic Besov and Triebel-Lizorkin spaces associated to general expansive matrices.

math.CA

Coorbit theory of warped time-frequency systems in $\mathbb{R}^d$

Warped time-frequency systems have recently been introduced as a class of structured continuous frames for functions on the real line. Herein, we generalize this framework to the setting of functions of arbitrary dimensionality. After showing that the basic properties of warped time-frequency representations carry over to higher dimensions, we determine conditions on the warping function which guarantee that the associated Gramian is well-localized, so that associated families of coorbit spaces can be constructed. We then show that discrete Banach frame decompositions for these coorbit spaces can be obtained by sampling the continuous warped time-frequency systems. In particular, this implies that sparsity of a given function $f$ in the discrete warped time-frequency dictionary is equivalent to membership of $f$ in the coorbit space. We put special emphasis on the case of radial warping functions, for which the relevant assumptions simplify considerably.

math.FA

Coorbit spaces associated to quasi-Banach function spaces and their molecular decomposition

This paper provides a self-contained exposition of coorbit spaces associated to integrable group representations and quasi-Banach function spaces, and at the same time extends and simplifies previous work. The main results provide an extension of the theory in [Studia Math., 180(3):237-253, 2007] from groups admitting a compact, conjugation-invariant unit neighborhood to arbitrary (possibly nonunimodular) locally compact groups. In addition, the present paper establishes the existence of molecular dual frames and Riesz sequences as in [J. Funct. Anal., 280(10):56, 2021] for the full scale of quasi-Banach function spaces. The theory is developed for possibly projective and reducible unitary representations in order to be easily applicable to well-studied function spaces not satisfying the classical assumptions of coorbit theory. Compared to the existing literature on quasi-Banach coorbit spaces, all our results apply under significantly weaker integrability conditions on the analyzing vectors, which allows for obtaining sharp results in concrete settings

math.FA

Optimal approximation using complex-valued neural networks

Complex-valued neural networks (CVNNs) have recently shown promising empirical success, for instance for increasing the stability of recurrent neural networks and for improving the performance in tasks with complex-valued inputs, such as in MRI fingerprinting. While the overwhelming success of Deep Learning in the real-valued case is supported by a growing mathematical foundation, such a foundation is still largely lacking in the complex-valued case. We thus analyze the expressivity of CVNNs by studying their approximation properties. Our results yield the first quantitative approximation bounds for CVNNs that apply to a wide class of activation functions including the popular modReLU and complex cardioid activation functions. Precisely, our results apply to any activation function that is smooth but not polyharmonic on some non-empty open set; this is the natural generalization of the class of smooth and non-polynomial activation functions to the complex setting. Our main result shows that the error for the approximation of $C^k$-functions scales as $m^{-k/(2n)}$ for $m \to \infty$ where $m$ is the number of neurons, $k$ the smoothness of the target function and $n$ is the (complex) input dimension. Under a natural continuity assumption, we show that this rate is optimal; we further discuss the optimality when dropping this assumption. Moreover, we prove that the problem of approximating $C^k$-functions using continuous approximation methods unavoidably suffers from the curse of dimensionality.

math.FA

Sampling numbers of smoothness classes via $\ell^1$-minimization

Using techniques developed recently in the field of compressed sensing we prove new upper bounds for general (nonlinear) sampling numbers of (quasi-)Banach smoothness spaces in $L^2$. In particular, we show that in relevant cases such as mixed and isotropic weighted Wiener classes or Sobolev spaces with mixed smoothness, sampling numbers in $L^2$ can be upper bounded by best $n$-term trigonometric widths in $L^\infty$. We describe a recovery procedure from $m$ function values based on $\ell^1$-minimization (basis pursuit denoising). With this method, a significant gain in the rate of convergence compared to recently developed linear recovery methods is achieved. In this deterministic worst-case setting we see an additional speed-up of $m^{-1/2}$ (up to log factors) compared to linear methods in case of weighted Wiener spaces. For their quasi-Banach counterparts even arbitrary polynomial speed-up is possible. Surprisingly, our approach allows to recover mixed smoothness Sobolev functions belonging to $S^r_pW(\mathbb{T}^d)$ on the $d$-torus with a logarithmically better rate of convergence than any linear method can achieve when $1 < p < 2$ and $d$ is large. This effect is not present for isotropic Sobolev spaces.

math.NA

Learning ReLU networks to high uniform accuracy is intractable

Statistical learning theory provides bounds on the necessary number of training samples needed to reach a prescribed accuracy in a learning problem formulated over a given target class. This accuracy is typically measured in terms of a generalization error, that is, an expected value of a given loss function. However, for several applications -- for example in a security-critical context or for problems in the computational sciences -- accuracy in this sense is not sufficient. In such cases, one would like to have guarantees for high accuracy on every input value, that is, with respect to the uniform norm. In this paper we precisely quantify the number of training samples needed for any conceivable training algorithm to guarantee a given uniform accuracy on any learning problem formulated over target classes containing (or consisting of) ReLU neural networks of a prescribed architecture. We prove that, under very general assumptions, the minimal number of training samples for this task scales exponentially both in the depth and the input dimension of the network architecture.

cs.LG

Anisotropic Triebel-Lizorkin spaces and wavelet coefficient decay over one-parameter dilation groups, I

This paper provides maximal function characterizations of anisotropic Triebel-Lizorkin spaces associated to general expansive matrices for the full range of parameters $p \in (0,\infty)$, $q \in (0,\infty]$ and $α\in \mathbb{R}$. The equivalent norm is defined in terms of the decay of wavelet coefficients, quantified by a Peetre-type space over a one-parameter dilation group. As an application, the existence of dual molecular frames and Riesz sequences is obtained; the wavelet systems are generated by translations and anisotropic dilations of a single function, where neither the translation nor dilation parameters are required to belong to a discrete subgroup. Explicit criteria for molecules are given in terms of mild decay, moment, and smoothness conditions.

math.FA

Anisotropic Triebel-Lizorkin spaces and wavelet coefficient decay over one-parameter dilation groups, II

Continuing previous work, this paper provides maximal characterizations of anisotropic Triebel-Lizorkin spaces $\dot{\mathbf{F}}^α_{p,q}$ for the endpoint case of $p = \infty$ and the full scale of parameters $α\in \mathbb{R}$ and $q \in (0,\infty]$. In particular, a Peetre-type characterization of the anisotropic Besov space $\dot{\mathbf{B}}^α_{\infty,\infty} = \dot{\mathbf{F}}^α_{\infty,\infty}$ is obtained. As a consequence, it is shown that there exist dual molecular frames and Riesz sequences in $\dot{\mathbf{F}}^α_{\infty,q}$.

math.FA