SearcharxivSearch

arXiv · 2303.16813

Optimal approximation using complex-valued neural networks

Abstract

Complex-valued neural networks (CVNNs) have recently shown promising empirical success, for instance for increasing the stability of recurrent neural networks and for improving the performance in tasks with complex-valued inputs, such as in MRI fingerprinting. While the overwhelming success of Deep Learning in the real-valued case is supported by a growing mathematical foundation, such a foundation is still largely lacking in the complex-valued case. We thus analyze the expressivity of CVNNs by studying their approximation properties. Our results yield the first quantitative approximation bounds for CVNNs that apply to a wide class of activation functions including the popular modReLU and complex cardioid activation functions. Precisely, our results apply to any activation function that is smooth but not polyharmonic on some non-empty open set; this is the natural generalization of the class of smooth and non-polynomial activation functions to the complex setting. Our main result shows that the error for the approximation of $C^k$-functions scales as $m^{-k/(2n)}$ for $m \to \infty$ where $m$ is the number of neurons, $k$ the smoothness of the target function and $n$ is the (complex) input dimension. Under a natural continuity assumption, we show that this rate is optimal; we further discuss the optimality when dropping this assumption. Moreover, we prove that the problem of approximating $C^k$-functions using continuous approximation methods unavoidably suffers from the curse of dimensionality.

Explore related subjects

Keep this discovery

BibTeXRIS

Paul Geuchen, Felix Voigtlaender. 2023-03-29. Optimal approximation using complex-valued neural networks. https://arxiv.org/abs/2303.16813

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Typical dynamical properties of operators on $\ell_p$

We investigate the typical dynamical properties of hypercyclic operators in $\mathcal{L}_M(X)$, the set of all bounded linear operators on $X$ whose norms are at most $M$, when $X=\ell_p$, $1< p<\infty$. We show that, with respect to SOT$^*$, a typical operator $T\in \mathcal{L}_M(X)$ is weakly mixing, is weakly disjoint from a given hypercyclic operator $S$, is not topologically ergodic, and satisfies $(T,T^2,\dotsc,T^k)$ is disjoint hypercyclic for any $k\geq 2$. We also study the typical dynamical properties for the concrete family $\mathcal{M}=\{I+B_w\in \mathcal{L}(X)\colon w\in c_0(\mathbb{Z})\}$, endowed with the norm topology, where $B_w$ is a bilateral weighted backward shift.

math.FA

A bi-Lipschitz characterization of strong minimum-attainment for Lipschitz maps

We completely characterize the denseness of strongly minimum-attaining Lipschitz functions, a minimum analogue for strongly norm-attaining Lipschitz functions, in terms of bi-Lipschitz embeddings. More precisely, our main result shows that the set of strongly minimum-attaining Lipschitz functions defined on a complete metric space $M$ fails the denseness if and only if $M$ is bi-Lipschitz equivalent to a subset of $\mathbb{R}$ with positive Lebesgue measure, or equivalently, if $M$ admits a bi-Lipschitz embedding into $\mathbb{R}$ and $M$ has positive 1-dimensional Hausdorff measure. As a consequence, we provide an isometric characterization of the pure 1-unrectifiability of $M$ in terms of strongly minimum-attaining Lipschitz maps defined on bi-Lipschitz copies of closed subsets of $M$. Several counterexamples showing that the main result cannot be naturally extended to the vector-valued setting are also presented.

math.FA

On weak dominance of t-conorms over t-norms

The weak dominance of aggregation operators, particularly between triangular norms (t-norms) and triangular conorms (t-conorms), has attracted considerable attention in aggregation operator theory. While several characterizations have been obtained for Archimedean and continuous cases, a general criterion for continuous t-conorms over continuous t-norms remains to be fully clarified. In this paper, we provide a complete characterization of a continuous t-conorm weakly dominating a continuous t-norm. We first reduce the problem for ordinal sum operators to that for their single Archimedean components, and then express the weak dominance condition entirely in terms of the additive generators of these components. Our approach covers both strict and nilpotent cases uniformly, and recovers the known results for Archimedean operators as a special case.

math.FA