Searcharxiv⌕ Search

arXiv subjects

Max Schölpple

Publications and source records attributed to Max Schölpple.

3 recordsLinked to original sources

Beyond ReLU: How Activations Affect Neural Kernels and Random Wide Networks

In recent years, the neural tangent kernel (NTK) and neural network Gaussian process kernel (NNGP) have given theoreticians tractable limiting cases of fully connected neural networks. However, the property of these kernels are poorly understood for activation functions other than powers of the ReLU. Our main contribution is a characterization of the RKHS of these kernels for activation functions whose only non-smoothness is at zero. This extends existing theory to numerous commonly used activation functions such as SELU, ELU, or LeakyReLU. Additionally, we analyze a broad set of special cases such as missing biases, two-layer networks, or polynomial activations. Our results show that a broad class of not infinitely smooth activations generate equivalent RKHSs at different network depths, depending only on the degree of the non-smoothness up to equivalence. On the other hand, the RKHS generated by polynomial activations depends on the network depth. Finally, we derive results for the smoothness of NNGP sample paths, characterizing the smoothness of infinitely wide neural networks at initialization.

stat.ML↗

Self-Regularized Learning Methods

We introduce a general framework for analyzing learning algorithms based on the notion of self-regularization, which captures implicit complexity control without requiring explicit regularization. This is motivated by previous observations that many algorithms, such as gradient-descent based learning, exhibit implicit regularization. In a nutshell, for a self-regularized algorithm the complexity of the predictor is inherently controlled by that of the simplest comparator achieving the same empirical risk. This framework is sufficiently rich to cover both classical regularized empirical risk minimization and gradient descent. Building on self-regularization, we provide a thorough statistical analysis of such algorithms including minmax-optimal rates, where it suffices to show that the algorithm is self-regularized -- all further requirements stem from the learning problem itself. Finally, we discuss the problem of data-dependent hyperparameter selection, providing a general result which yields minmax-optimal rates up to a double logarithmic factor and covers data-driven early stopping for RKHS-based gradient descent.

stat.ML↗

Which Spaces can be Embedded in Reproducing Kernel Hilbert Spaces?

Given a Banach space $E$ consisting of functions, we ask whether there exists a reproducing kernel Hilbert space $H$ with bounded kernel such that $E\subset H$. More generally, we consider the question, whether for a given Banach space consisting of functions $F$ with $E\subset F$, there exists an intermediate reproducing kernel Hilbert space $E\subset H\subset F$. We provide both sufficient and necessary conditions for this to hold. Moreover, we show that for typical classes of function spaces described by smoothness there is a strong dependence on the underlying dimension: the smoothness $s$ required for the space $E$ needs to grow \emph{proportional} to the dimension $d$ in order to allow for an intermediate reproducing kernel Hilbert space $H$.

math.FA↗