SearcharxivSearch

arXiv subjects

Jan Vybiral

Publications and source records attributed to Jan Vybiral.

At least 19 recordsLinked to original sources

Nonlocal techniques for the analysis of deep ReLU neural network approximations

Recently, Daubechies, DeVore, Foucart, Hanin, and Petrova introduced a system of piece-wise linear functions, which can be easily reproduced by artificial neural networks with the ReLU activation function and which form a Riesz basis of $L_2([0,1])$. This work was generalized by two of the authors to the multivariate setting. We show that this system serves as a Riesz basis also for Sobolev spaces $W^s([0,1]^d)$ and Barron classes ${\mathbb B}^s([0,1]^d)$ with smoothness $0<s<1$. We apply this fact to re-prove some recent results on the approximation of functions from these classes by deep neural networks. Our proof method avoids using local approximations and allows us to track also the implicit constants as well as to show that we can avoid the curse of dimension. Moreover, we also study how well one can approximate Sobolev and Barron functions by ANNs if only function values are known.

cs.LG

New lower bounds for the integration of periodic functions

We study the integration problem on Hilbert spaces of (multivariate) periodic functions. The standard technique to prove lower bounds for the error of quadrature rules uses bump functions and the pigeon hole principle. Recently, several new lower bounds have been obtained using a different technique which exploits the Hilbert space structure and a variant of the Schur product theorem. The purpose of this paper is to (a) survey the new proof technique, (b) show that it is indeed superior to the bump-function technique, and (c) sharpen and extend the results from the previous papers.

math.NA

Robust network formation with biological applications

We provide new results on the structure of optimal transportation networks obtained as minimizers of an energy cost functional consisting of a kinetic (pumping) and material (metabolic) cost terms, constrained by a local mass conservation law. In particular, we prove that every tree (i.e., graph without loops) represents a local minimizer of the energy with concave metabolic cost. For the linear metabolic cost, we prove that the set of minimizers contains a loop-free structure. Moreover, we enrich the energy functional such that it accounts also for robustness of the network, measured in terms of the Fiedler number of the graph with edge weights given by their conductivities. We examine fundamental properties of the modified functional, in particular, its convexity and differentiability. We provide analytical insights into the new model by considering two simple examples. Subsequently, we employ the projected subgradient method to find global minimizers of the modified functional numerically. We then present two numerical examples, illustrating how the optimal graph's structure and energy expenditure depend on the required robustness of the network.

math.OC

Lower bounds for integration and recovery in $L_2$

Function values are, in some sense, "almost as good" as general linear information for $L_2$-approximation (optimal recovery, data assimilation) of functions from a reproducing kernel Hilbert space. This was recently proved by new upper bounds on the sampling numbers under the assumption that the singular values of the embedding of this Hilbert space into $L_2$ are square-summable. Here we mainly prove new lower bounds. In particular we prove that the sampling numbers behave worse than the approximation numbers for Sobolev spaces with small smoothness. Hence there can be a logarithmic gap also in the case where the singular numbers of the embedding are square-summable. We first prove new lower bounds for the integration problem, again for rather classical Sobolev spaces of periodic univariate functions.

math.NA

Path regularity of Brownian motion and Brownian sheet

By the work of P. Lévy, the sample paths of the Brownian motion are known to satisfy a certain Hölder regularity condition almost surely. This was later improved by Ciesielski, who studied the regularity of these paths in Besov and Besov-Orlicz spaces. We review these results and propose new function spaces of Besov type, strictly smaller than those of Ciesielski and Lévy, where the sample paths of the Brownian motion lie in almost surely. In the same spirit, we review and extend the work of Kamont, who investigated the same question for the multivariate Brownian sheet and function spaces of dominating mixed smoothness.

math.PR

The minimal $k$-dispersion of point sets in high-dimensions

In this manuscript we introduce and study an extended version of the minimal dispersion of point sets, which has recently attracted considerable attention. Given a set $\mathscr P_n=\{x_1,\dots,x_n\}\subset [0,1]^d$ and $k\in\{0,1,\dots,n\}$, we define the $k$-dispersion to be the volume of the largest box amidst a point set containing at most $k$ points. The minimal $k$-dispersion is then given by the infimum over all possible point sets of cardinality $n$. We provide both upper and lower bounds for the minimal $k$-dispersion that coincide with the known bounds for the classical minimal dispersion for a surprisingly large range of $k$'s.

math.NA

Learning general sparse additive models from point queries in high dimensions

We consider the problem of learning a $d$-variate function $f$ defined on the cube $[-1,1]^d\subset {\mathbb R}^d$, where the algorithm is assumed to have black box access to samples of $f$ within this domain. Denote ${\mathcal S}_r \subset {[d] \choose r}; r=1,\dots,r_0$ to be sets consisting of unknown $r$-wise interactions amongst the coordinate variables. We then focus on the setting where $f$ has an additive structure, i.e., it can be represented as $$f = \sum_{{\mathbf j} \in {\mathcal S}_1} \phi_{{\mathbf j}} + \sum_{{\mathbf j} \in {\mathcal S}_2} \phi_{{\mathbf j}} + \dots + \sum_{{\mathbf j} \in {\mathcal S}_{r_0}} \phi_{{\mathbf j}},$$ where each $\phi_{{\mathbf j}}$; ${\mathbf j} \in {\cal S}_r$ is at most $r$-variate for $1 \leq r \leq r_0$. We derive randomized algorithms that query $f$ at carefully constructed set of points, and exactly recover each ${\mathcal S}_r$ with high probability. In contrary to the previous work, our analysis does not rely on numerical approximation of derivatives by finite order differences.

math.NA

Entropy numbers of embeddings of Schatten classes

Let $0<p,q \leq \infty$ and denote by $\mathcal S_p^N$ and $\mathcal S_q^N$ the corresponding finite-dimensional Schatten classes. We prove optimal bounds, up to constants only depending on $p$ and $q$, for the entropy numbers of natural embeddings between $\mathcal S_p^N$ and $\mathcal S_q^N$. This complements the known results in the classical setting of natural embeddings between finite-dimensional $\ell_p$ spaces due to Schütt, Edmunds-Triebel, Triebel and Guédon-Litvak/Kühn. We present a rather short proof that uses all the known techniques as well as a constructive proof of the upper bound in the range $N\leq n\leq N^2$ that allows deeper structural insight and is therefore interesting in its own right. Our main result can also be used to provide an alternative proof of recent lower bounds in the area of low-rank matrix recovery.

math.FA

Learning physical descriptors for materials science by compressed sensing

The availability of big data in materials science offers new routes for analyzing materials properties and functions and achieving scientific understanding. Finding structure in these data that is not directly visible by standard tools and exploitation of the scientific information requires new and dedicated methodology based on approaches from statistical learning, compressed sensing, and other recent methods from applied mathematics, computer science, statistics, signal processing, and information science. In this paper, we explain and demonstrate a compressed-sensing based methodology for feature selection, specifically for discovering physical descriptors, i.e., physical parameters that describe the material and its properties of interest, and associated equations that explicitly and quantitatively describe those relevant properties. As showcase application and proof of concept, we describe how to build a physical model for the quantitative prediction of the crystal structure of binary compound semiconductors.

cond-mat.mtrl-sci

Sparse Proteomics Analysis - A compressed sensing-based approach for feature selection and classification of high-dimensional proteomics mass spectrometry data

Background: High-throughput proteomics techniques, such as mass spectrometry (MS)-based approaches, produce very high-dimensional data-sets. In a clinical setting one is often interested in how mass spectra differ between patients of different classes, for example spectra from healthy patients vs. spectra from patients having a particular disease. Machine learning algorithms are needed to (a) identify these discriminating features and (b) classify unknown spectra based on this feature set. Since the acquired data is usually noisy, the algorithms should be robust against noise and outliers, while the identified feature set should be as small as possible. Results: We present a new algorithm, Sparse Proteomics Analysis (SPA), based on the theory of compressed sensing that allows us to identify a minimal discriminating set of features from mass spectrometry data-sets. We show (1) how our method performs on artificial and real-world data-sets, (2) that its performance is competitive with standard (and widely used) algorithms for analyzing proteomics data, and (3) that it is robust against random and systematic noise. We further demonstrate the applicability of our algorithm to two previously published clinical data-sets.

math.ST

Random Matrices and Matrix Completion

The aim of this note (as well as of the course itself) is to give a largely self-contained proof of two of the main results in the field of low-rank matrix recovery. This field aims for identification of low-rank matrices from only limited linear information exploiting in a crucial way their very special structure. As a crucial tool we develop also the basic statements of the theory of random matrices. The notes are based on a number of sources, which appeared in the last few years. As we give only the minimal amount of the subject needed for the application in mind, the reader is invited to study this further reading in detail.

math.FA

Carl's inequality for quasi-Banach spaces

We prove that for any two quasi-Banach spaces $X$ and $Y$ and any $α>0$ there exists a constant $γ_α>0$ such that $$ \sup_{1\le k\le n}k^αe_k(T)\le γ_α\sup_{1\le k\le n} k^αc_k(T) $$ holds for all linear and bounded operators $T:X\to Y$. Here $e_k(T)$ is the $k$-th entropy number of $T$ and $c_k(T)$ is the $k$-th Gelfand number of $T$. For Banach spaces $X$ and $Y$ this inequality is widely used and well-known as Carl's inequality. For general quasi-Banach spaces it is a new result.

math.FA

Big Data of Materials Science - Critical Role of the Descriptor

Statistical learning of materials properties or functions so far starts with a largely silent, non-challenged step: the choice of the set of descriptive parameters (termed descriptor). However, when the scientific connection between the descriptor and the actuating mechanisms is unclear, causality of the learned descriptor-property relation is uncertain. Thus, trustful prediction of new promising materials, identification of anomalies, and scientific advancement are doubtful. We analyse this issue and define requirements for a suited descriptor. For a classical example, the energy difference of zincblende/wurtzite and rocksalt semiconductors, we demonstrate how a meaningful descriptor can be found systematically.

physics.data-an

On some aspects of approximation of ridge functions

We present effective algorithms for uniform approximation of multivariate functions satisfying some prescribed inner structure. We extend in several directions the analysis of recovery of ridge functions $f(x)=g(\langle a,x\rangle)$ as performed earlier by one of the authors and his coauthors. We consider ridge functions defined on the unit cube $[-1,1]^d$ as well as recovery of ridge functions defined on the unit ball from noisy measurements. We conclude with the study of functions of the type $f(x)=g(\|a-x\|_{l_2^d}^2)$.

math.FA

Complex Interpolation of Weighted Besov- and Lizorkin-Triebel Spaces (long version)

We study complex interpolation of weighted Besov and Lizorkin-Triebel spaces. The used weights $w_0,w_1$ are local Muckenhoupt weights in the sense of Rychkov. As a first step we calculate the Calderón products of associated sequence spaces. Finally, as a corollary of these investigations, we obtain results on complex interpolation of radial subspaces of Besov and Lizorkin-Triebel spaces on $\R^d$.

math.FA

Entropy and sampling numbers of classes of ridge functions

We study properties of ridge functions $f(x)=g(a\cdot x)$ in high dimensions $d$ from the viewpoint of approximation theory. The considered function classes consist of ridge functions such that the profile $g$ is a member of a univariate Lipschitz class with smoothness $\alpha > 0$ (including infinite smoothness), and the ridge direction $a$ has $p$-norm $\|a\|_p \leq 1$. First, we investigate entropy numbers in order to quantify the compactness of these ridge function classes in $L_{\infty}$. We show that they are essentially as compact as the class of univariate Lipschitz functions. Second, we examine sampling numbers and face two extreme cases. In case $p=2$, sampling ridge functions on the Euclidean unit ball faces the curse of dimensionality. It is thus as difficult as sampling general multivariate Lipschitz functions, a result in sharp contrast to the result on entropy numbers. When we additionally assume that all feasible profiles have a first derivative uniformly bounded away from zero in the origin, then the complexity of sampling ridge functions reduces drastically to the complexity of sampling univariate Lipschitz functions. In between, the sampling problem's degree of difficulty varies, depending on the values of $\alpha$ and $p$. Surprisingly, we see almost the entire hierarchy of tractability levels as introduced in the recent monographs by Novak and Wo\'zniakowski.

math.NA

Weak and quasi-polynomial tractability of approximation of infinitely differentiable functions

We comment on recent results in the field of information based complexity, which state (in a number of different settings), that approximation of infinitely differentiable functions is intractable and suffers from the curse of dimensionality. We show that renorming the space of infinitely differentiable functions in a suitable way allows weakly tractable uniform approximation by using only function values. Moreover, the approximating algorithm is based on a simple application of Taylor's expansion at the center of the unit cube. We discuss also the approximation on the Euclidean ball and the approximation in the $L_1$-norm, which is closely related to the problem of numerical integration.

math.NA