SearcharxivSearch

arXiv subjects

Michael Gnewuch

Publications and source records attributed to Michael Gnewuch.

At least 19 recordsLinked to original sources

Embeddings of Reproducing Kernel Hilbert Spaces with General Weights

We study embeddings between reproducing kernel Hilbert spaces $H(K)$ of functions of $d \in \mathbb{N} \cup \{\infty\}$ variables. The kernels $K$ are superpositions of weighted finite tensor products of a fixed univariate kernel. The basic idea for the embeddings is to compensate a change of the univariate kernel by a suitable transformation of the weights. For the proofs we employ ($d \in \mathbb{N}$) and develop ($d = \infty$) a discrete calculus on the cone of all weights, where completely monotone weights play a particular role. We sketch how to apply the embedding results to computational problems, as, e.g., numerical integration or function recovery.

math.NA

Multi- and Infinite-variate Integration and $L^2$-Approximation on Hilbert Spaces with Gaussian Kernels

We study integration and $L^2$-approximation in the worst-case setting for deterministic linear algorithms based on function evaluations. The underlying function space is a reproducing kernel Hilbert space with a Gaussian kernel of tensor product form. In the infinite-variate case, for both computational problems, we establish matching upper and lower bounds for the polynomial convergence rate of the $n$-th minimal error. In the multivariate case, we improve several tractability results for the integration problem. For the proofs, we establish the following transference result together with an explicit construction: Each of the computational problems on a space with a Gaussian kernel is equivalent on the level of algorithms to the same problem on a Hermite space with suitable parameters.

math.NA

Poissonian pair correlations for dependent random variables

We consider Poissonian pair correlations (PPC) for uniformly distributed sequences of random numbers with a dependency structure. More specifically, we treat two classes of dependent random variables which have widely been studied in the literature, namely sequences of jittered samples and random walks on the torus. We show that for the former class, the PPC property depends on how the finite sample is extended to an infinite sequence. Moreover, we prove that, under some mild assumptions, the random walk on the torus generically has PPC.

math.NT

Data Compression using Rank-1 Lattices for Parameter Estimation in Machine Learning

The mean squared error and regularized versions of it are standard loss functions in supervised machine learning. However, calculating these losses for large data sets can be computationally demanding. Modifying an approach of J. Dick and M. Feischl [Journal of Complexity 67 (2021)], we present algorithms to reduce extensive data sets to a smaller size using rank-1 lattices. Rank-1 lattices are quasi-Monte Carlo (QMC) point sets that are, if carefully chosen, well-distributed in a multidimensional unit cube. The compression strategy in the preprocessing step assigns every lattice point a pair of weights depending on the original data and responses, representing its relative importance. As a result, the compressed data makes iterative loss calculations in optimization steps much faster. We analyze the errors of our QMC data compression algorithms and the cost of the preprocessing step for functions whose Fourier coefficients decay sufficiently fast so that they lie in certain Wiener algebras or Korobov spaces. In particular, we prove that our approach can lead to arbitrary high convergence rates as long as the functions are sufficiently smooth.

math.NA

QMC integration based on arbitrary (t,m,s)-nets yields optimal convergence rates on several scales of function spaces

We study the integration problem over the $s$-dimensional unit cube on four types of Banach spaces of integrands. First we consider Haar wavelet spaces, consisting of functions whose Haar wavelet coefficients exhibit a certain decay behavior measured by a parameter $\alpha >0$. We study the worst case error of integration over the norm unit ball and provide upper error bounds for quasi-Monte Carlo (QMC) cubature rules based on arbitrary $(t,m,s)$-nets as well as matching lower error bounds for arbitrary cubature rules. These results show that using arbitrary $(t,m,s)$-nets as sample points yields the best possible rate of convergence. Afterwards we study spaces of integrands of fractional smoothness $\alpha \in (0,1)$ and state a sharp Koksma-Hlawka-type inequality. More precisely, we show that on those spaces the worst case error of integration is equal to the corresponding fractional discrepancy. Those spaces can be continuously embedded into tensor product Bessel potential spaces, also known as Sobolev spaces of dominated mixed smoothness, with the same set of parameters. The latter spaces can be embedded into suitable Besov spaces of dominating mixed smoothness $\alpha$, which in turn can be embedded into the Haar wavelet spaces with the same set of parameters. Therefore our upper error bounds on Haar wavelet spaces for QMC cubatures based on $(t,m,s)$-nets transfer (with possibly different constants) to the corresponding spaces of integrands of fractional smoothness and to Sobolev and Besov spaces of dominating mixed smoothness. Moreover, known lower error bounds for periodic Sobolev and Besov spaces of dominating mixed smoothness show that QMC integration based on arbitrary $(t,m,s)$-nets yields the best possible convergence rate on periodic as well as on non-periodic Sobolev and Besov spaces of dominating smoothness.

math.NA

Improved bounds for the bracketing number of orthants or revisiting an algorithm of Thi\'{e}mard to compute bounds for the star discrepancy

We improve the best known upper bound for the bracketing number of $d$-dimensional axis-parallel boxes anchored in $0$ (or, put differently, of lower left orthants intersected with the $d$-dimensional unit cube $[0,1]^d$). More precisely, we provide a better upper bound for the cardinality of an algorithmic bracketing cover construction due to Eric Thi\'emard, which forms the core of his algorithm to approximate the star discrepancy of arbitrary point sets from [E. Thi\'emard, An algorithm to compute bounds for the star discrepancy, J.~Complexity 17 (2001), 850 -- 880]. Moreover, the new upper bound for the bracketing number of anchored axis-parallel boxes yields an improved upper bound for the bracketing number of arbitrary axis-parallel boxes in $[0,1]^d$. In our upper bounds all constants are fully explicit.

math.CO

Computable error bounds for quasi-Monte Carlo using points with non-negative local discrepancy

Let $f:[0,1]^d\to\mathbb{R}$ be a completely monotone integrand as defined by Aistleitner and Dick (2015) and let points $\boldsymbol{x}_0,\dots,\boldsymbol{x}_{n-1}\in[0,1]^d$ have a non-negative local discrepancy (NNLD) everywhere in $[0,1]^d$. We show how to use these properties to get a non-asymptotic and computable upper bound for the integral of $f$ over $[0,1]^d$. An analogous non-positive local discrepancy (NPLD) property provides a computable lower bound. It has been known since Gabai (1967) that the two dimensional Hammersley points in any base $b\ge2$ have non-negative local discrepancy. Using the probabilistic notion of associated random variables, we generalize Gabai's finding to digital nets in any base $b\ge2$ and any dimension $d\ge1$ when the generator matrices are permutation matrices. We show that permutation matrices cannot attain the best values of the digital net quality parameter when $d\ge3$. As a consequence the computable absolutely sure bounds we provide come with less accurate estimates than the usual digital net estimates do in high dimensions. We are also able to construct high dimensional rank one lattice rules that are NNLD. We show that those lattices do not have good discrepancy properties: any lattice rule with the NNLD property in dimension $d\ge2$ either fails to be projection regular or has all its points on the main diagonal. Complete monotonicity is a very strict requirement that for some integrands can be mitigated via a control variate.

math.NA

New Bounds for the Extreme and the Star Discrepancy of Double-Infinite Matrices

According to Aistleitner and Weimar, there exist two-dimensional (double) infinite matrices whose star-discrepancy $D_N^{*s}$ of the first $N$ rows and $s$ columns, interpreted as $N$ points in $[0,1]^s$, satisfies an inequality of the form $$D_N^{*s} \leq \sqrt{\alpha} \sqrt{A+B\frac{\ln(\log_2(N))}{s}}\sqrt{\frac{s}{N}}$$ with $\alpha = \zeta^{-1}(2) \approx 1.73, A=1165$ and $B=178$. These matrices are obtained by using i.i.d sequences, and the parameters $s$ and $N$ refer to the dimension and the sample size respectively. In this paper, we improve their result in two directions: First, we change the character of the equation so that the constant $A$ gets replaced by a value $A_s$ dependent on the dimension $s$ such that for $s>1$ we have $A_s<A$. Second, we generalize the result to the case of the (extreme) discrepancy. The paper is complemented by a section where we show numerical results for the dependence of the parameter $A_s$ on $s$.

math.NT

Infinite-dimensional integration and $L^2$-approximation on Hermite spaces

We study integration and $L^2$-approximation of functions of infinitely many variables in the following setting: The underlying function space is the countably infinite tensor product of univariate Hermite spaces and the probability measure is the corresponding product of the standard normal distribution. The maximal domain of the functions from this tensor product space is necessarily a proper subset of the sequence space $\mathbb{R}^\mathbb{N}$. We establish upper and lower bounds for the minimal worst case errors under general assumptions; these bounds do match for tensor products of well-studied Hermite spaces of functions with finite or with infinite smoothness. In the proofs we employ embedding results, and the upper bounds are attained constructively with the help of multivariate decomposition methods.

math.NA

Infinite-Variate $L^2$-Approximation with Nested Subspace Sampling

We consider $L^2$-approximation on weighted reproducing kernel Hilbert spaces of functions depending on infinitely many variables. We focus on unrestricted linear information, admitting evaluations of arbitrary continuous linear functionals. We distinguish between ANOVA and non-ANOVA spaces, where, by ANOVA spaces, we refer to function spaces whose norms are induced by an underlying ANOVA function decomposition. In ANOVA spaces, we provide an optimal algorithm to solve the approximation problem using linear information. We determine the upper and lower error bounds on the polynomial convergence rate of $n$-th minimal worst-case errors, which match if the weights decay regularly. For non-ANOVA spaces, we also establish upper and lower error bounds. Our analysis reveals that for weights with a regular and moderate decay behavior, the convergence rate of $n$-th minimal errors is strictly higher in ANOVA than in non-ANOVA spaces.

math.NA

On Negative Dependence Properties of Latin Hypercube Samples and Scrambled Nets

We study the notion of $γ$-negative dependence of random variables. This notion is a relaxation of the notion of negative orthant dependence (which corresponds to $1$-negative dependence), but nevertheless it still ensures concentration of measure and allows to use large deviation bounds of Chernoff-Hoeffding- or Bernstein-type. We study random variables based on random points $P$. These random variables appear naturally in the analysis of the discrepancy of $P$ or, equivalently, of a suitable worst-case integration error of the quasi-Monte Carlo cubature that uses the points in $P$ as integration nodes. We introduce the correlation number, which is the smallest possible value of $γ$ that ensures $γ$-negative dependence. We prove that the random variables of interest based on Latin hypercube sampling or on $(t,m,d)$-nets do, in general, not have a correlation number of $1$, i.e., they are not negative orthant dependent. But it is known that the random variables based on Latin hypercube sampling in dimension $d$ are actually $γ$-negatively dependent with $γ\le e^d$, and the resulting probabilistic discrepancy bounds do only mildly depend on the $γ$-value.

math.PR

A Generalized Faulhaber Inequality, Improved Bracketing Covers, and Applications to Discrepancy

We prove a generalized Faulhaber inequality to bound the sums of the $j$-th powers of the first $n$ (possibly shifted) natural numbers. With the help of this inequality we are able to improve the known bounds for bracketing numbers of $d$-dimensional axis-parallel boxes anchored in $0$ (or, put differently, of lower left orthants intersected with the $d$-dimensional unit cube $[0,1]^d$). We use these improved bracketing numbers to establish new bounds for the star-discrepancy of negatively dependent random point sets and its expectation. We apply our findings also to the weighted star-discrepancy.

math.CO

Discrepancy Bounds for a Class of Negatively Dependent Random Points Including Latin Hypercube Samples

We introduce a class of $γ$-negatively dependent random samples. We prove that this class includes, apart from Monte Carlo samples, in particular Latin hypercube samples and Latin hypercube samples padded by Monte Carlo. For a $γ$-negatively dependent $N$-point sample in dimension $d$ we provide probabilistic upper bounds for its star discrepancy with explicitly stated dependence on $N$, $d$, and $γ$. These bounds generalize the probabilistic bounds for Monte Carlo samples from [Heinrich et al., Acta Arith. 96 (2001), 279--302] and [C.~Aistleitner, J.~Complexity 27 (2011), 531--540], and they are optimal for Monte Carlo and Latin hypercube samples. In the special case of Monte Carlo samples the constants that appear in our bounds improve substantially on the constants presented in the latter paper and in [C.~Aistleitner, M.~T.~Hofer, Math. Comp.~83 (2014), 1373--1381].

math.ST

Randomized sparse grid algorithms for multivariate integration on Haar-Wavelet spaces

The \emph{deterministic} sparse grid method, also known as Smolyak's method, is a well-established and widely used tool to tackle multivariate approximation problems, and there is a vast literature on it. Much less is known about \emph{randomized} versions of the sparse grid method. In this paper we analyze randomized sparse grid algorithms, namely randomized sparse grid quadratures for multivariate integration on the $D$-dimensional unit cube $[0,1)^D$. Let $d,s \in \mathbb{N}$ be such that $D=d\cdot s$. The $s$-dimensional building blocks of the sparse grid quadratures are based on stratified sampling for $s=1$ and on scrambled $(0,m,s)$-nets for $s\ge 2$. The spaces of integrands and the error criterion we consider are Haar wavelet spaces with parameter $α$ and the randomized error (i.e., the worst case root mean square error), respectively. We prove sharp (i.e., matching) upper and lower bounds for the convergence rates of the $N$-th mininimal errors for all possible combinations of the parameters $d$ and $s$. Our upper error bounds still hold if we consider as spaces of integrands Sobolev spaces of mixed dominated smoothness with smoothness parameters $1/2< α< 1$ instead of Haar wavelet spaces.

math.NA

On Negatively Dependent Sampling Schemes, Variance Reduction, and Probabilistic Upper Discrepancy Bounds

We study some notions of negative dependence of a sampling scheme that can be used to derive variance bounds for the corresponding estimator or discrepancy bounds for the underlying random point set that are at least as good as the corresponding bounds for plain Monte Carlo sampling. We provide new pre-asymptotic bounds with explicit constants for the star discrepancy and the weighted star discrepancy of sampling schemes that satisfy suitable negative dependence properties. Furthermore, we compare the different notions of negative dependence and give several examples of negatively dependent sampling schemes.

math.NA

Explicit error bounds for randomized Smolyak algorithms and an application to infinite-dimensional integration

Smolyak's method, also known as hyperbolic cross approximation or sparse grid method, is a powerful tool to tackle multivariate tensor product problems solely with the help of efficient algorithms for the corresponding univariate problem. In this paper we study the randomized setting, i.e., we randomize Smolyak's method. We provide upper and lower error bounds for randomized Smolyak algorithms with explicitly given dependence on the number of variables and the number of information evaluations used. The error criteria we consider are the worst-case root mean square error (the typical error criterion for randomized algorithms, often referred to as "randomized error") and the root mean square worst-case error (often referred to as "worst-case error"). Randomized Smolyak algorithms can be used as building blocks for efficient methods such as multilevel algorithms, multivariate decomposition methods or dimension-wise quadrature methods to tackle successfully high-dimensional or even infnite-dimensional problems. As an example, we provide a very general and sharp result on the convergence rate of N-th minimal errors of infnite-dimensional integration on weighted reproducing kernel Hilbert spaces. Moreover, we are able to characterize the spaces for which randomized algorithms for infnte-dimensional integration are superior to deterministic ones. We illustrate our fndings for the special instance of weighted Korobov spaces. We indicate how these results can be extended, e.g., to spaces of functions whose smooth dependence on successive variables increases ("spaces of increasing smoothness") and to the problem of L2-approximation (function recovery).

math.NA

Note on Pairwise Negative Dependence of Randomized Rank-1 Lattices

In her recent paper [Negative dependence, scrambled nets, and variance bounds. Math. Oper. Res. 43 (2018), 228-251] Christiane Lemieux studied a framework to analyze the dependence structure of sampling schemes. The main goal of the framework is to determine conditions under which the negative dependence structure of a sampling scheme yields estimators with reduced variance compared to Monte Carlo estimators. For instance, she was able to show that in dimension d = 2 scrambled (0,m, d)-nets lead to randomized quasi-Monte Carlo estimators with variance no larger than the variance of Monte Carlo estimators for functions monotone in each variable. Her result relies on a pairwise negative dependence property that is, in particular, satisfied by (0,m, 2)-nets. In this note we establish that the same result holds true in arbitrary dimension d for a type of randomized lattice point sets that we call randomly shifted and jittered rank-1 lattices. We show that the details of the randomization are crucial and that already small modifications may destroy the pairwise negative dependence property.

math.NA

Probabilistic Lower Bounds for the Discrepancy of Latin Hypercube Samples

We provide probabilistic lower bounds for the star discrepancy of Latin hypercube samples. These bounds are sharp in the sense that they match the recent probabilistic upper bounds for the star discrepancy of Latin hypercube samples proved in [M.~Gnewuch, N.~Hebbinghaus. "Discrepancy bounds for a class of negatively dependent random points including Latin hypercube samples". Preprint 2016.]. Together, this result and our work implies that the discrepancy of Latin hypercube samples differs at most by constant factors from the discrepancy of uniformly sampled point sets.

math.NA