Searcharxiv⌕ Search

arXiv subjects

Björn Sprungk

Publications and source records attributed to Björn Sprungk.

18 recordsLinked to original sources

Delayed Acceptance Slice Sampling

Slice sampling is a well-established Markov chain Monte Carlo method for approximate sampling of target distributions which are only known up to a normalizing constant. The method is based on choosing a new state on a slice, i.e., a superlevel set of the given unnormalized target density (with respect to a reference measure). However, slice sampling algorithms usually require per step multiple evaluations of the target density, and thus can become computationally expensive. This is particularly the case for Bayesian inference with costly likelihoods. In this paper, we exploit deterministic approximations of the target density, which are relatively cheap to evaluate, and propose delayed acceptance versions of several common (hybrid) slice samplers. We show ergodicity of the resulting slice sampling methods, discuss the superiority of delayed acceptance (ideal) slice sampling over delayed acceptance Metropolis-Hastings algorithms, and illustrate the benefits of our novel approach in terms of improved computational efficiency in numerical experiments.

stat.CO↗

Weak Poincaré inequality comparisons for ideal and hybrid slice sampling

Using the framework of weak Poincaré inequalities, we provide a general comparison between Hybrid and Ideal Slice Sampling in terms of their corresponding Dirichlet forms. In particular, under suitable assumptions Hybrid Slice Sampling inherits fast convergence from Ideal Slice Sampling and conversely. We apply our results to analyse the convergence of the Independent Metropolis-Hastings, stepping-out and shrinkage, as well as Hit-and-Run within slice sampling algorithms.

stat.CO↗

Metropolis-adjusted interacting particle sampling

In recent years, various interacting particle samplers have been developed to sample from complex target distributions, such as those found in Bayesian inverse problems. These samplers are motivated by the mean-field limit perspective and implemented as ensembles of particles that move in the product state space according to coupled stochastic differential equations. The ensemble approximation and numerical time stepping used to simulate these systems can introduce bias and affect the invariance of the particle system with respect to the target distribution. To correct for this, we investigate the use of a Metropolization step, similar to the Metropolis-adjusted Langevin algorithm. We examine Metropolization of either the whole ensemble or smaller subsets of the ensemble, and prove basic convergence of the resulting ensemble Markov chain to the target distribution. Our numerical results demonstrate the benefits of this correction in numerical examples for popular interacting particle samplers such as ALDI, CBS, and stochastic SVGD.

stat.CO↗

Wasserstein convergence rates of increasingly concentrating probability measures

For $\ell\colon \mathbb{R}^d \to [0,\infty)$ we consider the sequence of probability measures $\left(μ_n\right)_{n \in \mathbb{N}}$, where $μ_n$ is determined by a density that is proportional to $\exp(-n\ell)$. We allow for infinitely many global minimal points of $\ell$, as long as they form a finite union of compact manifolds. In this scenario, we show estimates for the $p$-Wasserstein convergence of $\left(μ_n\right)_{n \in \mathbb{N}}$ to its limit measure. Imposing regularity conditions we obtain a speed of convergence of $n^{-1/(2p)}$ and adding a further technical assumption, we can improve this to a $p$-independent rate of $1/2$ for all orders $p\in\mathbb{N}$ of the Wasserstein distance.

math.PR↗

Wasserstein sensitivity of Risk and Uncertainty Propagation

When propagating uncertainty in the data of differential equations, the probability laws describing the uncertainty are typically themselves subject to uncertainty. We present a sensitivity analysis of uncertainty propagation for differential equations with random inputs to perturbations of the input measures. We focus on the elliptic diffusion equation with random coefficient and source term, for which the probability measure of the solution random field is shown to be Lipschitz-continuous in both total variation and Wasserstein distance. The result generalizes to the solution map of any differential equation with locally Hölder dependence on input parameters. In addition, these results extend to Lipschitz continuous quantities of interest of the solution as well as to coherent risk functionals of these applied to evaluate the impact of their uncertainty. Our analysis is based on the sensitivity of risk functionals and pushforward measures for locally Hölder mappings with respect to the Wasserstein distance of perturbed input distributions. The established results are applied, in particular, to the case of lognormal diffusion and the truncation of series representations of input random fields.

math.PR↗

Robust random walk-like Metropolis-Hastings algorithms for concentrating posteriors

Motivated by Bayesian inference with highly informative data we analyze the performance of random walk-like Metropolis-Hastings algorithms for approximate sampling of increasingly concentrating target distributions. We focus on Gaussian proposals which use a Hessian-based approximation of the target covariance. By means of pushforward transition kernels we show that for Gaussian target measures the spectral gap of the corresponding Metropolis-Hastings algorithm is independent of the concentration of the posterior, i.e., the noise level in the observational data that is used for Bayesian inference. Moreover, by exploiting the convergence of the concentrating posteriors to their Laplace approximation we extend the analysis to non-Gaussian target measures which either concentrate around a single point or along a linear manifold. In particular, in that setting we show that the average acceptance rate as well as the expected squared jump distance of suitable Metropolis-Hastings Markov chains do not deteriorate as the target concentrates.

stat.CO↗

Geometric convergence of elliptical slice sampling

For Bayesian learning, given likelihood function and Gaussian prior, the elliptical slice sampler, introduced by Murray, Adams and MacKay 2010, provides a tool for the construction of a Markov chain for approximate sampling of the underlying posterior distribution. Besides of its wide applicability and simplicity its main feature is that no tuning is necessary. Under weak regularity assumptions on the posterior density we show that the corresponding Markov chain is geometrically ergodic and therefore yield qualitative convergence guarantees. We illustrate our result for Gaussian posteriors as they appear in Gaussian process regression, as well as in a setting of a multi-modal distribution. Remarkably, our numerical experiments indicate a dimension-independent performance of elliptical slice sampling even in situations where our ergodicity result does not apply.

stat.ML↗

On the convergence of adaptive stochastic collocation for elliptic partial differential equations with affine diffusion

Convergence of an adaptive collocation method for the stationary parametric diffusion equation with finite-dimensional affine coefficient is shown. The adaptive algorithm relies on a recently introduced residual-based reliable a posteriori error estimator. For the convergence proof, a strategy recently used for a stochastic Galerkin method with an hierarchical error estimator is transferred to the collocation setting. Extensions to other variants of adaptive collocation methods (including the classical one proposed in the paper "Dimension-adaptive tensor-product quadratuture" Computing (2003) by T. Gerstner and M. Griebel) is explored.

math.NA↗

The linear conditional expectation in Hilbert space

The linear conditional expectation (LCE) provides a best linear (or rather, affine) estimate of the conditional expectation and hence plays an important rôle in approximate Bayesian inference, especially the Bayes linear approach. This article establishes the analytical properties of the LCE in an infinite-dimensional Hilbert space context. In addition, working in the space of affine Hilbert--Schmidt operators, we establish a regularisation procedure for this LCE. As an important application, we obtain a simple alternative derivation and intuitive justification of the conditional mean embedding formula, a concept widely used in machine learning to perform the conditioning of random variables by embedding them into reproducing kernel Hilbert spaces.

math.ST↗

Quantitative spectral gap estimate and Wasserstein contraction of simple slice sampling

We prove Wasserstein contraction of simple slice sampling for approximate sampling w.r.t. distributions with log-concave and rotational invariant Lebesgue densities. This yields, in particular, an explicit quantitative lower bound of the spectral gap of simple slice sampling. Moreover, this lower bound carries over to more general target distributions depending only on the volume of the (super-)level sets of their unnormalized density.

math.PR↗

Stability of doubly-intractable distributions

Doubly-intractable distributions appear naturally as posterior distributions in Bayesian inference frameworks whenever the likelihood contains a normalizing function $Z$. Having two such functions $Z$ and $\widetilde Z$ we provide estimates of the total variation and Wasserstein distance of the resulting posterior probability measures. As a consequence this leads to local Lipschitz continuity w.r.t. $Z$. In the more general framework of a random function $\widetilde Z$ we derive bounds on the expected total variation and expected Wasserstein distance. The applicability of the estimates is illustrated within the setting of two representative Monte Carlo recovery scenarios.

math.ST↗

On the Convergence of the Laplace Approximation and Noise-Level-Robustness of Laplace-based Monte Carlo Methods for Bayesian Inverse Problems

The Bayesian approach to inverse problems provides a rigorous framework for the incorporation and quantification of uncertainties in measurements, parameters and models. We are interested in designing numerical methods which are robust w.r.t. the size of the observational noise, i.e., methods which behave well in case of concentrated posterior measures. The concentration of the posterior is a highly desirable situation in practice, since it relates to informative or large data. However, it can pose a computational challenge for numerical methods based on the prior or reference measure. We propose to employ the Laplace approximation of the posterior as the base measure for numerical integration in this context. The Laplace approximation is a Gaussian measure centered at the maximum a-posteriori estimate and with covariance matrix depending on the logposterior density. We discuss convergence results of the Laplace approximation in terms of the Hellinger distance and analyze the efficiency of Monte Carlo methods based on it. In particular, we show that Laplace-based importance sampling and Laplace-based quasi-Monte-Carlo methods are robust w.r.t. the concentration of the posterior for large classes of posterior distributions and integrands whereas prior-based importance sampling and plain quasi-Monte Carlo are not. Numerical experiments are presented to illustrate the theoretical findings.

math.NA↗

On Expansions and Nodes for Sparse Grid Collocation of Lognormal Elliptic PDEs

This work is a follow-up to our previous contribution ("Convergence of sparse collocation for functions of countably many Gaussian random variables (with application to elliptic PDEs)", SIAM J. Numer. Anal., 2018), and contains further insights on some aspects of the solution of elliptic PDEs with lognormal diffusion coefficients using sparse grids. Specifically, we first focus on the choice of univariate interpolation rules, advocating the use of Gaussian Leja points as introduced by Narayan and Jakeman ("Adaptive Leja sparse grid constructions for stochastic collocation and high-dimensional approximation", SIAM J. Sci. Comput., 2014) and then discuss the possible computational advantages of replacing the standard Karhunen-Loève expansion of the diffusion coefficient with the Lévy-Ciesielski expansion, motivated by theoretical work of Bachmayr, Cohen, DeVore, and Migliorati ("Sparse polynomial approximation of parametric elliptic PDEs. part II: lognormal coefficients", ESAIM: M2AN, 2016). Our numerical results indicate that, for the problem under consideration, Gaussian Leja collocation points outperform Gauss-Hermite and Genz-Keister nodes for the sparse grid approximation and that the Karhunen-Loève expansion of the log diffusion coefficient is more appropriate than its Lévy-Ciesielski expansion for purpose of sparse grid collocation.

math.NA↗

On a Metropolis-Hastings importance sampling estimator

A classical approach for approximating expectations of functions w.r.t. partially known distributions is to compute the average of function values along a trajectory of a Metropolis-Hastings (MH) Markov chain. A key part in the MH algorithm is a suitable acceptance/rejection of a proposed state, which ensures the correct stationary distribution of the resulting Markov chain. However, the rejection of proposals causes highly correlated samples. In particular, when a state is rejected it is not taken any further into account. In contrast to that we consider a MH importance sampling estimator which explicitly incorporates all proposed states generated by the MH algorithm. The estimator satisfies a strong law of large numbers as well as a central limit theorem, and, in addition to that, we provide an explicit mean squared error bound. Remarkably, the asymptotic variance of the MH importance sampling estimator does not involve any correlation term in contrast to its classical counterpart. Moreover, although the analyzed estimator uses the same amount of information as the classical MH estimator, it can outperform the latter in scenarios of moderate dimensions as indicated by numerical experiments.

stat.CO↗

On the Local Lipschitz Stability of Bayesian Inverse Problems

In this note we consider the stability of posterior measures occuring in Bayesian inference w.r.t. perturbations of the prior measure and the log-likelihood function. This extends the well-posedness analysis of Bayesian inverse problems. In particular, we prove a general local Lipschitz continuous dependence of the posterior on the prior and the log-likelihood w.r.t. various common distances of probability measures. These include the total variation, Hellinger, and Wasserstein distance and the Kullback-Leibler divergence. We only assume the boundedness of the likelihoods and measure their perturbations in an $L^p$-norm w.r.t. the prior. The obtained stability yields under mild assumptions the well-posedness of Bayesian inverse problems, in particular, a well-posedness w.r.t. the Wasserstein distance. Moreover, our results indicate an increasing sensitivity of Bayesian inference as the posterior becomes more concentrated, e.g., due to more or more accurate data. This confirms and extends previous observations made in the sensitivity analysis of Bayesian inference.

math.ST↗

Convergence of Sparse Collocation for Functions of Countably Many Gaussian Random Variables (with Application to Elliptic PDEs)

We give a convergence proof for the approximation by sparse collocation of Hilbert-space-valued functions depending on countably many Gaussian random variables. Such functions appear as solutions of elliptic PDEs with lognormal diffusion coefficients. We outline a general $L^2$-convergence theory based on previous work by Bachmayr et al. (2016) and Chen (2016) and establish an algebraic convergence rate for sufficiently smooth functions assuming a mild growth bound for the univariate hierarchical surpluses of the interpolation scheme applied to Hermite polynomials. We verify specifically for Gauss-Hermite nodes that this assumption holds and also show algebraic convergence w.r.t. the resulting number of sparse grid points for this case. Numerical experiments illustrate the dimension-independent convergence rate.

math.NA↗

On a generalization of the preconditioned Crank-Nicolson Metropolis algorithm

Metropolis algorithms for approximate sampling of probability measures on infinite dimensional Hilbert spaces are considered and a generalization of the preconditioned Crank-Nicolson (pCN) proposal is introduced. The new proposal is able to incorporate information of the measure of interest. A numerical simulation of a Bayesian inverse problem indicates that a Metropolis algorithm with such a proposal performs independent of the state space dimension and the variance of the observational noise. Moreover, a qualitative convergence result is provided by a comparison argument for spectral gaps. In particular, it is shown that the generalization inherits geometric ergodicity from the Metropolis algorithm with pCN proposal.

stat.CO↗

Analysis of the Ensemble and Polynomial Chaos Kalman Filters in Bayesian Inverse Problems

We analyze the Ensemble and Polynomial Chaos Kalman filters applied to nonlinear stationary Bayesian inverse problems. In a sequential data assimilation setting such stationary problems arise in each step of either filter. We give a new interpretation of the approximations produced by these two popular filters in the Bayesian context and prove that, in the limit of large ensemble or high polynomial degree, both methods yield approximations which converge to a well-defined random variable termed the analysis random variable. We then show that this analysis variable is more closely related to a specific linear Bayes estimator than to the solution of the associated Bayesian inverse problem given by the posterior measure. This suggests limited or at least guarded use of these generalized Kalman filter methods for the purpose of uncertainty quantification.

math.NA↗