SearcharxivSearch

arXiv subjects

Philipp Wacker

Publications and source records attributed to Philipp Wacker.

28 records · Page 2Linked to original sources

Nested Sampling And Likelihood Plateaus

The main idea of nested sampling is to substitute the high-dimensional likelihood integral over the parameter space $Ω$ by an integral over the unit line $[0,1]$ by employing a push-forward with respect to a suitable transformation. For this substitution, it is often implicitly or explicitly assumed that samples from the prior are uniformly distributed along this unit line after having been mapped by this transformation. We show that this assumption is wrong, especially in the case of a likelihood function with plateaus. Nevertheless, we show that the substitution enacted by nested sampling works because of more interesting reasons which we lay out. Although this means that analytically, nested sampling can deal with plateaus in the likelihood function, the actual performance of the algorithm suffers under such a setting and the method fails to approximate the evidence, mean and variance appropriately. We suggest a robust implementation of nested sampling by a simple decomposition idea which demonstrably overcomes this issue.

math.ST

Pointwise defined version of conditional expectation with respect to a random variable

It is often of interest to condition on a singular event given by a random variable, e.g. $\{Y=y\}$ for a continuous random variable $Y$. Conditional measures with respect to this event are usually derived as a special case of the conditional expectation with respect to the random variables generating sigma algebra. The existence of the latter is usually proven via a non-constructive measure-theoretic argument which yields an only almost-everywhere defined quantity. In particular, the quantity $\mathbb E[f|Y]$ is initially only defined almost everywhere and conditioning on $Y=y$ corresponds to evaluating $\mathbb E[f|Y=y] = \mathbb E[f|Y]{Y=y}$, which is not meaningful because of $\mathbb E[f|Y]$ not being well-defined on such singular sets. This problem is not addressed by the introduction of regular conditional distributions, either. On the other hand it can be shown that the naively computed conditional density $f_{Z|Y=y}(z)$ (which is given by the ratio of joint and marginal densities) is a version of the conditional distribution, i.e. $\mathbb E[\{Z\in B\}|Y=y] = \int_B f_{Z|Y=y}(z) dz$ and this density can indeed be evaluated pointwise in $y$. This mismatch between mathematical theory (which generates an object which cannot produce what we need from it) and practical computation via the conditional density is an unfortunate fact. Furthermore, the classical approach does not allow a pointwise definition of conditional expectations of the form $\mathbb E[f|Y=y]$, only of conditional distributions $\mathbb E[\{Z\in B\}|Y=y]$. We propose a (as far as the author is aware) little known approach to obtaining a pointwise defined version of conditional expectation by use of the Lebesgue-Besicovich lemma without the need of additional topological arguments which are necessary in the usual derivation.

math.PR

On the Convergence of the Laplace Approximation and Noise-Level-Robustness of Laplace-based Monte Carlo Methods for Bayesian Inverse Problems

The Bayesian approach to inverse problems provides a rigorous framework for the incorporation and quantification of uncertainties in measurements, parameters and models. We are interested in designing numerical methods which are robust w.r.t. the size of the observational noise, i.e., methods which behave well in case of concentrated posterior measures. The concentration of the posterior is a highly desirable situation in practice, since it relates to informative or large data. However, it can pose a computational challenge for numerical methods based on the prior or reference measure. We propose to employ the Laplace approximation of the posterior as the base measure for numerical integration in this context. The Laplace approximation is a Gaussian measure centered at the maximum a-posteriori estimate and with covariance matrix depending on the logposterior density. We discuss convergence results of the Laplace approximation in terms of the Hellinger distance and analyze the efficiency of Monte Carlo methods based on it. In particular, we show that Laplace-based importance sampling and Laplace-based quasi-Monte-Carlo methods are robust w.r.t. the concentration of the posterior for large classes of posterior distributions and integrands whereas prior-based importance sampling and plain quasi-Monte Carlo are not. Numerical experiments are presented to illustrate the theoretical findings.

math.NA

Wavelet-based priors accelerate maximum-a-posteriori optimization in Bayesian inverse problems

Wavelet (Besov) priors are a promising way of reconstructing indirectly measured fields in a regularized manner. We demonstrate how wavelets can be used as a localized basis for reconstructing permeability fields with sharp interfaces from noisy pointwise pressure field measurements in the context of the elliptic inverse problem. For this we derive the adjoint method of minimizing the Besov-norm-regularized misfit functional (this corresponds to determining the maximum a posteriori point in the Bayesian point of view) in the Haar wavelet setting. As it turns out, choosing a wavelet--based prior allows for accelerated optimization compared to established trigonometrically--based priors.

math.NA

Well Posedness and Convergence Analysis of the Ensemble Kalman Inversion

The ensemble Kalman inversion is widely used in practice to estimate unknown parameters from noisy measurement data. Its low computational costs, straightforward implementation, and non-intrusive nature makes the method appealing in various areas of application. We present a complete analysis of the ensemble Kalman inversion with perturbed observations for a fixed ensemble size when applied to linear inverse problems. The well-posedness and convergence results are based on the continuous time scaling limits of the method. The resulting coupled system of stochastic differential equations allows to derive estimates on the long-time behaviour and provides insights into the convergence properties of the ensemble Kalman inversion. We view the method as a derivative free optimization method for the least-squares misfit functional, which opens up the perspective to use the method in various areas of applications such as imaging, groundwater flow problems, biological problems as well as in the context of the training of neural networks.

math.NA

A strongly convergent numerical scheme from Ensemble Kalman inversion

The Ensemble Kalman methodology in an inverse problems setting can be viewed as an iterative scheme, which is a weakly tamed discretization scheme for a certain stochastic differential equation (SDE). Assuming a suitable approximation result, dynamical properties of the SDE can be rigorously pulled back via the discrete scheme to the original Ensemble Kalman inversion. The results of this paper make a step towards closing the gap of the missing approximation result by proving a strong convergence result in a simplified model of a scalar stochastic differential equation. We focus here on a toy model with similar properties than the one arising in the context of Ensemble Kalman filter. The proposed model can be interpreted as a single particle filter for a linear map and thus forms the basis for further analysis. The difficulty in the analysis arises from the formally derived limiting SDE with non-globally Lipschitz continuous nonlinearities both in the drift and in the diffusion. Here the standard Euler-Maruyama scheme might fail to provide a strongly convergent numerical scheme and taming is necessary. In contrast to the strong taming usually used, the method presented here provides a weaker form of taming. We present a strong convergence analysis by first proving convergence on a domain of high probability by using a cut-off or localisation, which then leads, combined with bounds on moments for both the SDE and the numerical scheme, by a bootstrapping argument to strong convergence.

math.PR

Laplace's method in Bayesian inverse problems

In a Bayesian inverse problem setting, the solution consists of a posterior measure obtained by combining prior belief, information about the forward operator, and noisy observational data. This measure is most often given in terms of a density with respect to a reference measure in a high-dimensional (or infinite-dimensional) Banach space. Although Monte Carlo sampling methods provide a way of querying the posterior, the necessity of evaluating the forward operator many times (which will often be a costly PDE solver) prohibits this in practice. For this reason, many practitioners choose a suitable Gaussian approximation of the posterior measure, in a procedure called Laplace's method. Once generated, this Gaussian measure is a lot easier to sample from and properties like moments are immediately acquired. This paper derives Laplace's approximation of the posterior measure attributed to the inverse problem explicitly as the posterior measure of a second-order approximation of the data-misfit functional, specifically in the infinite-dimensional setting. By use of a reverse Cauchy-Schwarz inequality we are able to explicitly bound the Hellinger distance between the posterior and its approximation.

math.PR

Probabilistic Estimates of the Maximum Norm of Random Neumann Fourier Series

We study the maximum norm behavior of $L^2$-normalized random Fourier cosine series with a prescribed large wave number. Precise bounds of this type are an important technical tool in estimates for spinodal decomposition, the celebrated phase separation phenomenon in metal alloys. We derive rigorous asymptotic results as the wave number converges to infinity, and shed light on the behavior of the maximum norm for medium range wave numbers through numerical simulations. Finally, we develop a simplified model for describing the magnitude of extremal values of random Neumann Fourier series. The model describes key features of the development of maxima and can be used to predict them. This is achieved by decoupling magnitude and sign distribution, where the latter plays an important role for the study of the size of the maximum norm. Since we are considering series with Neumann boundary conditions, particular care has to be placed on understanding the behavior of the random sums at the boundary.

math.PR

Bayesian model selection for linear regression

In this note we introduce linear regression with basis functions in order to apply Bayesian model selection. The goal is to incorporate Occam's razor as provided by Bayes analysis in order to automatically pick the model optimally able to explain the data without overfitting.

math.ST

Pattern size in Gaussian fields from spinodal decomposition

We study the two-dimensional snake-like pattern that arises in phase separation of alloys described by spinodal decomposition in the Cahn-Hilliard model. These are somewhat universal patterns due to an overlay of eigenfunctions of the Laplacian with a similar wave-number. Similar structures appear in other models like reaction-diffusion systems describing animal coats' patterns or vegetation patterns in desertification. Our main result studies random functions given by cosine Fourier series with independent Gaussian coefficients, that dominate the dynamics in the Cahn-Hilliard model. This is not a cosine process, as the sum is taken over domains in Fourier space that not only grow and scale with a parameter of order $1/\varepsilon$, but also move to infinity. Moreover, the model under consideration is neither stationary nor isotropic. To study the pattern size of nodal domains we consider the density of zeros on any straight line through the spatial domain. Using a theorem by Edelman and Kostlan and weighted ergodic theorems that ensure the convergence of the moving sums, we show that the average distance of zeros is asymptotically of order $\varepsilon$ with a precisely given constant.

math.PR