SearcharxivSearch

arXiv subjects

Hanne Kekkonen

Publications and source records attributed to Hanne Kekkonen.

15 recordsLinked to original sources

Sharper Regret Bounds for Time-Varying Gaussian Process Bandits with Constant Exploration

We study Bayesian optimization in a time-varying environment where the unknown reward function evolves according to a Gaussian process drift model. Existing GP-UCB analyses in this setting typically require the exploration parameter to grow with the horizon to maintain uniform confidence bounds. Using per-round local confidence events, we show that GP-UCB can instead be run with a constant exploration parameter and obtain an expected-regret bound whose coefficient depends on the drift rate. We also derive a sharper time-varying maximum-information-gain bound. For the squared exponential kernel, it yields $\tildeγ_T/T=\widetilde{\mathcal O}(ε^{1/2})$ and expected average regret $\widetilde{\mathcal O}(ε^{1/4})$ in the persistent-drift regime. The same constant-exploration analysis also yields realized-regret guarantees. Simulations support the predicted logarithmic dependence of the bound-suggested exploration parameter on $1/ε$.

stat.ML

Piecewise Deterministic Markov Processes for Bayesian Inference of PDE Coefficients

We develop a general framework for piecewise deterministic Markov process (PDMP) samplers that enables efficient Bayesian inference in non-linear inverse problems with expensive likelihoods. The key ingredient is a surrogate-assisted thinning scheme in which a surrogate model provides a proposal event rate and a robust correction mechanism enforces an upper bound on the true rate by dynamically adjusting an additive offset whenever violations are detected. This construction is agnostic to the choice of surrogate and PDMP, and we demonstrate it for the Zig-Zag sampler and the Bouncy particle sampler with constant, Laplace, and Gaussian process (GP) surrogates, including gradient-informed and adaptively refined GP variants. As a representative application, we consider Bayesian inference of a spatially varying Young's modulus in a one-dimensional linear elasticity problem. Across dimensions, PDMP samplers equipped with GP-based surrogates achieve substantially higher accuracy and effective sample size per forward model evaluation than Random Walk Metropolis algorithm and the No-U-Turn sampler. The Bouncy particle sampler exhibits the most favorable overall efficiency and scaling, illustrating the potential of the proposed PDMP framework beyond this particular setting.

stat.CO

Random tree Besov priors: Data-driven regularisation parameter selection

We develop a data-driven algorithm for automatically selecting the regularisation parameter in Bayesian inversion under random tree Besov priors. One of the key challenges in Bayesian inversion is the construction of priors that are both expressive and computationally feasible. Random tree Besov priors, introduced in Kekkonen et al. (2023), provide a flexible framework for capturing local regularity properties and sparsity patterns in a wavelet basis. In this paper, we extend this approach by introducing a hierarchical model that enables data-driven selection of the wavelet density parameter, allowing the regularisation strength to adapt across scales while retaining computational efficiency. We focus on nonparametric regression and also present preliminary plug-and-play results for a deconvolution problem.

math.ST

Balancing Accuracy and Speed: A Multi-Fidelity Ensemble Kalman Filter with a Machine Learning Surrogate Model

Currently, more and more machine learning (ML) surrogates are being developed for computationally expensive physical models. In this work we investigate the use of a Multi-Fidelity Ensemble Kalman Filter (MF-EnKF) in which the low-fidelity model is such a machine learning surrogate model, instead of a traditional low-resolution or reduced-order model. The idea behind this is to use an ensemble of a few expensive full model runs, together with an ensemble of many cheap but less accurate ML model runs. In this way we hope to reach increased accuracy within the same computational budget. We investigate the performance by testing the approach on two common test problems, namely the Lorenz-2005 model and the Quasi-Geostrophic model. By keeping the original physical model in place, we obtain a higher accuracy than when we completely replace it by the ML model. Furthermore, the MF-EnKF reaches improved accuracy within the same computational budget. The ML surrogate has similar or improved accuracy compared to the low-resolution one, but it can provide a larger speed-up. Our method contributes to increasing the effective ensemble size in the EnKF, which improves the estimation of the initial condition and hence accuracy of the predictions in fields such as meteorology and oceanography.

cs.LG

Crocheting Mathematics

Crochet provides a superior method for the production of two-dimensional surfaces from one-dimensional material. Compared to any of the other known processes to generate constant flat, spherical or hyperbolic shapes, it is the most flexible and precise way to build a dynamical system with very simple local rules and with very high precision of the intrinsic curvature.

math.HO

Integration of Active Learning and MCMC Sampling for Efficient Bayesian Calibration of Mechanical Properties

Recent advancements in Markov chain Monte Carlo (MCMC) sampling and surrogate modelling have significantly enhanced the feasibility of Bayesian analysis across engineering fields. However, the selection and integration of surrogate models and cutting-edge MCMC algorithms, often depend on ad-hoc decisions. A systematic assessment of their combined influence on analytical accuracy and efficiency is notably lacking. The present work offers a comprehensive comparative study, employing a scalable case study in computational mechanics focused on the inference of spatially varying material parameters, that sheds light on the impact of methodological choices for surrogate modelling and sampling. We show that a priori training of the surrogate model introduces large errors in the posterior estimation even in low to moderate dimensions. We introduce a simple active learning strategy based on the path of the MCMC algorithm that is superior to all a priori trained models, and determine its training data requirements. We demonstrate that the choice of the MCMC algorithm has only a small influence on the amount of training data but no significant influence on the accuracy of the resulting surrogate model. Further, we show that the accuracy of the posterior estimation largely depends on the surrogate model, but not even a tailored surrogate guarantees convergence of the MCMC.Finally, we identify the forward model as the bottleneck in the inference process, not the MCMC algorithm. While related works focus on employing advanced MCMC algorithms, we demonstrate that the training data requirements render the surrogate modelling approach infeasible before the benefits of these gradient-based MCMC algorithms on cheap models can be reaped.

physics.comp-ph

Crocheting Bour's $\mathcal{B}_m$ minimal surfaces

Minimal surfaces can be though as a mathematical generalisation of surfaces formed by soap films. We consider Bour's minimal surfaces $\mathcal{B}_m$ that are intrinsically surfaces of revolution. We show how to generate crochet patterns for $\mathcal{B}_m$ surfaces using basic trigonometric identities to calculate required arc lengths. Three special cases of $\mathcal{B}_m$ surfaces are considered in more detail, namely Enneper's, Richmond's, and Bour's $\mathcal{B}_3$ surfaces, and we provide exact crochet instructions for the classical Enneper's surface.

math.HO

Exploring Mathematics with Curvagon Tiles

Building blocks and tiles are an excellent way of learning about geometry and mathematics in general. There are several versions of tiles that are either snapped together or connected with magnets that can be used to introduce topics like volume, tessellations, and Platonic solids. However, since these tiles are made of hard plastic, they are not very suitable for creating hyperbolic surfaces or shapes where the tiles need to bend. Curvagons are flexible regular polygon building blocks that allow you to quickly build anything from hyperbolic surfaces and tori to dinosaurs and shoes. They can be used to introduce mathematical concepts from Archimedean solids to Gauss-Bonnet theorem. You can also let your imagination run free and build whatever comes to mind.

math.HO

Do the Angles of a Triangle Add up to 180°? -- Introducing Non-Euclidean Geometry

How can we convince students, who have mainly learned to follow given mathematical rules, that mathematics can also be fascinating, creative, and beautiful? In this paper I discuss different ways of introducing non-Euclidean geometry to students and the general public using different physical models, including chalksphere, crocheted hyperbolic surfaces, curved folding, and polygon tilings. Spherical geometry offers a simple yet surprising introduction to the topic, whereas hyperbolic geometry is an entirely new and exciting concept to most. Non-Euclidean geometry demonstrates how crafts and art can be used to make complex mathematical concepts more accessible, and how mathematics itself can be beautiful, not just useful.

math.HO

Consistency of Bayesian inference with Gaussian process priors for a parabolic inverse problem

We consider the statistical nonlinear inverse problem of recovering the absorption term $f>0$ in the heat equation $$ \partial_tu-\frac{1}{2}Δu+fu=0 \quad \text{on $\mathcal{O}\times(0,\textbf{T})$}\quad u = g \quad \text{on $\partial\mathcal{O}\times(0,\textbf{T})$}\quad u(\cdot,0)=u_0 \quad \text{on $\mathcal{O}$}, $$ where $\mathcal{O}\in\mathbb{R}^d$ is a bounded domain, $\textbf{T}<\infty$ is a fixed time, and $g,u_0$ are given sufficiently smooth functions describing boundary and initial values respectively. The data consists of $N$ discrete noisy point evaluations of the solution $u_f$ on $\mathcal{O}\times(0,\textbf{T})$. We study the statistical performance of Bayesian nonparametric procedures based on a large class of Gaussian process priors. We show that, as the number of measurements increases, the resulting posterior distributions concentrate around the true parameter generating the data, and derive a convergence rate for the reconstruction error of the associated posterior means. We also consider the optimality of the contraction rates and prove a lower bound for the minimax convergence rate for inferring $f$ from the data, and show that optimal rates can be achieved with truncated Gaussian priors.

math.ST

Random tree Besov priors -- Towards fractal imaging

We propose alternatives to Bayesian a priori distributions that are frequently used in the study of inverse problems. Our aim is to construct priors that have similar good edge-preserving properties as total variation or Mumford-Shah priors but correspond to well defined infinite-dimensional random variables, and can be approximated by finite-dimensional random variables. We introduce a new wavelet-based model, where the non zero coefficient are chosen in a systematic way so that prior draws have certain fractal behaviour. We show that realisations of this new prior take values in some Besov spaces and have singularities only on a small set $τ$ that has a certain Hausdorff dimension. We also introduce an efficient algorithm for calculating the MAP estimator, arising from the the new prior, in denoising problem.

math.ST

Bernstein-von Mises theorems and uncertainty quantification for linear inverse problems

We consider the statistical inverse problem of recovering an unknown function $f$ from a linear measurement corrupted by additive Gaussian white noise. We employ a nonparametric Bayesian approach with standard Gaussian priors, for which the posterior-based reconstruction of $f$ corresponds to a Tikhonov regulariser $\bar f$ with a reproducing kernel Hilbert space norm penalty. We prove a semiparametric Bernstein-von Mises theorem for a large collection of linear functionals of $f$, implying that semiparametric posterior estimation and uncertainty quantification are valid and optimal from a frequentist point of view. The result is applied to study three concrete examples that cover both the mildly and severely ill-posed cases: specifically, an elliptic inverse problem, an elliptic boundary value problem and the heat equation. For the elliptic boundary value problem, we also obtain a nonparametric version of the theorem that entails the convergence of the posterior distribution to a prior-independent infinite-dimensional Gaussian probability measure with minimal covariance. As a consequence, it follows that the Tikhonov regulariser $\bar f$ is an efficient estimator of $f$, and we derive frequentist guarantees for certain credible balls centred at $\bar{f}$.

math.ST

Large Noise in Variational Regularization

In this paper we consider variational regularization methods for inverse problems with large noise that is in general unbounded in the image space of the forward operator. We introduce a Banach space setting that allows to define a reasonable notion of solutions for more general noise in a larger space provided one has sufficient mapping properties of the forward operators. A key observation, which guides us through the subsequent analysis, is that such a general noise model can be understood with the same setting as approximate source conditions (while a standard model of bounded noise is related directly to classical source conditions). Based on this insight we obtain a quite general existence result for regularized variational problems and derive error estimates in terms of Bregman distances. The latter are specialized for the particularly important cases of one- and p-homogeneous regularization functionals. As a natural further step we study stochastic noise models and in particular white noise, for which we derive error estimates in terms of the expectation of the Bregman distance. The finiteness of certain expectations leads to a novel class of abstract smoothness conditions on the forward operator, which can be easily interpreted in the Hilbert space case. We finally exemplify the approach and in particular the conditions for popular examples of regularization functionals given by squared norm, Besov norm and total variation, respectively.

math.NA

Analysis of regularized inversion of data corrupted by white Gaussian noise

Tikhonov regularization is studied in the case of linear pseudodifferential operator as the forward map and additive white Gaussian noise as the measurement error. The measurement model for an unknown function $u(x)$ is \begin{eqnarray*} m(x) = Au(x) + δ\hspace{.2mm}\varepsilon(x), \end{eqnarray*} where $δ>0$ is the noise magnitude. If $\varepsilon$ was an $L^2$-function, Tikhonov regularization gives an estimate \begin{eqnarray*} T_α(m) = \text{argmin}_{u\in H^r}\big\{\|A u-m\|_{L^2}^2+ α\|u\|_{H^r}^2 \big\}\end{eqnarray*} for $u$ where $α=α(δ)$ is the regularization parameter. Here penalization of the Sobolev norm $ \|u\|_{H^r}$ covers the cases of standard Tikhonov regularization ($r=0$) and first derivative penalty ($r=1$). Realizations of white Gaussian noise are almost never in $L^2$, but do belong to $H^s$ with probability one if $s<0$ is small enough. A modification of Tikhonov regularization theory is presented, covering the case of white Gaussian measurement noise. Furthermore, the convergence of regularized reconstructions to the correct solution as $δ\rightarrow 0$ is proven in appropriate function spaces using microlocal analysis. The convergence of the related finite-dimensional problems to the infinite-dimensional problem is also analysed.

math.AP

Posterior consistency and convergence rates for Bayesian inversion with hypoelliptic operators

Bayesian approach to inverse problems is studied in the case where the forward map is a linear hypoelliptic pseudodifferential operator and measurement error is additive white Gaussian noise. The measurement model for an unknown Gaussian random variable $U(x,ω)$ is \begin{eqnarray*} M(y,ω) = A(U(x,ω) )+ δ\hspace{.2mm}\mathcal{E}(y,ω), \end{eqnarray*} where $A$ is a finitely many times smoothing linear hypoelliptic operator and $δ>0$ is the noise magnitude. The covariance operator $C_U$ of $U$ is $2r$ times smoothing, self-adjoint, injective and elliptic pseudodifferential operator. If $\mathcal{E}$ was taking values in $L^2$ then in Gaussian case solving the conditional mean (and maximum a posteriori) estimate is linked to solving the minimisation problem \begin{eqnarray*} T_δ(M) = \text{argmin}_{u\in H^r} \big\{\|A u-m\|_{L^2}^2+ δ^2\|C_U^{-1/2}u\|_{L^2}^2 \big\}. \end{eqnarray*} However, Gaussian white noise does not take values in $L^2$ but in $H^{-s}$ where $s>0$ is big enough. A modification of the above approach to solve the inverse problem is presented, covering the case of white Gaussian measurement noise. Furthermore, the convergence of conditional mean estimate to the correct solution as $δ\rightarrow 0$ is proven in appropriate function spaces using microlocal analysis. Also the contraction of the confidence regions is studied.

math.ST