Searcharxiv⌕ Search

arXiv subjects

Sebastian Krumscheid

Publications and source records attributed to Sebastian Krumscheid.

35 records · Page 2Linked to original sources

Structure-Preserving Operator Learning: Modeling the Collision Operator of Kinetic Equations

This work explores the application of deep operator learning principles to a problem in statistical physics. Specifically, we consider the linear kinetic equation, consisting of a differential advection operator and an integral collision operator, which is a powerful yet expensive mathematical model for interacting particle systems with ample applications, e.g., in radiation transport. We investigate the capabilities of the Deep Operator network (DeepONet) approach to modelling the high dimensional collision operator of the linear kinetic equation. This integral operator has crucial analytical structures that a surrogate model, e.g., a DeepONet, needs to preserve to enable meaningful physical simulation. We propose several DeepONet modifications to encapsulate essential structural properties of this integral operator in a DeepONet model. To be precise, we adapt the architecture of the trunk-net so the DeepONet has the same collision invariants as the theoretical kinetic collision operator, thus preserving conserved quantities, e.g., mass, of the modeled many-particle system. Further, we propose an entropy-inspired data-sampling method tailored to train the modified DeepONet surrogates without requiring an excessive expensive simulation-based data generation.

math.NA↗

Copula modeling and uncertainty propagation in field-scale simulation of CO$_2$ fault leakage

Subsurface storage of CO$_2$ is an important means to mitigate climate change, and to investigate the fate of CO$_2$ over several decades in vast reservoirs, numerical simulation based on realistic models is essential. Faults and other complex geological structures introduce modeling challenges as their effects on storage operations are uncertain due to limited data. In this work, we present a computational framework for forward propagation of uncertainty, including stochastic upscaling and copula representation of flow functions for a CO$_2$ storage site using the Vette fault zone in the Smeaheia formation in the North Sea as a test case. The upscaling method leads to a reduction of the number of stochastic dimensions and the cost of evaluating the reservoir model. A viable model that represents the upscaled data needs to capture dependencies between variables, and allow sampling. Copulas provide representation of dependent multidimensional random variables and a good fit to data, allow fast sampling, and coupling to the forward propagation method via independent uniform random variables. The non-stationary correlation within some of the upscaled flow function are accurately captured by a data-driven transformation model. The uncertainty in upscaled flow functions and other parameters are propagated to uncertain leakage estimates using numerical reservoir simulation of a two-phase system. The expectations of leakage are estimated by an adaptive stratified sampling technique, where samples are sequentially concentrated to regions of the parameter space to greedily maximize variance reduction. We demonstrate cost reduction compared to standard Monte Carlo of one or two orders of magnitude for simpler test cases with only fault and reservoir layer permeabilities assumed uncertain, and factors 2--8 cost reduction for stochastic multi-phase flow properties and more complex stochastic models.

math.NA↗

Sequential Estimation using Hierarchically Stratified Domains with Latin Hypercube Sampling

Quantifying the effect of uncertainties in systems where only point evaluations in the stochastic domain but no regularity conditions are available is limited to sampling-based techniques. This work presents an adaptive sequential stratification estimation method that uses Latin Hypercube Sampling within each stratum. The adaptation is achieved through a sequential hierarchical refinement of the stratification, guided by previous estimators using local (i.e., stratum-dependent) variability indicators based on generalized polynomial chaos expansions and Sobol decompositions. For a given total number of samples $N$, the corresponding hierarchically constructed sequence of Stratified Sampling estimators combined with Latin Hypercube sampling is adequately averaged to provide a final estimator with reduced variance. Numerical experiments illustrate the procedure's efficiency, indicating that it can offer a variance decay proportional to $N^{-2}$ in some cases.

stat.ME↗

Scalable method for Bayesian experimental design without integrating over posterior distribution

We address the computational efficiency in solving the A-optimal Bayesian design of experiments problems for which the observational map is based on partial differential equations and, consequently, is computationally expensive to evaluate. A-optimality is a widely used and easy-to-interpret criterion for Bayesian experimental design. This criterion seeks the optimal experimental design by minimizing the expected conditional variance, which is also known as the expected posterior variance. This study presents a novel likelihood-free approach to the A-optimal experimental design that does not require sampling or integrating the Bayesian posterior distribution. The expected conditional variance is obtained via the variance of the conditional expectation using the law of total variance, and we take advantage of the orthogonal projection property to approximate the conditional expectation. We derive an asymptotic error estimation for the proposed estimator of the expected conditional variance and show that the intractability of the posterior distribution does not affect the performance of our approach. We use an artificial neural network (ANN) to approximate the nonlinear conditional expectation in the implementation of our method. We then extend our approach for dealing with the case that the domain of experimental design parameters is continuous by integrating the training process of the ANN into minimizing the expected conditional variance. Through numerical experiments, we demonstrate that our method greatly reduces the number of observation model evaluations compared with widely used importance sampling-based approaches. This reduction is crucial, considering the high computational cost of the observational models. Code is available at https://github.com/vinh-tr-hoang/DOEviaPACE.

math.NA↗

A Fully Parallelized and Budgeted Multi-Level Monte Carlo Method and the Application to Acoustic Waves

We present a novel variant of the multi-level Monte Carlo method that effectively utilizes a reserved computational budget on a high-performance computing system to minimize the mean squared error. Our approach combines concepts of the continuation multi-level Monte Carlo method with dynamic programming techniques following Bellman's optimality principle, and a new parallelization strategy based on a single distributed data structure. Additionally, we establish a theoretical bound on the error reduction on a parallel computing cluster and provide empirical evidence that the proposed method adheres to this bound. We implement, test, and benchmark the approach on computationally demanding problems, focusing on its application to acoustic wave propagation in high-dimensional random media.

math.NA↗

On the equivalence of different adaptive batch size selection strategies for stochastic gradient descent methods

In this study, we demonstrate that the norm test and inner product/orthogonality test presented in \cite{Bol18} are equivalent in terms of the convergence rates associated with Stochastic Gradient Descent (SGD) methods if $ε^2=θ^2+ν^2$ with specific choices of $θ$ and $ν$. Here, $ε$ controls the relative statistical error of the norm of the gradient while $θ$ and $ν$ control the relative statistical error of the gradient in the direction of the gradient and in the direction orthogonal to the gradient, respectively. Furthermore, we demonstrate that the inner product/orthogonality test can be as inexpensive as the norm test in the best case scenario if $θ$ and $ν$ are optimally selected, but the inner product/orthogonality test will never be more computationally affordable than the norm test if $ε^2=θ^2+ν^2$. Finally, we present two stochastic optimization problems to illustrate our results.

math.OC↗

Quantifying uncertain system outputs via the multi-level Monte Carlo method -- distribution and robustness measures

In this work, we consider the problem of estimating the probability distribution, the quantile or the conditional expectation above the quantile, the so called conditional-value-at-risk, of output quantities of complex random differential models by the MLMC method. We follow the approach of (reference), which recasts the estimation of the above quantities to the computation of suitable parametric expectations. In this work, we present novel computable error estimators for the estimation of such quantities, which are then used to optimally tune the MLMC hierarchy in a continuation type adaptive algorithm. We demonstrate the efficiency and robustness of our adaptive continuation-MLMC in an array of numerical test cases.

stat.CO↗

Dynamical Stability Indicator based on Autoregressive Moving-Average Models: Critical Transitions and the Atlantic Meridional Overturning Circulation

A statistical indicator for dynamic stability known as the $Υ$ indicator is used to gauge the stability and hence detect approaching tipping points of simulation data from a reduced 5-box model of the North-Atlantic Meridional Overturning Circulation (AMOC) exposed to a time dependent hosing function. The hosing function simulates the influx of fresh water due to the melting of the Greenland ice sheet and increased precipitation in the North Atlantic. The $Υ$ indicator is designed to detect changes in the memory properties of the dynamics, and is based on fitting ARMA (auto-regressive moving-average) models in a sliding window approach to time series data. An increase in memory properties is interpreted as a sign of dynamical instability. The performance of the indicator is tested on time series subject to different types of tipping, namely bifurcation-induced, noise-induced and rate-induced tipping. The numerical analysis show that the indicator indeed responds to the different types of induced instabilities. Finally, the indicator is applied to two AMOC time series from a full complexity Earth systems model (CESM2). Compared with the doubling CO$_2$ scenario, the quadrupling CO$_2$ scenario results in stronger dynamical instability of the AMOC during its weakening phase.

stat.AP↗

Machine learning-based conditional mean filter: a generalization of the ensemble Kalman filter for nonlinear data assimilation

This paper presents the machine learning-based ensemble conditional mean filter (ML-EnCMF) -- a filtering method based on the conditional mean filter (CMF) previously introduced in the literature. The updated mean of the CMF matches that of the posterior, obtained by applying Bayes' rule on the filter's forecast distribution. Moreover, we show that the CMF's updated covariance coincides with the expected conditional covariance. Implementing the EnCMF requires computing the conditional mean (CM). A likelihood-based estimator is prone to significant errors for small ensemble sizes, causing the filter divergence. We develop a systematical methodology for integrating machine learning into the EnCMF based on the CM's orthogonal projection property. First, we use a combination of an artificial neural network (ANN) and a linear function, obtained based on the ensemble Kalman filter (EnKF), to approximate the CM, enabling the ML-EnCMF to inherit EnKF's advantages. Secondly, we apply a suitable variance reduction technique to reduce statistical errors when estimating loss function. Lastly, we propose a model selection procedure for element-wisely selecting the applied filter, i.e., either the EnKF or ML-EnCMF, at each updating step. We demonstrate the ML-EnCMF performance using the Lorenz-63 and Lorenz-96 systems and show that the ML-EnCMF outperforms the EnKF and the likelihood-based EnCMF.

cs.LG↗

Adaptive stratified sampling for non-smooth problems

Science and engineering problems subject to uncertainty are frequently both computationally expensive and feature nonsmooth parameter dependence, making standard Monte Carlo too slow, and excluding efficient use of accelerated uncertainty quantification methods relying on strict smoothness assumptions. To remedy these challenges, we propose an adaptive stratification method suitable for nonsmooth problems and with significantly reduced variance compared to Monte Carlo sampling. The stratification is iteratively refined and samples are added sequentially to satisfy an allocation criterion combining the benefits of proportional and optimal sampling. Theoretical estimates are provided for the expected performance and probability of failure to correctly estimate essential statistics. We devise a practical adaptive stratification method with strata of the same kind of geometrical shapes, cost-effective refinement satisfying a greedy variance reduction criterion. Numerical experiments corroborate the theoretical findings and exhibit speedups of up to three orders of magnitude compared to standard Monte Carlo sampling.

math.NA↗

Statistical learning of nonlinear stochastic differential equations from non-stationary time series using variational clustering

Parameter estimation for non-stationary stochastic differential equations (SDE) with an arbitrary nonlinear drift, and nonlinear diffusion is accomplished in combination with a non-parametric clustering methodology. Such a model-based clustering approach includes a quadratic programming (QP) problem with equality and inequality constraints. We couple the QP problem to a closed-form likelihood function approach based on suitable Hermite expansion to approximate the parameter values of the SDE model. The classification problem provides a smooth indicator function, which enables us to recover the underlying temporal parameter modulation of the one-dimensional SDE. The numerical examples show that the clustering approach recovers a hidden functional relationship between the SDE model parameters and an additional auxiliary process. The study builds upon this functional relationship to develop closed-form, non-stationary, data-driven stochastic models for multiscale dynamical systems in real-world applications.

math.OC↗

Plateau Proposal Distributions for Adaptive Component-wise Multiple-Try Metropolis

Markov chain Monte Carlo (MCMC) methods are sampling methods that have become a commonly used tool in statistics, for example to perform Monte Carlo integration. As a consequence of the increase in computational power, many variations of MCMC methods exist for generating samples from arbitrary, possibly complex, target distributions. The performance of an MCMC method is predominately governed by the choice of the so-called proposal distribution used. In this paper, we introduce a new type of proposal distribution for the use in MCMC methods that operates component-wise and with multiple trials per iteration. Specifically, the novel class of proposal distributions, called Plateau distributions, do not overlap, thus ensuring that the multiple trials are drawn from different regions of the state space. Furthermore, the Plateau proposal distributions allow for a bespoke adaptation procedure that lends itself to a Markov chain with efficient problem dependent state space exploration and improved burn-in properties. Simulation studies show that our novel MCMC algorithm outperforms competitors when sampling from distributions with a complex shape, highly correlated components or multiple modes.

stat.CO↗

Detecting regime transitions of the nocturnal and Polar near-surface temperature inversion

Many natural systems undergo critical transitions, i.e. sudden shifts from one dynamical regime to another. In the climate system, the atmospheric boundary layer can experience sudden transitions between fully turbulent states and quiescent, quasi-laminar states. Such rapid transitions are observed in Polar regions or at night when the atmospheric boundary layer is stably stratified, and they have important consequences in the strength of mixing with the higher levels of the atmosphere. To analyze the stable boundary layer, many approaches rely on the identification of regimes that are commonly denoted as weakly and very stable regimes. Detecting transitions between the regimes is crucial for modeling purposes. In this work a combination of methods from dynamical systems and statistical modeling is applied to study these regime transitions and to develop an early-warning signal that can be applied to non-stationary field data. The presented metric aims at detecting nearing transitions by statistically quantifying the deviation from the dynamics expected when the system is close to a stable equilibrium. An idealized stochastic model of near-surface inversions is used to evaluate the potential of the metric as an indicator of regime transitions. In this stochastic system, small-scale perturbations can be amplified due to the nonlinearity, resulting in transitions between two possible equilibria of the temperature inversion. The simulations show such noise-induced regime transitions, successfully identified by the indicator. The indicator is further applied to time series data from nocturnal and Polar meteorological measurements.

physics.ao-ph↗

Central limit theorems for multilevel Monte Carlo methods

In this work, we show that uniform integrability is not a necessary condition for central limit theorems (CLT) to hold for normalized multilevel Monte Carlo (MLMC) estimators and we provide near optimal weaker conditions under which the CLT is achieved. In particular, if the variance decay rate dominates the computational cost rate (i.e., $β> γ$), we prove that the CLT applies to the standard (variance minimizing) MLMC estimator. For other settings where the CLT may not apply to the standard MLMC estimator, we propose an alternative estimator, called the mass-shifted MLMC estimator, to which the CLT always applies. This comes at a small efficiency loss: the computational cost of achieving mean square approximation error $\mathcal{O}(ε^2)$ is at worst a factor $\mathcal{O}(\log(1/ε))$ higher with the mass-shifted estimator than with the standard one.

math.PR↗

Perturbation-based inference for diffusion processes: Obtaining effective models from multiscale data

We consider the inference problem for parameters in stochastic differential equation models from discrete time observations (e.g. experimental or simulation data). Specifically, we study the case where one does not have access to observations of the model itself, but only to a perturbed version which converges weakly to the solution of the model. Motivated by this perturbation argument, we study the convergence of estimation procedures from a numerical analysis point of view. More precisely, we introduce appropriate consistency, stability, and convergence concepts and study their connection. It turns out that standard statistical techniques, such as the maximum likelihood estimator, are not convergent methodologies in this setting, since they fail to be stable. Due to this shortcoming, we introduce and analyse a novel inference procedure for parameters in stochastic differential equation models which turns out to be convergent. As such, the method is particularly suited for the estimation of parameters in effective (i.e. coarse-grained) models from observations of the corresponding multiscale process. We illustrate these theoretical findings via several numerical examples.

math.NA↗

A new framework for extracting coarse-grained models from time series with multiscale structure

In many applications it is desirable to infer coarse-grained models from observational data. The observed process often corresponds only to a few selected degrees of freedom of a high-dimensional dynamical system with multiple time scales. In this work we consider the inference problem of identifying an appropriate coarse-grained model from a single time series of a multiscale system. It is known that estimators such as the maximum likelihood estimator or the quadratic variation of the path estimator can be strongly biased in this setting. Here we present a novel parametric inference methodology for problems with linear parameter dependency that does not suffer from this drawback. Furthermore, we demonstrate through a wide spectrum of examples that our methodology can be used to derive appropriate coarse-grained models from time series of partial observations of a multiscale system in an effective and systematic fashion.

math.ST↗

Semi-Parametric Drift and Diffusion Estimation for Multiscale Diffusions

We consider the problem of statistical inference for the effective dynamics of multiscale diffusion processes with (at least) two widely separated characteristic time scales. More precisely, we seek to determine parameters in the effective equation describing the dynamics on the longer diffusive time scale, i.e. in a homogenization framework. We examine the case where both the drift and the diffusion coefficients in the effective dynamics are space-dependent and depend on multiple unknown parameters. It is known that classical estimators, such as Maximum Likelihood and Quadratic Variation of the Path Estimators, fail to obtain reasonable estimates for parameters in the effective dynamics when based on observations of the underlying multiscale diffusion. We propose a novel algorithm for estimating both the drift and diffusion coefficients in the effective dynamics based on a semi-parametric framework. We demonstrate by means of extensive numerical simulations of a number of selected examples that the algorithm performs well when applied to data from a multiscale diffusion. These examples also illustrate that the algorithm can be used effectively to obtain accurate and unbiased estimates.

math.ST↗