SearcharxivSearch

arXiv subjects

Sebastian Krumscheid

Publications and source records attributed to Sebastian Krumscheid.

At least 19 recordsLinked to original sources

Rethinking Likelihood distributions: Student's t Likelihood Boosts Bayesian Neural Network Performance

In Bayesian neural networks (BNNs), variational inference is a widely adopted framework for modeling uncertainty in a distributional way, with the evidence lower bound (ELBO) serving as the standard objective function. Several distributions contribute to the ELBO loss, such as the prior, approximated posterior, and likelihood distribution. Typically, these distributions are all approximated by a Gaussian distribution, since it is easy to compute, allows for reparameterized gradients, and provides a closed-form loss for training. However, several works have highlighted that this assumption may not generally hold, posing the risk of model misspecification. Alternative distributions have been proposed for the prior specifically, while the effect of distribution choice on the likelihood distribution remains unexplored. In this work, our aim is to close this gap by investigating whether alternative assumptions for the likelihood distribution can outperform the commonly used Gaussian. We compare several likelihood distribution assumptions, such as skewed or heavy-tailed, across regression tasks on both artificial and real-world datasets using standard multilayer perceptrons (MLPs). Our findings demonstrate that Student's t yields better predictive performance than a Gaussian likelihood distribution, independent of the data distribution and MLP architecture (depth and width). In some cases, Student's t can also lead to shorter training times, while still being easy to implement.

cs.LG

A Dynamical Low-rank Multilevel Monte Carlo Estimator for High-Dimensional Kinetic Equations

Kinetic equations are used to model a wide range of phenomena important for real-world applications. Their applications span astrophysics, nuclear physics, engineering, and social sciences. Due to their high-dimensional phase space, modelling and quantifying uncertainties, relevant for applications, poses a significant challenge even for modern computing infrastructure. In recent years, dynamical low-rank approximation (DLRA) has gained popularity for making fine grid simulations of high-dimensional problems feasible by evolving the solution of a time-dependent PDE as a low-rank factorization. This reduces the computational and memory requirements significantly. In this work, we propose a low-rank multilevel Monte Carlo estimator for kinetic equations based on a probabilistic rank-adaptive DLRA time integrator. The level hierarchy of the low-rank multilevel estimator is constructed through spatial refinement and by ensuring that the low-rank error remains below the spatial discretization error. We demonstrate the efficacy of the estimator through several numerical experiments from radiation transport, radiation therapy, and shallow water flow.

math.NA

A Posteriori Error Analysis of Runge-Kutta Discontinuous Galerkin Schemes with SIAC Post-Processing for Nonlinear Convection-Diffusion Systems

We develop reliable a posteriori error estimators for fully discrete Runge-Kutta discontinuous Galerkin approximations of nonlinear convection-diffusion systems endowed with a convex entropy in multiple spatial dimensions on the flat torus T^d, with a focus on the convection-dominated regime. In order to use the relative entropy method, we reconstruct the numerical solution via tensor-product Smoothness-Increasing Accuracy-Conserving (SIAC) filtering which has superconvergence properties. We then derive reliable a posteriori error estimators for the difference between the entropy weak solution and the reconstruction, with constants that are uniform in the vanishing viscosity limit. Our numerical experiments show that the a posteriori error bounds converge with the same order as the error of the reconstructed numerical solution.

math.NA

PurSAMERE: Reliable Adversarial Purification via Sharpness-Aware Minimization of Expected Reconstruction Error

We propose a novel deterministic purification method to improve adversarial robustness by mapping a potentially adversarial sample toward a nearby sample that lies close to a mode of the data distribution, where classifiers are more reliable. We design the method to be deterministic to ensure reliable test accuracy and to prevent the degradation of effective robustness observed in stochastic purification approaches when the adversary has full knowledge of the system and its randomness. We employ a score model trained by minimizing the expected reconstruction error of noise-corrupted data, thereby learning the structural characteristics of the input data distribution. Given a potentially adversarial input, the method searches within its local neighborhood for a purified sample that minimizes the expected reconstruction error under noise corruption and then feeds this purified sample to the classifier. During purification, sharpness-aware minimization is used to guide the purified samples toward flat regions of the expected reconstruction error landscape, thereby enhancing robustness. We further show that, as the noise level decreases, minimizing the expected reconstruction error biases the purified sample toward local maximizers of the Gaussian-smoothed density; under additional local assumptions on the score model, we prove recovery of a local maximizer in the small-noise limit. Experimental results demonstrate significant gains in adversarial robustness over state-of-the-art methods under strong deterministic white-box attacks.

cs.LG

Nonparametric estimation of homogenized invariant measures from multiscale data via Hermite expansion

We consider the problem of density estimation in the context of multiscale Langevin diffusion processes, where a single-scale homogenized surrogate model can be derived. In particular, our aim is to learn the density of the invariant measure of the homogenized dynamics from a continuous-time trajectory generated by the full multiscale system. We propose a spectral method based on a truncated Fourier expansion with Hermite functions as orthonormal basis. The Fourier coefficients are computed directly from the data owing to the ergodic theorem. We prove that the resulting density estimator is robust and converges to the invariant density of the homogenized model as the scale separation parameter vanishes, provided the time horizon and the number of Fourier modes are suitably chosen in relation to the multiscale parameter. The accuracy and reliability of this methodology is further demonstrated through a series of numerical experiments.

math.NA

regTPS-KLE: A Novel Approach To Approximate A Gaussian Random Field for Bayesian Spatial Modeling

Gaussian random field is a ubiquitous model for spatial phenomena in diverse scientific disciplines. Its approximation is often crucial for computational feasibility in simulation, inference, and uncertainty quantification. The Karhunen-Lo\`eve Expansion provides a theoretically optimal basis for representing a Gaussian random field as a sum of deterministic orthonormal functions weighted by uncorrelated random variables. While this is a well-established method for dimension reduction and approximation of (spatial) stochastic processes, its practical application depends on the explicit or implicit definition of the covariance structure. In this work, we propose a novel approach, referred to as regTPS-KLE, for approximating a Gaussian random field by explicitly constructing its covariance via a regularized thin plate spline (TPS) kernel. Because TPS kernels are conditionally positive definite and lack a direct spectral decomposition, we formulate the covariance as the inverse of a regularized elliptic operator. To evaluate its statistical performance, we compare its predictive accuracy and computational efficiency with a Gaussian random field approximation constructed using the stochastic partial differential equations (SPDE) method and implemented within an MCMC algorithm. In simulation studies, the predictive differences between the SPDE and regTPS-KLE models were minimal when the spatial field was generated using Mat\`ern and exponential covariance functions, while regTPS-KLE models consistently outperformed the SPDE approach in terms of computational efficiency. In a real data application, regTPS-KLE exhibits superior predictive accuracy compared with SPDE models based on leave-one-out cross-validation while also achieving improved computational efficiency.

stat.CO

An Investigation into the Distribution of Ratios of Particle Solver-based Likelihoods

We investigate the use of the Metropolis-Hastings algorithm to sample posterior distribution in a Bayesian inverse problem, where the likelihood function is random. Concretely, we consider the case where one has full field observations of a PDE solution, in case a one-dimensional diffusion equation, subject to a Gaussian observation error. Assuming one uses a particle-based Monte Carlo simulation when approximating the likelihood function, one gets an approximate likelihood with additive Gaussian noise in the log-likelihood. We study how these two Gaussian distributions affect the distribution of ratios of approximate likelihood evaluations, as required when evaluating acceptance probabilities in the Metropolis-Hastings algorithm. We do so through both theoretical analysis and numerical experiments.

math.NA

A Minimum Distance Estimator Approach for Misspecified Ergodic Processes

We propose a minimum distance estimator (MDE) for parameter identification in misspecified models characterized by a sequence of ergodic stochastic processes that converge weakly to the model of interest. The data is generated by the sequence of processes, and we are interested in inferring parameters for the limiting processes. We define a general statistical setting for parameter estimation under such model misspecification and prove the robustness of the MDE. Furthermore, we prove the asymptotic normality of the MDE for multiscale diffusion processes with a well-defined homogenized limit. A tractable numerical implementation of the MDE is provided and realized in the programming language Julia.

stat.ME

Limit Theorems for One-Dimensional Homogenized Diffusion Processes

We present two limit theorems, a mean ergodic and a central limit theorem, for a specific class of one-dimensional diffusion processes that depend on a small-scale parameter $\varepsilon$ and converge weakly to a homogenized diffusion process in the limit $\varepsilon \rightarrow 0$. In these results, we allow for the time horizon to blow up such that $T_\varepsilon \rightarrow \infty$ as $\varepsilon \rightarrow 0$. The novelty of the results arises from the circumstance that many quantities are unbounded for $\varepsilon \rightarrow 0$, so that formerly established theory is not directly applicable here and a careful investigation of all relevant $\varepsilon$-dependent terms is required. As a mathematical application, we then use these limit theorems to prove asymptotic properties of a minimum distance estimator for parameters in a homogenized diffusion equation.

math.PR

Cooperative games defined by multi-objective optimization in competition for subsurface resources

We propose a novel decision making framework for forming potential collaboration among otherwise competing agents in subsurface systems. The agents can be, e.g., groundwater, CO$_2$, or hydrogen injectors and extractors with conflicting goals on a geophysically connected system. The operations of a given agent affect the other agents by induced pressure buildup that may jeopardize system integrity. In this work, such a situation is modeled as a cooperative game where the set of agents is partitioned into disjoint coalitions that define the collaborations. The games are in partition function form with externalities, i.e., the value of a coalition depends on both the coalition itself and on the actions of external agents. We investigate the class of cooperative games where the coalition values are the total injection volumes as given by Pareto optimal solutions to multi-objective optimization problems subject to arbitrary physical constraints. For this class of games, we prove that the Pareto set of any coalition structure is a subset of any other coalition structure obtained by splitting coalitions of the first coalition structure. Furthermore, the hierarchical structure of the Pareto sets is used to reduce the computational cost in an algorithm to hierarchically compute the entire Pareto fronts of all possible coalition structures. We demonstrate the framework on a pumping wells groundwater example, and nonlinear and realistic CO$_2$ injection cases, displaying a wide range of possible outcomes. Numerical cost reduction is demonstrated for the proposed algorithm with hierarchically computed Pareto fronts compared to independently solving the multi-objective optimization problems.

math.OC

The Probability of Tiered Benefit: Partial Identification with Robust and Stable Inference

We define the Probability of Tiered Benefit in scenarios with a binary exposure and an outcome that is either categorical with $K \geq 2$ ordered tiers or continuous partitioned by $K-1$ fixed thresholds into disjoint intervals. Similarly to other pure counterfactual queries, this parameter is not $g$-identifiable without additional assumptions. We demonstrate that strong monotonicity does not suffice for point identification when $K \geq 3$ and provide sharp bounds both with and without such constraint. Inference and uncertainty quantification for these bounds are challenging tasks due to potential nonregularity induced by ambiguities in the underlying individualized optimization problems. Such ambiguities can arise from immunities or null treatment effects in subpopulations with positive probability, affecting the lower bound estimate and hindering conservative inference. To address these issues, we extend the available Stabilized One-Step Correction (S1S) procedure by incorporating stratum-specific stabilizing matrices. Through simulations, we illustrate the benefits of this approach over existing alternatives. We apply our method to estimate bounds on the probabilities of tiered benefit and harm from pharmacological treatment for ADHD upon academic achievement, employing observational data from diagnosed Norwegian schoolchildren. Our findings indicate that while girls and children with low prior test performance could have moderate chances of both benefit and harm from treatment, a clear-cut recommendation remains uncertain across all strata.

stat.ME

Low-rank variance reduction for uncertain radiative transfer with control variates

The radiative transfer equation models various physical processes ranging from plasma simulations to radiation therapy. In practice, these phenomena are often subject to uncertainties. Modeling and propagating these uncertainties requires accurate and efficient solvers for the radiative transfer equations. Due to the equation's high-dimensional phase space, fine-grid solutions of the radiative transfer equation are computationally expensive and memory-intensive. In recent years, dynamical low-rank approximation has become a popular method for solving kinetic equations due to the development of computationally inexpensive, memory-efficient and robust algorithms like the augmented basis update \& Galerkin integrator. In this work, we propose a low-rank Monte Carlo estimator and combine it with a control variate strategy based on multi-fidelity low-rank approximations for variance reduction. We investigate the error analytically and numerically and find that a joint approach to balance rank and grid size is necessary. Numerical experiments further show that the efficiency of estimators can be improved using dynamical low-rank approximation, especially in the context of control variates.

math.NA

Reliable Uncertainty Quantification for Fiber Orientation in Composite Molding Processes using Multilevel Polynomial Surrogates

Fiber orientation is decisive for the mechanical performance of composite materials. During manufacturing, variations in material and process parameters can influence fiber orientation. We employ multilevel polynomial surrogates to model the propagation of uncertain material properties in the injection molding process. To ensure reliable uncertainty quantification, a key focus is deriving novel error bounds for statistical measures of a quantity of interest. Numerical experiments employ the Cross-WLF viscosity model and Hagen-Poiseuille flow to investigate the impact of uncertainties in fiber length and matrix temperature on the fractional anisotropy of fiber orientation. The Folgar-Tucker equation and the improved anisotropic rotary diffusion model, incorporating analytical solutions, are used for verification. Results show that the method improves significantly upon standard Monte Carlo estimation, while also providing error guarantees. These findings offer the first step toward a reliable and practical tool for optimizing fiber-reinforced polymer manufacturing processes in the future.

cs.CE

Multilevel randomized quasi-Monte Carlo estimator for nested integration

Nested integration problems arise in various scientific and engineering applications, including Bayesian experimental design, financial risk assessment, and uncertainty quantification. These nested integrals take the form $\int f\left(\int g(\boldsymbol{y},\boldsymbol{x})\mathrm{d}\boldsymbol{x}\right)\mathrm{d}\boldsymbol{y}$, for nonlinear $f$, making them computationally challenging, particularly in high-dimensional settings. Although widely used for single integrals, traditional Monte Carlo (MC) methods can be inefficient when encountering complexities of nested integration. This work introduces a novel multilevel estimator, combining deterministic and randomized quasi-MC (rQMC) methods to handle nested integration problems efficiently. In this context, the inner number of samples and the discretization accuracy of the inner integrand evaluation constitute the level. We provide a comprehensive theoretical analysis of the estimator, deriving error bounds demonstrating significant reductions in bias and variance compared with standard methods. The proposed estimator is particularly effective in scenarios where the integrand is evaluated approximately, as it adapts to different levels of resolution without compromising precision. We verify the performance of our method via numerical experiments, focusing on estimating the expected information gain of experiments. When applied to Gaussian noise in the experiment, a truncation scheme ensures finite error bounds. The results reveal that the proposed multilevel rQMC estimator outperforms existing MC and rQMC approaches, offering a substantial reduction in computational costs and offering a powerful tool for practitioners dealing with complex, nested integration problems across various domains.

math.NA

Non-parametric Inference for Diffusion Processes: A Computational Approach via Bayesian Inversion for PDEs

In this paper, we present a theoretical and computational workflow for the non-parametric Bayesian inference of drift and diffusion functions of autonomous diffusion processes. We base the inference on the partial differential equations arising from the infinitesimal generator of the underlying process. Following a problem formulation in the infinite-dimensional setting, we discuss optimization- and sampling-based solution methods. As preliminary results, we showcase the inference of a single-scale, as well as a multiscale process from trajectory data.

cs.CE

Multi-objective optimization for multi-agent injection strategies in subsurface CO$_2$ storage

We propose a novel framework for optimizing injection strategies in large-scale CO$_2$ storage combining multi-agent models with multi-objective optimization, and reservoir simulation. We investigate whether agents should form coalitions for collaboration to maximize the outcome of their storage activities. In multi-agent systems, it is typically assumed that the optimal strategy for any given coalition structure is already known, and it remains to identify which coalition structure is optimal according to some predefined criterion. For any coalition structure in this work, the optimal CO$_2$ injection strategy is not a priori known, and needs to be found by a combination of reservoir simulation and a multi-objective optimization problem. The multi-objective optimization problems all come with the numerical challenges of repeated evaluations of complex-physics models. We use versatile evolutionary algorithms to solve the multi-objective optimization problems, where the solution is a set of values, e.g., a Pareto front. The Pareto fronts are first computed using the so-called weighted sum method that transforms the multi-objective optimization problem into a set of single-objective optimization problems. Results based on two different Pareto front selection criteria are presented. Then a truly multi-objective optimization method is used to obtain the Pareto fronts, and compared to the previous weighted sum method. We demonstrate the proposed framework on the Bjarmeland formation, a pressure-limited prospective storage site in the Barents Sea. The problem is constrained by the maximum sustainable pressure buildup and a supply of CO$_2$ that can vary over time. In addition to identifying the optimal coalitions, the methodology shows how distinct suboptimal coalitions perform in comparison to the optimum.

math.NA

Capturing the Variability of the Nocturnal Boundary Layer through Localized Perturbation Modeling

A single-column model is used to investigate regime transitions within the stable atmospheric boundary layer, focusing on the role of small-scale fluctuations in wind and temperature dynamics and of turbulence intermittency as triggers for these transitions. Previous studies revealed abrupt near-surface temperature inversion transitions within a limited wind speed range. However, representing these transitions in numerical weather prediction (NWP) and climate models is a known difficulty. To shed light on boundary layer processes that explain these abrupt transitions, the Ekman layer height and its correlation with regime shifts are analyzed. A sensitivity study is performed with several types of perturbations of the wind and temperature tendencies, as well as with the inclusion of intermittent turbulent mixing through a stochastic stability function, to quantify the effect of small fluctuations of the dynamics on regime transitions. The combined results for all tested perturbation types indicate that small-scale phenomena can drive persistent regime transitions from very to weakly stable regimes, but for the opposite direction, no evidence of persistent regime transitions was found. The inclusion of intermittency prevents the model from getting trapped in the very stable regime, thus preventing the so-called "runaway cooling", an issue for commonly used short-tail stability functions. The findings suggest that using stochastic parameterizations of boundary layer processes, either through stochastically perturbed tendencies or parameters, is an effective approach to represent sharp transitions in the boundary layer regimes and is, therefore, a promising avenue to improve the representation of stable boundary layers in NWP and climate models.

physics.ao-ph

Structure-Preserving Operator Learning: Modeling the Collision Operator of Kinetic Equations

This work explores the application of deep operator learning principles to a problem in statistical physics. Specifically, we consider the linear kinetic equation, consisting of a differential advection operator and an integral collision operator, which is a powerful yet expensive mathematical model for interacting particle systems with ample applications, e.g., in radiation transport. We investigate the capabilities of the Deep Operator network (DeepONet) approach to modelling the high dimensional collision operator of the linear kinetic equation. This integral operator has crucial analytical structures that a surrogate model, e.g., a DeepONet, needs to preserve to enable meaningful physical simulation. We propose several DeepONet modifications to encapsulate essential structural properties of this integral operator in a DeepONet model. To be precise, we adapt the architecture of the trunk-net so the DeepONet has the same collision invariants as the theoretical kinetic collision operator, thus preserving conserved quantities, e.g., mass, of the modeled many-particle system. Further, we propose an entropy-inspired data-sampling method tailored to train the modified DeepONet surrogates without requiring an excessive expensive simulation-based data generation.

math.NA