SearcharxivSearch

arXiv subjects

Mathias Oster

Publications and source records attributed to Mathias Oster.

13 recordsLinked to original sources

Preconditioning and Numerical Stability in Neural Network Training for Parametric PDEs

In the context of training neural network-based approximations of solutions of parameter-dependent PDEs, we investigate the effect of preconditioning via well-conditioned frame representations of operators and demonstrate a significant improvement on the performance of standard training methods. We also observe that standard representations of preconditioned matrices are insufficient for obtaining numerical stability and propose a generally applicable form of stable representations that enables computations with single- and half-precision floating point numbers without loss of precision.

math.NA

Anderson Mixing in Bures Wasserstein Space of Gaussian Measures

Various statistical tasks, including sampling or computing Wasserstein barycenters, can be reformulated as fixed-point problems for operators on probability distributions. Accelerating standard fixed-point iteration schemes provides a promising novel approach to the design of efficient numerical methods for these problems. The Wasserstein geometry on the space of probability measures, although not precisely Riemannian, allows us to define various useful Riemannian notions, such as tangent spaces, exponential maps and parallel transport, motivating the adaptation of Riemannian numerical methods. We demonstrate this by developing and implementing the Riemannian Anderson Mixing (RAM) method for Gaussian distributions. The method reuses the history of the residuals and improves the iteration complexity, and we argue that the additional costs, compared to Picard method, are negligible. We show that certain open balls in the Bures-Wasserstein manifold satisfy the requirements for convergence of RAM. The numerical experiments show a significant acceleration compared to a Picard iteration, and performance on par with Riemannian Gradient Descent and Conjugate Gradient methods.

math.OC

Neural network approximation of regularized density functionals

Density functional theory is one of the most efficient and widely used computational methods of quantum mechanics, especially in fields such as solid state physics and quantum chemistry. From the theoretical perspecive, its central object is the universal density functional which contains all intrinsic information about the quantum system in question. Once the external potential is provided, in principle one can obtain the exact ground-state energy via a simple minimization. However, the universal density functional is a very complicated mathematical object and almost always it is replaced with its approximate variants. So far, no ``first principles'', mathematically consistent and convergent approximation procedure has been devised that has general applicability. In this paper, we propose such a procedure by first applying Moreau--Yosida regularization to make the exact functionals continuous (even differentiable) and then approximate the regularized functional by a neural network. The resulting neural network preserves the positivity and convexity of the exact functionals. More importantly, it is differentiable, so it can be directly used in a Kohn--Sham calculation.

physics.chem-ph

Optimization-Free Diffusion Model -- A Perturbation Theory Approach

Diffusion models have emerged as a powerful framework in generative modeling, typically relying on optimizing neural networks to estimate the score function via forward SDE simulations. In this work, we propose an alternative method that is both optimization-free and forward SDE-free. By expanding the score function in a sparse set of eigenbasis of the backward Kolmogorov operator associated with the diffusion process, we reformulate score estimation as the solution to a linear system, avoiding iterative optimization and time-dependent sample generation. We analyze the approximation error using perturbation theory and demonstrate the effectiveness of our method on high-dimensional Boltzmann distributions and real-world datasets.

math.NA

S-SOS: Stochastic Sum-Of-Squares for Parametric Polynomial Optimization

Global polynomial optimization is an important tool across applied mathematics, with many applications in operations research, engineering, and physical sciences. In various settings, the polynomials depend on external parameters that may be random. We discuss a stochastic sum-of-squares (S-SOS) algorithm based on the sum-of squares hierarchy that constructs a series of semidefinite programs to jointly find strict lower bounds on the global minimum and extract candidates for parameterized global minimizers. We prove quantitative convergence of the hierarchy as the degree increases and use it to solve unconstrained and constrained polynomial optimization problems parameterized by random variables. By employing $n$-body priors from condensed matter physics to induce sparsity, we can use S-SOS to produce solutions and uncertainty intervals for sensor network localization problems containing up to 40 variables and semidefinite matrix sizes surpassing $800 \times 800$.

math.OC

Variationally Correct Neural Residual Regression for Parametric PDEs: On the Viability of Controlled Accuracy

This paper is about learning the parameter-to-solution map for systems of partial differential equations (PDEs) that depend on a potentially large number of parameters covering all PDE types for which a stable variational formulation (SVF) can be found. A central constituent is the notion of variationally correct residual loss function meaning that its value is always uniformly proportional to the squared solution error in the norm determined by the SVF, hence facilitating rigorous a posteriori accuracy control. It is based on a single variational problem, associated with the family of parameter dependent fiber problems, employing the notion of direct integrals of Hilbert spaces. Since in its original form the loss function is given as a dual test norm of the residual a central objective is to develop equivalent computable expressions. A first critical role is played by hybrid hypothesis classes, whose elements are piecewise polynomial in (low-dimensional) spatio-temporal variables with parameter-dependent coefficients that can be represented, e.g. by neural networks. Second, working with first order SVFs, we distinguish two scenarios: (i) the test space can be chosen as an $L_2$-space (e.g. for elliptic or parabolic problems) so that residuals live in $L_2$ and can be evaluated directly; (ii) when trial and test spaces for the fiber problems (e.g. for transport equations) depend on the parameters, we use ultraweak formulations. In combination with Discontinuous Petrov Galerkin concepts the hybrid format is then instrumental to arrive at variationally correct computable residual loss functions. Our findings are illustrated by numerical experiments representing (i) and (ii), namely elliptic boundary value problems with piecewise constant diffusion coefficients and pure transport equations with parameter dependent convection field.

math.NA

A comparison study of supervised learning techniques for the approximation of high dimensional functions and feedback control

Approximation of high dimensional functions is in the focus of machine learning and data-based scientific computing. In many applications, empirical risk minimisation techniques over nonlinear model classes are employed. Neural networks, kernel methods and tensor decomposition techniques are among the most popular model classes. We provide a numerical study comparing the performance of these methods on various high-dimensional functions with focus on optimal control problems, where the collection of the dataset is based on the application of the State-Dependent Riccati Equation.

math.NA

Coupled cluster theory: Towards an algebraic geometry formulation

Coupled cluster theory produced arguably the most widely used high-accuracy computational quantum chemistry methods. Despite the approach's overall great computational success, its mathematical understanding is so far limited to results within the realm of functional analysis. The coupled cluster amplitudes, which are the targeted objects in coupled cluster theory, correspond to solutions to the coupled cluster equations, which is a system of polynomial equations of at most degree four. The high dimensionality of the electronic Schr\"odinger equation and the non-linearity of the coupled cluster ansatz have so far stalled a formal analysis of this polynomial system. In this article, we present algebraic investigations that shed light on the coupled cluster equations and the root structure of this ansatz. This is of importance for the a posteriori evaluation of coupled cluster calculations. To that end, we investigate the root structure by means of Newton polytopes. We derive a general v-description, which is subsequently turned into an h-description for explicit examples. This perspective reveals an apparent connection between Pauli's exclusion principle and the geometrical structure of the Newton polytopes. We also propose an alternative characterization of the coupled cluster equations projected onto singles and doubles as cubic polynomials on an algebraic variety with certain sparsity patterns. Moreover, we provide numerical simulations of two computationally tractable systems, namely, the two electrons in four spin-orbitals system and the three electrons in six spin-orbitals system. These simulations provide novel insight into the root structure of the coupled cluster solutions when the coupled cluster ansatz is truncated.

math.AG

Some suggestions concerning the conjecture in: 'Tractable semi-algebraic approximation using Christoffel-Darboux kernel'

In 'Tractable semi-algebraic approximation using Christoffel-Darboux kernel' Marx, Pauwels, Weisser, Henrion and Lasserre conjectured, that the approximation rate $\mathcal O(\frac 1 {\sqrt(d)})$ of a Lipschitz functions by a semi-algebraic function induced by a Christoffel- Darboux kernel of degree $d$ in the $L^1$ norm can be improved for more regular functions. Here we will show, that for semi-algebraic and definable functions the results can be strengthened to a rational approximation rate in the $L^\infty$ norm.

math.CA

Re-entrant tensegrity: A three-periodic, chiral, tensegrity structure that is auxetic

We present a three-periodic, chiral, tensegrity structure and demonstrate that it is auxetic. Our tensegrity structure is constructed using the chiral symmetry {\Pi}+ cylinder packing, transforming the cylinders themselves to the elastic elements and cylinder contacts to incompressible rods. The resulting structure displays local re-entrant geometry at its vertices, and is shown to be auxetic when modelled as an equilibrium configuration of spatial constraints subject to a quasi-static deformation. When the tensegrity structure is subsequently modelled as a lattice material with elastic elements, the auxetic behaviour is again confirmed through finite element modelling as well as demonstrated with a 3d printed example. The cubic symmetry of the structure means that the behaviour is independent of the chosen stretching direction and the auxetic behaviour is observed in both perpendicular directions. This structure could be the simplest three-dimensional analogue to the two-dimensional re-entrant honeycomb. This, alongside the chirality of the structure, make it an interesting design target for multifunctional materials.

cond-mat.mtrl-sci

Approximating optimal feedback controllers of finite horizon control problems using hierarchical tensor formats

Controlling systems of ordinary differential equations (ODEs) is ubiquitous in science and engineering. For finding an optimal feedback controller, the value function and associated fundamental equations such as the Bellman equation and the Hamilton-Jacobi-Bellman (HJB) equation are essential. The numerical treatment of these equations poses formidable challenges due to their non-linearity and their (possibly) high-dimensionality. In this paper we consider a finite horizon control system with associated Bellman equation. After a time-discretization, we obtain a sequence of short time horizon problems which we call local optimal control problems. For solving the local optimal control problems we apply two different methods, one being the well-known policy iteration, where a fixed-point iteration is required for every time step. The other algorithm borrows ideas from Model Predictive Control (MPC), by solving the local optimal control problem via open-loop control methods on a short time horizon, allowing us to replace the fixed-point iteration by an adjoint method. For high-dimensional systems we apply low rank hierarchical tensor product approximation/tree-based tensor formats, in particular tensor trains (TT tensors) and multi-polynomials, together with high-dimensional quadrature, e.g. Monte-Carlo. We prove a linear error propagation with respect to the time discretization and give numerical evidence by controlling a diffusion equation with unstable reaction term and an Allen-Kahn equation.

math.OC

Approximative Policy Iteration for Exit Time Feedback Control Problems driven by Stochastic Differential Equations using Tensor Train format

We consider a stochastic optimal exit time feedback control problem. The Bellman equation is solved approximatively via the Policy Iteration algorithm on a polynomial ansatz space by a sequence of linear equations. As high degree multi-polynomials are needed, the corresponding equations suffer from the curse of dimensionality even in moderate dimensions. We employ tensor-train methods to account for this problem. The approximation process within the Policy Iteration is done via a Least-Squares ansatz and the integration is done via Monte-Carlo methods. Numerical evidences are given for the (multi dimensional) double well potential and a three-hole potential.

math.OC

Approximating the Stationary Bellman Equation by Hierarchical Tensor Products

We treat infinite horizon optimal control problems by solving the associated stationary Hamilton-Jacobi-Bellman (HJB) equation numerically to compute the value function and an optimal feedback law. The dynamical systems under consideration are spatial discretizations of non linear parabolic partial differential equations (PDE), which means that the HJB is non linear and suffers from the curse of dimensionality. Its non linearity is handled by the Policy Iteration algorithm, where the problem is reduced to a sequence of linear, hyperbolic PDEs. These equations remain the computational bottleneck due to their high dimensions. By the method of characteristics these linearized HJB equations can be reformulated via the Koopman operator in the spirit of dynamic programming. The resulting operator equations are solved using a minimal residual method. To overcome numerical infeasability we use low rank hierarchical tensor product approximation/tree-based tensor formats, in particular tensor trains (TT tensors), and multi-polynomials, together with high dimensional quadrature, e.g. Monte-Carlo. By controlling a destabilized version of viscous Burgers and a diffusion equation with unstable reaction term numerical evidences are given.

math.OC