Searcharxiv⌕ Search

arXiv subjects

Lisa Maria Kreusser

Publications and source records attributed to Lisa Maria Kreusser.

At least 19 recordsLinked to original sources

Equidistribution-based training of Free Knot Splines and ReLU Neural Networks

We consider the problem of univariate nonlinear function approximation using shallow neural networks (NN) with a rectified linear unit (ReLU) activation function. We show that the $L_2$ based approximation problem is ill-conditioned and the behaviour of optimisation algorithms used in training these networks degrades rapidly as the width of the network increases. This can lead to significantly poorer approximation in practice than expected from the theoretical expressivity of the ReLU architecture and traditional methods such as univariate Free Knot Splines (FKS). Univariate shallow ReLU NNs and FKS span the same function space, and thus have the same theoretical expressivity. However, the FKS representation remains well-conditioned as the number of knots increases. We leverage the theory of optimal piecewise linear interpolants to improve the training procedure for ReLU NNs. Using the equidistribution principle, we propose a two-level procedure for training the FKS by first solving the nonlinear problem of finding the optimal knot locations of the interpolating FKS, and then determine the optimal weights and knots of the FKS by solving a nearly linear, well-conditioned problem. The training of the FKS gives insights into how we can train a ReLU NN effectively, with an equally accurate approximation. We combine the training of the ReLU NN with an equidistribution-based loss to find the breakpoints of the ReLU functions. This is then combined with preconditioning the ReLU NN approximation to find the scalings of the ReLU functions. This fast, well-conditioned and reliable method finds an accurate shallow ReLU NN approximation to a univariate target function. We test this method on a series of regular, singular, and rapidly varying target functions and obtain good results, realising the expressivity of the shallow ReLU network in all cases. We then extend our results to deeper networks.

cs.LG↗

Models for information propagation on graphs

We propose and unify classes of different models for information propagation over graphs. In a first class, propagation is modelled as a wave which emanates from a set of \emph{known} nodes at an initial time, to all other \emph{unknown} nodes at later times with an ordering determined by the arrival time of the information wave front. A second class of models is based on the notion of a travel time along paths between nodes. The time of information propagation from an initial \emph{known} set of nodes to a node is defined as the minimum of a generalised travel time over subsets of all admissible paths. A final class is given by imposing a local equation of an eikonal form at each \emph{unknown} node, with boundary conditions at the \emph{known} nodes. The solution value of the local equation at a node is coupled to those of neighbouring nodes with lower values. We provide precise formulations of the model classes and prove equivalences between them. Finally we apply the front propagation models on graphs to semi-supervised learning via label propagation and information propagation on trust networks.

math.NA↗

Parallel-in-Time Solutions with Random Projection Neural Networks

This paper considers one of the fundamental parallel-in-time methods for the solution of ordinary differential equations, Parareal, and extends it by adopting a neural network as a coarse propagator. We provide a theoretical analysis of the convergence properties of the proposed algorithm and show its effectiveness for several examples, including Lorenz and Burgers' equations. In our numerical simulations, we further specialize the underpinning neural architecture to Random Projection Neural Networks (RPNNs), a 2-layer neural network where the first layer weights are drawn at random rather than optimized. This restriction substantially increases the efficiency of fitting RPNN's weights in comparison to a standard feedforward network without negatively impacting the accuracy, as demonstrated in the SIR system example.

math.NA↗

A Rule-Based Epidemiological Modelling Framework

Motivated by chemical reaction rules, we introduce a rule-based epidemiological framework for the systematic mathematical modelling of future pandemics. Here we stress that we do not have a specific model in mind, but a whole collection of models which can be transformed into each other, or represent different aspects of a pandemic, and these aspects can change during the course of the emergency, as happened during the Covid-19 pandemic. As conditions for outbreaks in the modern world change on different time-scales, some rapidly, epidemiology has few 'laws', besides perhaps the fundamental infection process described by Kermack-McKendrick. Each single of our variety of models, called framework, is based on a mathematical formulation that we call a rule-based system. They have several advantages, for example that they can be both interpreted stochastically and deterministically, without changing the model structure. Rule-based systems should be easier to communicate to non-specialists, when compared to differential equations. Due to their combinatorial nature, the rule-based model framework we propose is ideal for systematic mathematical modelling, systematic links to statistics, data analysis in general and also machine learning leading to artificial intelligence.

q-bio.PE↗

A general formulation of reweighted least squares fitting

We present a generalized formulation for reweighted least squares approximations. The goal of this article is twofold: firstly, to prove that the solution of such problem can be expressed as a convex combination of certain interpolants when the solution is sought in any finite-dimensional vector space; secondly, to provide a general strategy to iteratively update the weights according to the approximation error and apply it to the spline fitting problem. In the experiments, we provide numerical examples for the case of polynomials and splines spaces. Subsequently, we evaluate the performance of our fitting scheme for spline curve and surface approximation, including adaptive spline constructions.

math.NA↗

Closing the ODE-SDE gap in score-based diffusion models through the Fokker-Planck equation

Score-based diffusion models have emerged as one of the most promising frameworks for deep generative modelling, due to their state-of-the art performance in many generation tasks while relying on mathematical foundations such as stochastic differential equations (SDEs) and ordinary differential equations (ODEs). Empirically, it has been reported that ODE based samples are inferior to SDE based samples. In this paper we rigorously describe the range of dynamics and approximations that arise when training score-based diffusion models, including the true SDE dynamics, the neural approximations, the various approximate particle dynamics that result, as well as their associated Fokker--Planck equations and the neural network approximations of these Fokker--Planck equations. We systematically analyse the difference between the ODE and SDE dynamics of score-based diffusion models, and link it to an associated Fokker--Planck equation. We derive a theoretical upper bound on the Wasserstein 2-distance between the ODE- and SDE-induced distributions in terms of a Fokker--Planck residual. We also show numerically that conventional score-based diffusion models can exhibit significant differences between ODE- and SDE-induced distributions which we demonstrate using explicit comparisons. Moreover, we show numerically that reducing the Fokker--Planck residual by adding it as an additional regularisation term leads to closing the gap between ODE- and SDE-induced distributions. Our experiments suggest that this regularisation can improve the distribution generated by the ODE, however that this can come at the cost of degraded SDE sample quality.

cs.LG↗

Analysis of nonlinear poroviscoelastic flows with discontinuous porosities

Existence and uniqueness of solutions is shown for a class of viscoelastic flows in porous media with particular attention to problems with nonsmooth porosities. The considered models are formulated in terms of the time-dependent nonlinear interaction between porosity and effective pressure, which in certain cases leads to porosity waves. In particular, conditions for well-posedness in the presence of initial porosities with jump discontinuities are identified.

math.AP↗

$Γ$-Convergence of an Ambrosio-Tortorelli approximation scheme for image segmentation

Given an image $u_0$, the aim of minimising the Mumford-Shah functional is to find a decomposition of the image domain into sub-domains and a piecewise smooth approximation $u$ of $u_0$ such that $u$ varies smoothly within each sub-domain. Since the Mumford-Shah functional is highly non-smooth, regularizations such as the Ambrosio-Tortorelli approximation can be considered which is one of the most computationally efficient approximations of the Mumford-Shah functional for image segmentation. While very impressive numerical results have been achieved in a large range of applications when minimising the functional, no analytical results are currently available for minimizers of the functional in the piecewise smooth setting, and this is the goal of this work. Our main result is the $Γ$-convergence of the Ambrosio-Tortorelli approximation of the Mumford-Shah functional for piecewise smooth approximations. This requires the introduction of an appropriate function space. As a consequence of our $Γ$-convergence result, we can infer the convergence of minimizers of the respective functionals.

math.OC↗

Generalised Gillespie Algorithms for Simulations in a Rule-Based Epidemiological Model Framework

Rule-based models have been successfully used to represent different aspects of the COVID-19 pandemic, including age, testing, hospitalisation, lockdowns, immunity, infectivity, behaviour, mobility and vaccination of individuals. These rule-based approaches are motivated by chemical reaction rules which are traditionally solved numerically with the standard Gillespie algorithm proposed in the context of molecular dynamics. When applying reaction system type of approaches to epidemiology, generalisations of the Gillespie algorithm are required due to the time-dependency of the problems. In this article, we present different generalisations of the standard Gillespie algorithm which address discrete subtypes (e.g., incorporating the age structure of the population), time-discrete updates (e.g., incorporating daily imposed change of rates for lockdowns) and deterministic delays (e.g., given waiting time until a specific change in types such as release from isolation occurs). These algorithms are complemented by relevant examples in the context of the COVID-19 pandemic and numerical results.

q-bio.PE↗

Wasserstein GANs Work Because They Fail (to Approximate the Wasserstein Distance)

Wasserstein GANs are based on the idea of minimising the Wasserstein distance between a real and a generated distribution. We provide an in-depth mathematical analysis of differences between the theoretical setup and the reality of training Wasserstein GANs. In this work, we gather both theoretical and empirical evidence that the WGAN loss is not a meaningful approximation of the Wasserstein distance. Moreover, we argue that the Wasserstein distance is not even a desirable loss function for deep generative models, and conclude that the success of Wasserstein GANs can in truth be attributed to a failure to approximate the Wasserstein distance.

stat.ML↗

Autophosphorylation and the dynamics of the activation of Lck

Lck (lymphocyte-specific protein tyrosine kinase) is an enzyme which plays a number of important roles in the function of immune cells. It belongs to the Src family of kinases which are known to undergo autophosphorylation. It turns out that this leads to a remarkable variety of dynamical behaviour which can occur during their activation. We prove that in the presence of autophosphorylation one phenomenon, bistability, already occurs in a mathematical model for a protein with a single phosphorylation site. We further show that a certain model of Lck exhibits oscillations. Finally we discuss the relations of these results to models in the literature which involve Lck and describe specific biological processes, such as the early stages of T cell activation and the stimulation of T cell responses resulting from the suppression of PD-1 signalling which is important in immune checkpoint therapy for cancer.

q-bio.MN↗

Equilibria of an anisotropic nonlocal interaction equation: Analysis and numerics

In this paper, we study the equilibria of an anisotropic, nonlocal aggregation equation with nonlinear diffusion which does not possess a gradient flow structure. Here, the anisotropy is induced by an underlying tensor field. Anisotropic forces cannot be associated with a potential in general and stationary solutions of anisotropic aggregation equations generally cannot be regarded as minimizers of an energy functional. We derive equilibrium conditions for stationary line patterns in the setting of spatially homogeneous tensor fields. The stationary solutions can be regarded as the minimizers of a regularised energy functional depending on a scalar potential. A dimension reduction from the two- to the one-dimensional setting allows us to study the associated one-dimensional problem instead of the two-dimensional setting. We establish $Γ$-convergence of the regularised energy functionals as the diffusion coefficient vanishes, and prove the convergence of minimisers of the regularised energy functional to minimisers of the non-regularised energy functional. Further, we investigate properties of stationary solutions on the torus, based on known results in one spatial dimension. Finally, we prove weak convergence of a numerical scheme for the numerical solution of the anisotropic, nonlocal aggregation equation with nonlinear diffusion and any underlying tensor field, and show numerical results.

math.AP↗

On anisotropic diffusion equations for label propagation

In many problems in data classification one wishes to assign labels to points in a point cloud with a certain number of them being already correctly labeled. In this paper, we propose a microscopic ODE approach, in which information about correct labels is propagated to neighboring points. Its dynamics are based on alignment mechanisms, which are commonly used in large interacting agent systems in consensus formation. We derive the respective continuum description, which corresponds to an anisotropic diffusion equation with reaction term. Solutions of the continuum model on the bounded domain inherit certain properties of the underlying point cloud. We discuss these analytic properties and exemplify the results with micro- and macroscopic simulations.

math.AP↗

A Deterministic Gradient-Based Approach to Avoid Saddle Points

Loss functions with a large number of saddle points are one of the major obstacles for training modern machine learning models efficiently. First-order methods such as gradient descent are usually the methods of choice for training machine learning models. However, these methods converge to saddle points for certain choices of initial guesses. In this paper, we propose a modification of the recently proposed Laplacian smoothing gradient descent [Osher et al., arXiv:1806.06317], called modified Laplacian smoothing gradient descent (mLSGD), and demonstrate its potential to avoid saddle points without sacrificing the convergence rate. Our analysis is based on the attraction region, formed by all starting points for which the considered numerical scheme converges to a saddle point. We investigate the attraction region's dimension both analytically and numerically. For a canonical class of quadratic functions, we show that the dimension of the attraction region for mLSGD is floor((n-1)/2), and hence it is significantly smaller than that of the gradient descent whose dimension is n-1.

cs.LG↗

Mean-field optimal control for biological pattern formation

We propose a mean-field optimal control problem for the parameter identification of a given pattern. The cost functional is based on the Wasserstein distance between the probability measures of the modeled and the desired patterns. The first-order optimality conditions corresponding to the optimal control problem are derived using a Lagrangian approach on the mean-field level. Based on these conditions we propose a gradient descent method to identify relevant parameters such as angle of rotation and force scaling which may be spatially inhomogeneous. We discretize the first-order optimality conditions in order to employ the algorithm on the particle level. Moreover, we prove a rate for the convergence of the controls as the number of particles used for the discretization tends to infinity. Numerical results for the spatially homogeneous case demonstrate the feasibility of the approach.

math.OC↗

Detection of high codimensional bifurcations in variational PDEs

We derive bifurcation test equations for A-series singularities of nonlinear functionals and, based on these equations, we propose a numerical method for detecting high codimensional bifurcations in parameter-dependent PDEs such as parameter-dependent semilinear Poisson equations. As an example, we consider a Bratu-type problem and show how high codimensional bifurcations such as the swallowtail bifurcation can be found numerically. In particular, our original contributions are (1) the use of the Infinite-dimensional Splitting Lemma, (2) the unified and simplified treatment of all A-series bifurcations, (3) the presentation in Banach spaces, i.e. our results apply both to the PDE and its (variational) discretization, (4) further simplifications for parameter-dependent semilinear Poisson equations (both continuous and discrete), and (5) the unified treatment of the continuous problem and its discretisation.

math.NA↗

Auxin transport model for leaf venation

The plant hormone auxin controls many aspects of the development of plants. One striking dynamical feature is the self-organisation of leaf venation patterns which is driven by high levels of auxin within vein cells. The auxin transport is mediated by specialised membrane-localised proteins. Many venation models have been based on polarly localised efflux-mediator proteins of the PIN family. Here, we investigate a modeling framework for auxin transport with a positive feedback between auxin fluxes and transport capacities that are not necessarily polar, i.e.\ directional across a cell wall. Our approach is derived from a discrete graph-based model for biological transportation networks, where cells are represented by graph nodes and intercellular membranes by edges. The edges are not a-priori oriented and the direction of auxin flow is determined by its concentration gradient along the edge. We prove global existence of solutions to the model and the validity of Murray's law for its steady states. Moreover, we demonstrate with numerical simulations that the model is able connect an auxin source-sink pair with a mid-vein and that it can also produce branching vein patterns. A significant innovative aspect of our approach is that it allows the passage to a formal macroscopic limit which can be extended to include network growth. We perform mathematical analysis of the macroscopic formulation, showing the global existence of weak solutions for an appropriate parameter range.

math.DS↗

Stability analysis of line patterns of an anisotropic interaction model

Motivated by the formation of fingerprint patterns we consider a class of interacting particle models with anisotropic, repulsive-attractive interaction forces whose orientations depend on an underlying tensor field. This class of models can be regarded as a generalization of a gradient flow of a nonlocal interaction potential which has a local repulsion and a long-range attraction structure. In addition, the underlying tensor field introduces an anisotropy leading to complex patterns which do not occur in isotropic models. Central to this pattern formation are straight line patterns. For a given spatially homogeneous tensor field, we show that there exists a preferred direction of straight lines, i.e.\ straight vertical lines can be stable for sufficiently many particles, while many other rotations of the straight lines are unstable steady states, both for a sufficiently large number of particles and in the continuum limit. For straight vertical lines we consider specific force coefficients for the stability analysis of steady states, show that stability can be achieved for exponentially decaying force coefficients for a sufficiently large number of particles and relate these results to the Kücken-Champod model for simulating fingerprint patterns. The mathematical analysis of the steady states is completed with numerical results.

math.DS↗