Searcharxiv⌕ Search

arXiv subjects

Brynjulf Owren

Publications and source records attributed to Brynjulf Owren.

At least 19 recordsLinked to original sources

1-Lipschitz Neural Networks on Hadamard Manifolds

Controlling the Lipschitz constant of a neural network is a standard way to promote robustness and stability. Most existing constraining strategies are designed for Euclidean spaces. In this work, we construct and analyze a class of 1-Lipschitz neural networks on Hadamard manifolds. Our layers are of gradient-descent type, $1$-Lipschitz, and quasi-$α$-firmly nonexpansive. The core building blocks of the proposed architecture are Busemann functions, and we exploit the properties of Busemann gradient flows to design $1$-Lipschitz geometry-preserving layers. We provide explicit constructions and examples for hyperbolic manifolds and the manifold of symmetric positive definite (SPD) matrices. We test the proposed architecture in two numerical experiments: robust classification on the Poincaré disk and masked-Wishart covariance reconstruction. On the Poincaré disk, the proposed networks yield robust classifiers under hyperbolic perturbations. On the SPD manifold, we train SPD-valued denoisers and adopt them as a Plug-and-Play prior for a masked-Wishart covariance reconstruction problem. We show improved results from the nonexpansive denoiser over static, data-only, and Log-Euclidean denoising baselines, and empirically test its convergence properties.

math.NA↗

A Unified Discrete Gradient-SAV Framework for Structure-Preserving Integration

We present a framework combining discrete gradient (DG) methods with the Scalar Auxiliary Variable (SAV) approach to construct structure-preserving integrators for dissipative and conservative systems. The key observation is that SAV quadratization lifts the dynamics to an extended state space on which the modified energy has an exact discrete-gradient identity. This viewpoint yields three integrators with different accuracy and cost profiles: a first-order semi-implicit Forward Euler scheme, a second-order self-adjoint Midpoint scheme, and a second-order Predictive scheme with reduced implicitness. The construction extends to almost-Poisson systems and preserves selected Casimir invariants under an enforceable discrete condition. Numerical experiments cover the Allen--Cahn equation, an Ohta--Kawasaki-type nonlocal gradient flow, a double-well Hamiltonian oscillator, and a Poisson system with a nonlinear cubic Casimir.

math.NA↗

Learning Forced Multibody Dynamics on Lie Groups

We propose an architecture for learning the dynamics of mechanical systems based on discrete forced Euler-Lagrange equations on Lie groups using only position data. By formulating the dynamics directly on manifold-valued configuration spaces, the method naturally respects the geometric structure of the systems and preserves geometric invariants and conservation laws. The reliance on position measurements alone makes the framework applicable in settings where velocity data are unavailable or noisy. The approach extends naturally to multibody systems, accommodates external control inputs, and demonstrates strong performance on both synthetic and real-world datasets.

cs.LG↗

Mixed Precision Training of Neural ODEs

Exploiting low-precision computations has become a standard strategy in deep learning to address the growing computational costs imposed by ever larger models and datasets. However, naively performing all computations in low precision can lead to roundoff errors and instabilities. Therefore, mixed precision training schemes usually store the weights in high precision and use low-precision computations only for whitelisted operations. Despite their success, these principles are currently not reliable for training continuous-time architectures such as neural ordinary differential equations (Neural ODEs). This paper presents a mixed precision training framework for neural ODEs consisting of explicit ODE solvers and a custom backpropagation scheme and shows their effectiveness in a range of learning tasks. Our scheme uses low-precision computations for evaluating the velocity, parameterized by the neural network, and for storing intermediate states, while numerical reliability is provided by custom dynamic adjoint scaling and by accumulating the solution and gradients in higher precision. These contributions address two key challenges in training neural ODEs: the computational cost of repeated network evaluations and the growth of memory requirements with the number of time steps or layers. Along with the paper we publish our extendable, open-source PyTorch package \texttt{rampde}, whose syntax resembles that of leading packages to provide a drop-in replacement in existing codes. We demonstrate the reliability and effectiveness of our scheme using challenging test cases and on neural ODE applications in image classification and generative models, achieving approximately 50\% memory reduction and up to 2x speedup while maintaining accuracy comparable to single-precision training.

cs.LG↗

Approximation properties of neural ODEs

We study the approximation properties of neural ordinary differential equations (neural ODEs) in the space of continuous functions. Since a neural ODE requires input and output dimensions to be the same, while input and output dimensions of a continuous function are generally different, we need to embed an input into the latent space of the neural ODE, and to project the output of the neural ODE into the output space. By composing the neural ODE flow map with such embedding and projection operations, we get a shallow neural network whose activation function is defined as the flow map of the neural ODE at the final time of the integration interval. Thus, the study of the approximation properties of neural ODEs leads to the study of the approximation properties of shallow neural networks with a particular choice of activation function. We prove the universal approximation property (UAP) of such shallow neural networks in the space of continuous functions. Furthermore, we investigate the approximation properties of shallow neural networks whose parameters satisfy specific constraints. In particular, we constrain the Lipschitz constant of the neural ODE's flow map and the norms of the weights to increase the network's stability. We prove that the UAP holds if we consider either constraint independently. When both are enforced, there is a loss of expressiveness, and we derive approximation bounds that quantify how accurately such a constrained network can approximate a continuous function.

math.NA↗

Conditional Stability of the Euler Method on Riemannian Manifolds

We derive nonlinear stability results for numerical integrators on Riemannian manifolds, by imposing conditions on the ODE vector field and the step size that makes the numerical solution non-expansive whenever the exact solution is non-expansive over the same time step. Our model case is a geodesic version of the explicit Euler method. Precise bounds are obtained in the case of Riemannian manifolds of constant sectional curvature. The approach is based on a cocoercivity property of the vector field adapted to manifolds from Euclidean space. It allows us to compare the new results to the corresponding well-known results in flat spaces, and in general we find that a non-zero curvature will deteriorate the stability region of the geodesic Euler method. The step size bounds depend on the distance traveled over a step from the initial point. Numerical examples for spheres and hyperbolic 2-space confirm that the bounds are tight.

math.NA↗

Predictions Based on Pixel Data: Insights from PDEs and Finite Differences

As supported by abundant experimental evidence, neural networks are state-of-the-art for many approximation tasks in high-dimensional spaces. Still, there is a lack of a rigorous theoretical understanding of what they can approximate, at which cost, and at which accuracy. One network architecture of practical use, especially for approximation tasks involving images, is (residual) convolutional networks. However, due to the locality of the linear operators involved in these networks, their analysis is more complicated than that of fully connected neural networks. This paper deals with approximation of time sequences where each observation is a matrix. We show that with relatively small networks, we can represent exactly a class of numerical discretizations of PDEs based on the method of lines. We constructively derive these results by exploiting the connections between discrete convolution and finite difference operators. Our network architecture is inspired by those typically adopted in the approximation of time sequences. We support our theoretical results with numerical experiments simulating the linear advection, heat, and Fisher equations.

math.NA↗

Neural networks for the approximation of Euler's elastica

Euler's elastica is a classical model of flexible slender structures, relevant in many industrial applications. Static equilibrium equations can be derived via a variational principle. The accurate approximation of solutions of this problem can be challenging due to nonlinearity and constraints. We here present two neural network based approaches for the simulation of this Euler's elastica. Starting from a data set of solutions of the discretised static equilibria, we train the neural networks to produce solutions for unseen boundary conditions. We present a $\textit{discrete}$ approach learning discrete solutions from the discrete data. We then consider a $\textit{continuous}$ approach using the same training data set, but learning continuous solutions to the problem. We present numerical evidence that the proposed neural networks can effectively approximate configurations of the planar Euler's elastica for a range of different boundary conditions.

math.NA↗

Designing Stable Neural Networks using Convex Analysis and ODEs

Motivated by classical work on the numerical integration of ordinary differential equations we present a ResNet-styled neural network architecture that encodes non-expansive (1-Lipschitz) operators, as long as the spectral norms of the weights are appropriately constrained. This is to be contrasted with the ordinary ResNet architecture which, even if the spectral norms of the weights are constrained, has a Lipschitz constant that, in the worst case, grows exponentially with the depth of the network. Further analysis of the proposed architecture shows that the spectral norms of the weights can be further constrained to ensure that the network is an averaged operator, making it a natural candidate for a learned denoiser in Plug-and-Play algorithms. Using a novel adaptive way of enforcing the spectral norm constraints, we show that, even with these constraints, it is possible to train performant networks. The proposed architecture is applied to the problem of adversarially robust image classification, to image denoising, and finally to the inverse problem of deblurring.

cs.LG↗

B-stability of numerical integrators on Riemannian manifolds

We propose a generalization of nonlinear stability of numerical one-step integrators to Riemannian manifolds in the spirit of Butcher's notion of B-stability. Taking inspiration from Simpson-Porco and Bullo, we introduce non-expansive systems on such manifolds and define B-stability of integrators. In this first exposition, we provide concrete results for a geodesic version of the Implicit Euler (GIE) scheme. We prove that the GIE method is B-stable on Riemannian manifolds with non-positive sectional curvature. We show through numerical examples that the GIE method is expansive when applied to a certain non-expansive vector field on the 2-sphere, and that the GIE method does not necessarily possess a unique solution for large enough step sizes. Finally, we derive a new improved global error estimate for general Lie group integrators.

math.NA↗

Using aromas to search for preserved measures and integrals in Kahan's method

The numerical method of Kahan applied to quadratic differential equations is known to often generate integrable maps in low dimensions and can in more general situations exhibit preserved measures and integrals. Computerized methods based on discrete Darboux polynomials have recently been used for finding these measures and integrals. However, if the differential system contains many parameters, this approach can lead to highly complex results that can be difficult to interpret and analyze. But this complexity can in some cases be substantially reduced by using aromatic series. These are a mathematical tool introduced independently by Chartier and Murua and by Iserles, Quispel and Tse. We develop an algorithm for this purpose and derive some necessary conditions for the Kahan map to have preserved measures and integrals expressible in terms of aromatic functions. An important reason for the success of this method lies in the equivariance of the map from vector fields to their aromatic funtions. We demonstrate the algorithm on a number of examples showing a great reduction in complexity compared to what had been obtained by a fixed basis such as monomials.

math.NA↗

Dynamical systems' based neural networks

Neural networks have gained much interest because of their effectiveness in many applications. However, their mathematical properties are generally not well understood. If there is some underlying geometric structure inherent to the data or to the function to approximate, it is often desirable to take this into account in the design of the neural network. In this work, we start with a non-autonomous ODE and build neural networks using a suitable, structure-preserving, numerical time-discretisation. The structure of the neural network is then inferred from the properties of the ODE vector field. Besides injecting more structure into the network architectures, this modelling procedure allows a better theoretical understanding of their behaviour. We present two universal approximation results and demonstrate how to impose some particular properties on the neural networks. A particular focus is on 1-Lipschitz architectures including layers that are not 1-Lipschitz. These networks are expressive and robust against adversarial attacks, as shown for the CIFAR-10 and CIFAR-100 datasets.

cs.LG↗

Learning Hamiltonians of constrained mechanical systems

Recently, there has been an increasing interest in modelling and computation of physical systems with neural networks. Hamiltonian systems are an elegant and compact formalism in classical mechanics, where the dynamics is fully determined by one scalar function, the Hamiltonian. The solution trajectories are often constrained to evolve on a submanifold of a linear vector space. In this work, we propose new approaches for the accurate approximation of the Hamiltonian function of constrained mechanical systems given sample data information of their solutions. We focus on the importance of the preservation of the constraints in the learning strategy by using both explicit Lie group integrators and other classical schemes.

math.NA↗

Detecting and determining preserved measures and integrals of birational maps

In this paper we use the method of discrete Darboux polynomials to calculate preserved measures and integrals of rational maps. The approach is based on the use of cofactors and Darboux polynomials and relies on the use of symbolic algebra tools. Given sufficient computing power, most, if not all, rational preserved integrals can be found (and even some non-rational ones). We show, in a number of examples, how it is possible to use this method to both determine and detect preserved measures and integrals of the considered rational maps. Many of the examples arise from the Kahan-Hirota-Kimura discretization of completely integrable systems of ordinary differential equations.

math.NA↗

Lie Group integrators for mechanical systems

Since they were introduced in the 1990s, Lie group integrators have become a method of choice in many application areas. These include multibody dynamics, shape analysis, data science, image registration and biophysical simulations. Two important classes of intrinsic Lie group integrators are the Runge--Kutta--Munthe--Kaas methods and the commutator free Lie group integrators. We give a short introduction to these classes of methods. The Hamiltonian framework is attractive for many mechanical problems, and in particular we shall consider Lie group integrators for problems on cotangent bundles of Lie groups where a number of different formulations are possible. There is a natural symplectic structure on such manifolds and through variational principles one may derive symplectic Lie group integrators. We also consider the practical aspects of the implementation of Lie group integrators, such as adaptive time stepping. The theory is illustrated by applying the methods to two nontrivial applications in mechanics. One is the N-fold spherical pendulum where we introduce the restriction of the adjoint action of the group $SE(3)$ to $TS^2$, the tangent bundle of the two-dimensional sphere. Finally, we show how Lie group integrators can be applied to model the controlled path of a payload being transported by two rotors. This problem is modeled on $\mathbb{R}^6\times \left(SO(3)\times \mathfrak{so}(3)\right)^2\times (TS^2)^2$ and put in a format where Lie group integrators can be applied.

math.NA↗

Dynamics of the N-fold Pendulum in the framework of Lie Group Integrators

Since their introduction, Lie group integrators have become a method of choice in many application areas. Various formulations of these integrators exist, and in this work we focus on Runge--Kutta--Munthe--Kaas methods. First, we briefly introduce this class of integrators, considering some of the practical aspects of their implementation, such as adaptive time stepping. We then present some mathematical background that allows us to apply them to some families of Lagrangian mechanical systems. We conclude with an application to a nontrivial mechanical system: the N-fold 3D pendulum.

math.NA↗

Computational geometric methods for preferential clustering of particle suspensions

A geometric numerical method for simulating suspensions of spherical and non-spherical particles with Stokes drag is proposed. The method combines divergence-free matrix-valued radial basis function interpolation of the fluid velocity field with a splitting method integrator that preserves the sum of the Lyapunov spectrum while mimicking the centrifuge effect of the exact solution. We discuss how breaking the divergence-free condition in the interpolation step can erroneously affect how the volume of the particulate phase evolves under numerical methods. The methods are tested on suspensions of $10^4$ particles evolving in discrete cellular flow field. The results are that the proposed geometric methods generate more accurate and cost-effective particle distributions compared to conventional methods.

physics.comp-ph↗

Equivariant neural networks for inverse problems

In recent years the use of convolutional layers to encode an inductive bias (translational equivariance) in neural networks has proven to be a very fruitful idea. The successes of this approach have motivated a line of research into incorporating other symmetries into deep learning methods, in the form of group equivariant convolutional neural networks. Much of this work has been focused on roto-translational symmetry of $\mathbf R^d$, but other examples are the scaling symmetry of $\mathbf R^d$ and rotational symmetry of the sphere. In this work, we demonstrate that group equivariant convolutional operations can naturally be incorporated into learned reconstruction methods for inverse problems that are motivated by the variational regularisation approach. Indeed, if the regularisation functional is invariant under a group symmetry, the corresponding proximal operator will satisfy an equivariance property with respect to the same group symmetry. As a result of this observation, we design learned iterative methods in which the proximal operators are modelled as group equivariant convolutional neural networks. We use roto-translationally equivariant operations in the proposed methodology and apply it to the problems of low-dose computerised tomography reconstruction and subsampled magnetic resonance imaging reconstruction. The proposed methodology is demonstrated to improve the reconstruction quality of a learned reconstruction method with a little extra computational cost at training time but without any extra cost at test time.

cs.LG↗