SearcharxivSearch

arXiv subjects

Michael Feischl

Publications and source records attributed to Michael Feischl.

At least 19 recordsLinked to original sources

High Probability Derivative Bounds for Random tanh Neural Networks on a Hypercube

We establish high-probability bounds for mixed input derivatives of wide random neural networks whose activation derivatives satisfy a factorial growth bound. Our main result specializes these estimates to $\tanh$ networks with Xavier initialization. A direct deterministic analysis based on Euclidean operator norms of the weight matrices yields derivative bounds that generally grow exponentially with the depth. We show that this growth can be substantially improved for sufficiently wide Gaussian networks by isolating the term that is linear in the highest-order derivative and controlling the corresponding tangent directions by measurable finite nets. For scalar-output $\tanh$ networks with Gaussian weights and Xavier initialization, we prove that there exist constants $C,C_0,C_1>0$ such that, whenever the common hidden width satisfies $n \geq C\left(L^3n_0^2(1+\log n_0)+L^2\left(1+\log(L/\eta)\right)\right)$, then, with probability at least $1-\eta$, the estimate $\left|D^u\mathcal{R}_{\Phi^{(L)}}(x)\right| \leq C_0 |u|! (C_1L)^{|u|-1}\prod_{j\in u}\beta_j(\eta,n_0)$ holds simultaneously for every non-empty $u\subseteq[n_0]$ and every $x\in[0,1]^{n_0}$. Thus, the first-order derivative bound is independent of the depth, while a square-free mixed derivative of order $|u|$ grows at most polynomially as $L^{|u|-1}$, apart from the coordinate factors. As consequences, we obtain high-probability bounds for the Euclidean Lipschitz constant and for weighted Sobolev norms of the network realization. The latter connect the derivative estimates to quasi-Monte Carlo integration and indicate how such regularity can enter the analysis of QMC-based training.

cs.LG

BDF2-type integrator for Landau-Lifshitz-Gilbert equation in micromagnetics: a-priori error estimates

We consider the Landau-Lifshitz-Gilbert equation (LLG), which models time-dependent micromagnetic phenomena. We analyze a fully discrete scheme that combines first-order finite elements in space with a BDF2 method in time. The method requires the solution of only one linear system of equations per time step and does not enforce the pointwise unit-length constraint of the magnetization. While unconditional weak convergence has been analyzed in an earlier work, we now prove optimal-order convergence rates under sufficient regularity assumptions on the exact solution and the external field. In combination with our previous work, this establishes the first higher-order-in-time and linear integrator that converges both to weak and strong solutions of LLG. Numerical experiments confirm first-order convergence in space and second-order convergence in time.

math.NA

Optimal Time-Adaptivity for Parabolic Problems

Since the first optimality proofs for adaptive mesh refinement algorithms in the early 2000s, the theory of optimal mesh refinement for PDEs was inherently limited to stationary problems. The reason for this is that time-dependent problems usually do not exhibit the necessary coercive structure that is used in optimality proofs to show a certain quasi-orthogonality, which is crucial for the theory. Recently, by using a new equivalence between quasi-orthogonality and inf-sup stability of the underlying problem, it was shown that an adaptive Crank-Nicolson scheme for the heat equation is optimal under a severe step size restriction. In this work, we use this new approach towards quasi-orthogonality together with Radau IIA methods of any order larger than one to obtain the first adaptive time stepping method for non-stationary PDEs that is provably rate optimal with respect to number of time steps vs. approximation error.

math.NA

BDF2-type integrator for Landau-Lifshitz-Gilbert equation in micromagnetics: unconditional weak convergence to weak solutions

We consider the Landau-Lifshitz-Gilbert equation (LLG) that models time-dependent micromagnetic phenomena. We propose a full discretization that employs first-order finite elements in space and a BDF2-type two-step method in time. In each time step, only one linear system of equations has to be solved. We employ linear interpolation in time to reconstruct the discrete space-time magnetization. We prove that the integrator is unconditionally stable and thus guarantees that a subsequence of the reconstructed magnetization converges weakly in $H^1$ towards a weak solution of LLG in the space-time domain. Numerical experiments verify that the proposed integrator is indeed first-order in space and second-order in time.

math.NA

A Reduced Basis Method for the Stochastic Landau-Lifshitz-Gilbert Equation

In this work, we consider the construction of efficient surrogates for the stochastic version of the Landau-Lifshitz-Gilbert (LLG) equation using model order reduction techniques, in particular, the Reduced Basis (RB) method. The Stochastic LLG (SLLG) equation is a widely used phenomenological model for the time evolution of the magnetization field confined to a ferromagnetic body while taking into account the effect of random heat perturbations. This phenomenon is mathematically formulated as a nonlinear parabolic problem, where the stochastic component is represented as a parameter-dependent datum depending on a non-compact and high-dimensional parameter. In an $\textit{offline}$ phase, we use Proper Orthogonal Decomposition (POD) on high-fidelity samples of the unbounded parameter space. To that end, we use the so-called $\textit{tangent plane scheme}$. For the $\textit{online}$ phase of the RB method, we again employ the tangent plane scheme in the RB space. This is possible due to our particular construction that reduces both spaces of the magnetization and of its time derivative. Due to the saddle-point nature of this scheme, a stabilization that appropriately enriches the RB space is required. Numerical experiments show a clear advantage over earlier approaches using sparse grid interpolation. In a complementary approach, we test a sparse grid approximation of the reduced coefficients in a purely data-driven method, exhibiting the weaknesses of earlier sparse grid approaches, but benefiting from increased stability.

math.NA

Optimal adaptive implicit time stepping

We revisit adaptive time stepping, one of the classical topics of numerical analysis and computational engineering. While widely used in application and subject of many theoretical works, a complete understanding is still missing. Apart from special cases, there does not exist a complete theory that shows how to choose the time steps such that convergence towards the exact solution is guaranteed with the optimal convergence rate. In this work, we use recent advances in adaptive mesh refinement to propose an adaptive time stepping algorithm that is mathematically guaranteed to be optimal in the sense that it achieves the best possible convergence of the error with respect to the number of time steps, and it can be implemented using a time stepping scheme as a black box.

math.NA

Computational Math with Neural Networks is Hard

We show that under some widely believed assumptions, there are no higher-order algorithms for basic tasks in computational mathematics such as: Computing integrals with neural network integrands, computing solutions of a Poisson equation with neural network source term, and computing the matrix-vector product with a neural network encoded matrix. We show that this is already true for very simple feed-forward networks with at least three hidden layers, bounded weights, bounded realization, and sparse connectivity, even if the algorithms are allowed to access the weights of the network. The fundamental idea behind these results is that it is already very hard to check whether a given neural network represents the zero function. The non-locality of the problems above allow us to reduce the approximation setting to deciding whether the input is zero or not. We demonstrate sharpness of our results by providing fast quadrature algorithms for one-layer networks and giving numerical evidence that quasi-Monte Carlo methods achieve the best possible order of convergence for quadrature with neural networks.

math.NA

Well-Posedness of Discretizations for Fractional Elasto-Plasticity

We consider a fractional plasticity model based on linear isotropic and kinematic hardening as well as a standard von-Mises yield function, where the flow rule is replaced by a Riesz--Caputo fractional derivative. The resulting mathematical model is typically non-local and non-smooth. Our numerical algorithm is based on the well-known radial return mapping and exploits that the kernel is finitely supported. We propose explicit and implicit discretizations of the model and show the well-posedness of the explicit in time discretization in combination with a standard finite element approach in space. Our numerical results in 2D and 3D illustrate the performance of the algorithm and the influence of the fractional parameter.

math.NA

Towards optimal hierarchical training of neural networks

We propose a hierarchical training algorithm for standard feed-forward neural networks that adaptively extends the network architecture as soon as the optimization reaches a stationary point. By solving small (low-dimensional) optimization problems, the extended network provably escapes any local minimum or stationary point. Under some assumptions on the approximability of the data with stable neural networks, we show that the algorithm achieves an optimal convergence rate s in the sense that loss is bounded by the number of parameters to the -s. As a byproduct, we obtain computable indicators which judge the optimality of the training state of a given network and derive a new notion of generalization error.

math.NA

Regularized dynamical parametric approximation

This paper studies the numerical approximation of evolution equations by nonlinear parametrizations $u(t)=\Phi(\param(t))$ with time-dependent parameters $\param(t)$, which are to be determined in the computation. The motivation comes from approximations in quantum dynamics by multiple Gaussians and approximations of various dynamical problems by tensor networks and neural networks. In all these cases, the parametrization is typically irregular: the derivative $\Phi'(\param)$ can have arbitrarily small singular values and may have varying rank. We derive approximation results for a regularized approach in the time-continuous case as well as in time-discretized cases. For the latter, there is a nontrivial interplay between the regularization parameter and the time stepsize, depending also on the defect size and local bounds of the second derivative of the parametrization map $\Phi$. When this is appropriately taken into account, the approach can be successfully applied in irregular situations, even though it runs counter to the basic principle in numerical analysis to avoid solving ill-posed subproblems when aiming for a stable algorithm. The paper also includes two theoretical case studies: regularized parametric steepest descent for optimization and regularized parametric Strang splitting for the time-dependent Schr\"odinger equation. Numerical experiments with sums of Gaussians for approximating quantum dynamics and with neural networks for approximating the flow map of a system of ordinary differential equations illustrate and complement the theoretical results.

math.NA

On full linear convergence and optimal complexity of adaptive FEM with inexact solver

The ultimate goal of any numerical scheme for partial differential equations (PDEs) is to compute an approximation of user-prescribed accuracy at quasi-minimal computational time. To this end, algorithmically, the standard adaptive finite element method (AFEM) integrates an inexact solver and nested iterations with discerning stopping criteria balancing the different error components. The analysis ensuring optimal convergence order of AFEM with respect to the overall computational cost critically hinges on the concept of R-linear convergence of a suitable quasi-error quantity. This work tackles several shortcomings of previous approaches by introducing a new proof strategy. First, the algorithm requires several fine-tuned parameters in order to make the underlying analysis work. A redesign of the standard line of reasoning and the introduction of a summability criterion for R-linear convergence allows us to remove restrictions on those parameters. Second, the usual assumption of a (quasi-)Pythagorean identity is replaced by the generalized notion of quasi-orthogonality from [Feischl, Math. Comp., 91 (2022)]. Importantly, this paves the way towards extending the analysis to general inf-sup stable problems beyond the energy minimization setting. Numerical experiments investigate the choice of the adaptivity parameters.

math.NA

Sparse grid approximation of nonlinear SPDEs: The Landau--Lifshitz--Gilbert equation

We show convergence rates for a sparse grid approximation of the distribution of solutions of the stochastic Landau-Lifshitz-Gilbert equation. Beyond being a frequently studied equation in engineering and physics, the stochastic Landau-Lifshitz-Gilbert equation poses many interesting challenges that do not appear simultaneously in previous works on uncertainty quantification: The equation is strongly non-linear, time-dependent, and has a non-convex side constraint. Moreover, the parametrization of the stochastic noise features countably many unbounded parameters and low regularity compared to other elliptic and parabolic problems studied in uncertainty quantification. We use a novel technique to establish uniform holomorphic regularity of the parameter-to-solution map based on a Gronwall-type estimate and the implicit function theorem. This method is very general and based on a set of abstract assumptions. Thus, it can be applied beyond the Landau-Lifshitz-Gilbert equation as well. We demonstrate numerically the feasibility of approximating with sparse grid and show a clear advantage of a multilevel sparse grid scheme.

math.NA

Adaptive Image Compression via Optimal Mesh Refinement

The JPEG algorithm is a defacto standard for image compression. We investigate whether adaptive mesh refinement can be used to optimize the compression ratio and propose a new adaptive image compression algorithm. We prove that it produces a quasi-optimal subdivision grid for a given error norm with high probability. This subdivision can be stored with very little overhead and thus leads to an efficient compression algorithm. We demonstrate experimentally, that the new algorithm can achieve better compression ratios than standard JPEG compression with no visible loss of quality on many images. The mathematical core of this work shows that Binev's optimal tree approximation algorithm is applicable to image compression with high probability, when we assume small additive Gaussian noise on the pixels of the image.

math.NA

Adaptive mesh refinement for the Landau-Lifshitz-Gilbert equation

We propose a new adaptive algorithm for the approximation of the Landau-Lifshitz-Gilbert equation via a higher-order tangent plane scheme. We show that the adaptive approximation satisfies an energy inequality and demonstrate numerically, that the adaptive algorithm outperforms uniform approaches.

math.NA

Inf-sup stability implies quasi-orthogonality

We prove new optimality results for adaptive mesh refinement algorithms for non-symmetric, indefinite, and time-dependent problems by proposing a generalization of quasi-orthogonality which follows directly from the inf-sup stability of the underlying problem. This completely removes a central technical difficulty in modern proofs of optimal convergence of adaptive mesh refinement algorithms and leads to simple optimality proofs for the Taylor-Hood discretization of the stationary Stokes problem, a finite-element/boundary-element discretization of an unbounded transmission problem, and an adaptive time-stepping scheme for parabolic equations. The main technical tool are new stability bounds for the LU-factorization of matrices together with a recently established connection between quasi-orthogonality and matrix factorization.

math.NA

A quasi-Monte Carlo data compression algorithm for machine learning

We introduce an algorithm to reduce large data sets using so-called digital nets, which are well distributed point sets in the unit cube. These point sets together with weights, which depend on the data set, are used to represent the data. We show that this can be used to reduce the computational effort needed in finding good parameters in machine learning algorithms. To illustrate our method we provide some numerical examples for neural networks.

math.NA

Convergence of adaptive stochastic collocation with finite elements

We consider an elliptic partial differential equation with a random diffusion parameter discretized by a stochastic collocation method in the parameter domain and a finite element method in the spatial domain. We prove convergence of an adaptive algorithm which adaptively enriches the parameter space as well as refines the finite element meshes.

math.NA

Recurrent Neural Networks as Optimal Mesh Refinement Strategies

We show that an optimal finite element mesh refinement algorithm for a prototypical elliptic PDE can be learned by a recurrent neural network with a fixed number of trainable parameters independent of the desired accuracy and the input size, i.e., number of elements of the mesh. Moreover, for a general class of PDEs with solutions which are well-approximated by deep neural networks, we show that an optimal mesh refinement strategy can be learned by recurrent neural networks. This includes problems for which no optimal adaptive strategy is known yet.

math.NA