Searcharxiv⌕ Search

arXiv subjects

Reinhold Schneider

Publications and source records attributed to Reinhold Schneider.

At least 19 recordsLinked to original sources

On low-rank tensor train approximability for linear nearest neighbor systems

Low-rank tensor methods are an important tool in the numerical treatment of equations with a high-dimensional state space. Nearest neighbor interaction systems like the Ising model or more general Markov jump processes, as well as 1D finite-state quantum systems are examples of such problems. While low-rank tensor train/matrix product state models have been shown to be highly efficient for the simulation of such systems, providing theoretical justification for this remains a challenging task. One approach for obtaining estimates on required ranks for certain accuracies is to investigate the rank increase in Krylov subspace methods for solving the problem at hand. In the context of area laws for ground states of 1D spin systems, nontrivial results on rank-increasing properties of nearest neighbor operator polynomials have been obtained in work of Arad et al. [arXiv:1301.1162] by studying the partial commutativity of local operators. In the present work, this technique is applied to polynomial methods for definite linear equations and dissipative linear ODEs with nearest neighbor structure. This allows to derive corresponding low-rank approximability statements for solutions of such problems which are independent of the system size. Numerical simulations of high-dimensional nearest neighbor systems illustrate the theoretical findings.

math.NA↗

Low-rank approximation of continuous functions in Sobolev spaces with dominating mixed smoothness

Let $Ω_i\subset\mathbb{R}^{n_i}$, $i=1,\ldots,m$, be given domains. In this article, we study the low-rank approximation with respect to $L^2(Ω_1\times\dots\timesΩ_m)$ of functions from Sobolev spaces with dominating mixed smoothness. To this end, we first estimate the rank of a bivariate approximation, i.e., the rank of the continuous singular value decomposition. In comparison to the case of functions from Sobolev spaces with isotropic smoothness, compare \cite{GH14,GH19}, we obtain improved results due to the additional mixed smoothness. This convergence result is then used to study the tensor train decomposition as a method to construct multivariate low-rank approximations of functions from Sobolev spaces with dominating mixed smoothness. We show that this approach is able to beat the curse of dimension.

math.NA↗

Some suggestions concerning the conjecture in: 'Tractable semi-algebraic approximation using Christoffel-Darboux kernel'

In 'Tractable semi-algebraic approximation using Christoffel-Darboux kernel' Marx, Pauwels, Weisser, Henrion and Lasserre conjectured, that the approximation rate $\mathcal O(\frac 1 {\sqrt(d)})$ of a Lipschitz functions by a semi-algebraic function induced by a Christoffel- Darboux kernel of degree $d$ in the $L^1$ norm can be improved for more regular functions. Here we will show, that for semi-algebraic and definable functions the results can be strengthened to a rational approximation rate in the $L^\infty$ norm.

math.CA↗

Approximation by tree tensor networks in high dimensions: Sobolev and compositional functions

This paper is concerned with convergence estimates for fully discrete tree tensor network approximations of high-dimensional functions from several model classes. For functions having standard or mixed Sobolev regularity, new estimates generalizing and refining known results are obtained, based on notions of linear widths of multivariate functions. In the main results of this paper, such techniques are applied to classes of functions with compositional structure, which are known to be particularly suitable for approximation by deep neural networks. As shown here, such functions can also be approximated by tree tensor networks without a curse of dimensionality -- however, subject to certain conditions, in particular on the depth of the underlying tree. In addition, a constructive encoding of compositional functions in tree tensor networks is given.

math.NA↗

Dynamical low-rank approximations of solutions to the Hamilton-Jacobi-Bellman equation

We present a novel method to approximate optimal feedback laws for nonlinear optimal control based on low-rank tensor train (TT) decompositions. The approach is based on the Dirac-Frenkel variational principle with the modification that the optimisation uses an empirical risk. Compared to current state-of-the-art TT methods, our approach exhibits a greatly reduced computational burden while achieving comparable results. A rigorous description of the numerical scheme and demonstrations of its performance are provided.

math.OC↗

Approximating the Stationary Bellman Equation by Hierarchical Tensor Products

We treat infinite horizon optimal control problems by solving the associated stationary Hamilton-Jacobi-Bellman (HJB) equation numerically to compute the value function and an optimal feedback law. The dynamical systems under consideration are spatial discretizations of non linear parabolic partial differential equations (PDE), which means that the HJB is non linear and suffers from the curse of dimensionality. Its non linearity is handled by the Policy Iteration algorithm, where the problem is reduced to a sequence of linear, hyperbolic PDEs. These equations remain the computational bottleneck due to their high dimensions. By the method of characteristics these linearized HJB equations can be reformulated via the Koopman operator in the spirit of dynamic programming. The resulting operator equations are solved using a minimal residual method. To overcome numerical infeasability we use low rank hierarchical tensor product approximation/tree-based tensor formats, in particular tensor trains (TT tensors), and multi-polynomials, together with high dimensional quadrature, e.g. Monte-Carlo. By controlling a destabilized version of viscous Burgers and a diffusion equation with unstable reaction term numerical evidences are given.

math.OC↗

Convergence bounds for empirical nonlinear least-squares

We consider best approximation problems in a nonlinear subset $\mathcal{M}$ of a Banach space of functions $(\mathcal{V},\|\bullet\|)$. The norm is assumed to be a generalization of the $L^2$-norm for which only a weighted Monte Carlo estimate $\|\bullet\|_n$ can be computed. The objective is to obtain an approximation $v\in\mathcal{M}$ of an unknown function $u \in \mathcal{V}$ by minimizing the empirical norm $\|u-v\|_n$. We consider this problem for general nonlinear subsets and establish error bounds for the empirical best approximation error. Our results are based on a restricted isometry property (RIP) which holds in probability and is independent of the nonlinear least squares setting. Several model classes are examined where analytical statements can be made about the RIP and the results are compared to existing sample complexity bounds from the literature. We find that for well-studied model classes our general bound is weaker but exhibits many of the same properties as these specialized bounds. Notably, we demonstrate the advantage of an optimal sampling density (as known for linear spaces) for sets of functions with sparse representations.

math.NA↗

A block-sparse Tensor Train Format for sample-efficient high-dimensional Polynomial Regression

Low-rank tensors are an established framework for high-dimensional least-squares problems. We propose to extend this framework by including the concept of block-sparsity. In the context of polynomial regression each sparsity pattern corresponds to some subspace of homogeneous multivariate polynomials. This allows us to adapt the ansatz space to align better with known sample complexity results. The resulting method is tested in numerical experiments and demonstrates improved computational resource utilization and sample efficiency.

math.NA↗

Approximating optimal feedback controllers of finite horizon control problems using hierarchical tensor formats

Controlling systems of ordinary differential equations (ODEs) is ubiquitous in science and engineering. For finding an optimal feedback controller, the value function and associated fundamental equations such as the Bellman equation and the Hamilton-Jacobi-Bellman (HJB) equation are essential. The numerical treatment of these equations poses formidable challenges due to their non-linearity and their (possibly) high-dimensionality. In this paper we consider a finite horizon control system with associated Bellman equation. After a time-discretization, we obtain a sequence of short time horizon problems which we call local optimal control problems. For solving the local optimal control problems we apply two different methods, one being the well-known policy iteration, where a fixed-point iteration is required for every time step. The other algorithm borrows ideas from Model Predictive Control (MPC), by solving the local optimal control problem via open-loop control methods on a short time horizon, allowing us to replace the fixed-point iteration by an adjoint method. For high-dimensional systems we apply low rank hierarchical tensor product approximation/tree-based tensor formats, in particular tensor trains (TT tensors) and multi-polynomials, together with high-dimensional quadrature, e.g. Monte-Carlo. We prove a linear error propagation with respect to the time discretization and give numerical evidence by controlling a diffusion equation with unstable reaction term and an Allen-Kahn equation.

math.OC↗

Approximative Policy Iteration for Exit Time Feedback Control Problems driven by Stochastic Differential Equations using Tensor Train format

We consider a stochastic optimal exit time feedback control problem. The Bellman equation is solved approximatively via the Policy Iteration algorithm on a polynomial ansatz space by a sequence of linear equations. As high degree multi-polynomials are needed, the corresponding equations suffer from the curse of dimensionality even in moderate dimensions. We employ tensor-train methods to account for this problem. The approximation process within the Policy Iteration is done via a Least-Squares ansatz and the integration is done via Monte-Carlo methods. Numerical evidences are given for the (multi dimensional) double well potential and a three-hole potential.

math.OC↗

A Theoretical Analysis of Deep Neural Networks and Parametric PDEs

We derive upper bounds on the complexity of ReLU neural networks approximating the solution maps of parametric partial differential equations. In particular, without any knowledge of its concrete shape, we use the inherent low-dimensionality of the solution manifold to obtain approximation rates which are significantly superior to those provided by classical neural network approximation results. Concretely, we use the existence of a small reduced basis to construct, for a large variety of parametric partial differential equations, neural networks that yield approximations of the parametric solution maps in such a way that the sizes of these networks essentially only depend on the size of the reduced basis.

math.NA↗

Numerical Solution of the Parametric Diffusion Equation by Deep Neural Networks

We perform a comprehensive numerical study of the effect of approximation-theoretical results for neural networks on practical learning problems in the context of numerical analysis. As the underlying model, we study the machine-learning-based solution of parametric partial differential equations. Here, approximation theory predicts that the performance of the model should depend only very mildly on the dimension of the parameter space and is determined by the intrinsic dimension of the solution manifold of the parametric partial differential equation. We use various methods to establish comparability between test-cases by minimizing the effect of the choice of test-cases on the optimization and sampling aspects of the learning problem. We find strong support for the hypothesis that approximation-theoretical effects heavily influence the practical behavior of learning problems in numerical analysis.

math.NA↗

Analysis of The Tailored Coupled-Cluster Method in Quantum Chemistry

In quantum chemistry, one of the most important challenges is the static correlation problem when solving the electronic Schrödinger equation for molecules in the Born--Oppenheimer approximation. In this article, we analyze the tailored coupled-cluster method (TCC), one particular and promising method for treating molecular electronic-structure problems with static correlation. The TCC method combines the single-reference coupled-cluster (CC) approach with an approximate reference calculation in a subspace [complete active space (CAS)] of the considered Hilbert space that covers the static correlation. A one-particle spectral gap assumption is introduced, separating the CAS from the remaining Hilbert space. This replaces the nonexisting or nearly nonexisting gap between the highest occupied molecular orbital and the lowest unoccupied molecular orbital usually encountered in standard single-reference quantum chemistry. The analysis covers, in particular, CC methods tailored by tensor-network states (TNS-TCC methods). The problem is formulated in a nonlinear functional analysis framework, and, under certain conditions such as the aforementioned gap, local uniqueness and existence are proved using Zarantonello's lemma. From the Aubin--Nitsche-duality method, a quadratic error bound valid for TNS-TCC methods is derived, e.g., for linear-tensor-network TCC schemes using the density matrix renormalization group method.

math.NA↗

Adaptive stochastic Galerkin FEM for lognormal coefficients in hierarchical tensor representations

Stochastic Galerkin methods for non-affine coefficient representations are known to cause major difficulties from theoretical and numerical points of view. In this work, an adaptive Galerkin FE method for linear parametric PDEs with lognormal coefficients discretized in Hermite chaos polynomials is derived. It employs problem-adapted function spaces to ensure solvability of the variational formulation. The inherently high computational complexity of the parametric operator is made tractable by using hierarchical tensor representations. For this, a new tensor train format of the lognormal coefficient is derived and verified numerically. The central novelty is the derivation of a reliable residual-based a posteriori error estimator. This can be regarded as a unique feature of stochastic Galerkin methods. It allows for an adaptive algorithm to steer the refinements of the physical mesh and the anisotropic Wiener chaos polynomial degrees. For the evaluation of the error estimator to become feasible, a numerically efficient tensor format discretization is developed. Benchmark examples with unbounded lognormal coefficient fields illustrate the performance of the proposed Galerkin discretization and the fully adaptive algorithm.

math.NA↗

Variational Monte Carlo - Bridging Concepts of Machine Learning and High Dimensional Partial Differential Equations

A statistical learning approach for parametric PDEs related to Uncertainty Quantification is derived. The method is based on the minimization of an empirical risk on a selected model class and it is shown to be applicable to a broad range of problems. A general unified convergence analysis is derived, which takes into account the approximation and the statistical errors. By this, a combination of theoretical results from numerical analysis and statistics is obtained. Numerical experiments illustrate the performance of the method with the model class of hierarchical tensors.

math.NA↗

Numerical and Theoretical Aspects of the DMRG-TCC Method Exemplified by the Nitrogen Dimer

In this article, we investigate the numerical and theoretical aspects of the coupled-cluster method tailored by matrix-product states. We investigate chemical properties of the used method, such as energy size extensivity and the equivalence of linked and unlinked formulation. The existing mathematical analysis is here elaborated in a quantum chemical framework. In particular, we highlight the use of a so-called CAS-ext gap describing the basis splitting between the complete active space and the external part. Moreover, the behavior of the energy error as a function of the optimal basis splitting is discussed. We show numerical investigations on the robustness with respect to the bond dimensions of the single orbital entropy and the mutual information, which are quantities that are used to choose the complete active space. Furthermore, we extend the mathematical analysis with a numerical study on the complete active space dependence of the error.

physics.chem-ph↗

A Randomized Tensor Train Singular Value Decomposition

The hierarchical SVD provides a quasi-best low rank approximation of high dimensional data in the hierarchical Tucker framework. Similar to the SVD for matrices, it provides a fundamental but expensive tool for tensor computations. In the present work we examine generalizations of randomized matrix decomposition methods to higher order tensors in the framework of the hierarchical tensors representation. In particular we present and analyze a randomized algorithm for the calculation of the hierarchical SVD (HSVD) for the tensor train (TT) format.

math.NA↗

Low rank tensor recovery via iterative hard thresholding

We study extensions of compressive sensing and low rank matrix recovery (matrix completion) to the recovery of low rank tensors of higher order from a small number of linear measurements. While the theoretical understanding of low rank matrix recovery is already well-developed, only few contributions on the low rank tensor recovery problem are available so far. In this paper, we introduce versions of the iterative hard thresholding algorithm for several tensor decompositions, namely the higher order singular value decomposition (HOSVD), the tensor train format (TT), and the general hierarchical Tucker decomposition (HT). We provide a partial convergence result for these algorithms which is based on a variant of the restricted isometry property of the measurement operator adapted to the tensor decomposition at hand that induces a corresponding notion of tensor rank. We show that subgaussian measurement ensembles satisfy the tensor restricted isometry property with high probability under a certain almost optimal bound on the number of measurements which depends on the corresponding tensor format. These bounds are extended to partial Fourier maps combined with random sign flips of the tensor entries. Finally, we illustrate the performance of iterative hard thresholding methods for tensor recovery via numerical experiments where we consider recovery from Gaussian random measurements, tensor completion (recovery of missing entries), and Fourier measurements for third order tensors.

cs.IT↗