Searcharxiv⌕ Search

arXiv subjects

Steffen Börm

Publications and source records attributed to Steffen Börm.

At least 19 recordsLinked to original sources

$\mathcal{H}^2$-matrices for translation-invariant kernel functions

Boundary element methods for elliptic partial differential equations typically lead to boundary integral operators with translation-invariant kernel functions. Taking advantage of this property is fairly simple for particle methods, e.g., Nystrom-type discretizations, but more challenging if the supports of basis functions have to be taken into account. In this article, we present a modified construction for $\mathcal{H}^2$-matrices that uses translation-invariance to significantly reduce the storage requirements. Due to the uniformity of the boxes used for the construction, we need only a few uncomplicated assumptions to prove estimates for the resulting storage complexity.

math.NA↗

Adaptive multiplication of rank-structured matrices in linear complexity

Hierarchical matrices approximate a given matrix by a decomposition into low-rank submatrices that can be handled efficiently in factorized form. $\mathcal{H}^2$-matrices refine this representation following the ideas of fast multipole methods in order to achieve linear, i.e., optimal complexity for a variety of important algorithms. The matrix multiplication, a key component of many more advanced numerical algorithms, has so far proven tricky: the only linear-time algorithms known so far either require the very special structure of HSS-matrices or need to know a suitable basis for all submatrices in advance. In this article, a new and fairly general algorithm for multiplying $\mathcal{H}^2$-matrices in linear complexity with adaptively constructed bases is presented. The algorithm consists of two phases: first an intermediate representation with a generalized block structure is constructed, then this representation is re-compressed in order to match the structure prescribed by the application. The complexity and accuracy are analysed and numerical experiments indicate that the new algorithm can indeed be significantly faster than previous attempts.

math.NA↗

Adaptive multiplication of $\mathcal{H}^2$-matrices with block-relative error control

The discretization of non-local operators, e.g., solution operators of partial differential equations or integral operators, leads to large densely populated matrices. $\mathcal{H}^2$-matrices take advantage of local low-rank structures in these matrices to provide an efficient data-sparse approximation that allows us to handle large matrices efficiently, e.g., to reduce the storage requirements to $\mathcal{O}(n k)$ for $n$-dimensional matrices with local rank $k$, and to reduce the complexity of the matrix-vector multiplication to $\mathcal{O}(n k)$ operations. In order to perform more advanced operations, e.g., to construct efficient preconditioners or evaluate matrix functions, we require algorithms that take $\mathcal{H}^2$-matrices as input and approximate the result again by $\mathcal{H}^2$-matrices, ideally with controllable accuracy. In this manuscript, we introduce an algorithm that approximates the product of two $\mathcal{H}^2$-matrices and guarantees block-relative error estimates for the submatrices of the result. It uses specialized tree structures to represent the exact product in an intermediate step, thereby allowing us to apply mathematically rigorous error control strategies.

math.NA↗

Memory-efficient compression of $\mathcal{DH}^2$-matrices for high-frequency problems

Directional interpolation is a fast and efficient compression technique for high-frequency Helmholtz boundary integral equations, but it requires a very large amount of storage in its original form. Algebraic recompression can significantly reduce the storage requirements and speed up the solution process accordingly. During the recompression process, weight matrices are required to correctly measure the influence of different basis vectors on the final result, and for highly accurate approximations, these weight matrices require more storage than the final compressed matrix. We present a compression method for the weight matrices and demonstrate that it introduces only a controllable error to the overall approximation. Numerical experiments show that the new method leads to a significant reduction in storage requirements.

math.NA↗

Computing Maxwell eigenmodes with Bloch boundary conditions

Our goal is to predict the band structure of photonic crystals. This task requires us to compute a number of the smallest non-zero eigenvalues of the time-harmonic Maxwell operator depending on the chosen Bloch boundary conditions. We propose to use a block inverse iteration preconditioned with a suitably modified geometric multigrid method. Since we are only interested in non-zero eigenvalues, we eliminate the large null space by combining a lifting operator and a secondary multigrid method. To obtain suitable initial guesses for the iteration, we employ a generalized extrapolation technique based on the minimization of the Rayleigh quotient that significantly reduces the number of iteration steps and allows us to treat families of very large eigenvalue problems efficiently.

math.NA↗

Topological photonics by breaking the degeneracy of line node singularities in semimetal-like photonic crystals

Degeneracy is an omnipresent phenomenon in various physical systems, which has its roots in the preservation of geometrical symmetry. In electronic and photonic crystal systems, very often this degeneracy can be broken by virtue of strong interactions between photonic modes of the same energy, where the level repulsion and the hybridization between modes causes the emergence of photonic bandgaps. However, most often this phenomenon does not lead to a complete and inverted bandgap formation over the entire Brillouin zone. Here, by systematically breaking the symmetry of a two-dimensional square photonic crystal, we investigate the formation of Dirac points, line node singularities, and inverted bandgaps. The formation of this complete bandgap is due to the level repulsion between degenerate modes along the line nodes of a semimetal-like photonic crystal, over the entire Brillouin zone. Our numerical experimentations are performed by a home-build numerical framework based on a multigrid finite element method. The developed numerical toolbox and our observations pave the way towards designing complete bandgap photonic crystals and exploring the role of symmetry on the optical behaviour of even more complicated orders in photonic crystal systems.

physics.optics↗

Distributed $\mathcal{H}^2$-matrices for boundary element methods

Standard discretization techniques for boundary integral equations, e.g., the Galerkin boundary element method, lead to large densely populated matrices that require fast and efficient compression techniques like the fast multipole method or hierarchical matrices. If the underlying mesh is very large, running the corresponding algorithms on a distributed computer is attractive, e.g., since distributed computers frequently are cost-effective and offer a high accumulated memory bandwidth. Compared to the closely related particle methods, for which distributed algorithms are well-established, the Galerkin discretization poses a challenge, since the supports of the basis functions influence the block structure of the matrix and therefore the flow of data in the corresponding algorithms. This article introduces distributed $\mathcal{H}^2$-matrices, a class of hierarchical matrices that is closely related to fast multipole methods and particularly well-suited for distributed computing. While earlier efforts required the global tree structure of the $\mathcal{H}^2$-matrix to be stored in every node of the distributed system, the new approach needs only local multilevel information that can be obtained via a simple distributed algorithm, allowing us to scale to significantly larger systems. Experiments show that this approach can handle very large meshes with more than $130$ million triangles efficiently.

math.NA↗

On iterated interpolation

Matrices resulting from the discretization of a kernel function, e.g., in the context of integral equations or sampling probability distributions, can frequently be approximated by interpolation. In order to improve the efficiency, a multi-level approach can be employed that involves interpolating the kernel function and its approximations multiple times. This article presents a new approach to analyze the error incurred by these iterated interpolation procedures that is considerably more elegant than its predecessors and allows us to treat not only the kernel function itself, but also its derivatives.

math.NA↗

Hybrid matrix compression for high-frequency problems

Boundary element methods for the Helmholtz equation lead to large dense matrices that can only be handled if efficient compression techniques are used. Directional compression techniques can reach good compression rates even for high-frequency problems. Currently there are two approaches to directional compression: analytic methods approximate the kernel function, while algebraic methods approximate submatrices. Analytic methods are quite fast and proven to be robust, while algebraic methods yield significantly better compression rates. We present a hybrid method that combines the speed and reliability of analytic methods with the good compression rates of algebraic methods.

math.NA↗

Dielectric breakdown prediction with GPU-accelerated BEM

The prediction of a dielectric breakdown in a high-voltage device is based on criteria that evaluate the electric field along field lines. Therefore it is necessary to efficiently compute the electric field at arbitrary points in space. A boundary element method (BEM) based on an indirect formulation, realized with MPI-parallel collocation, has proven to cope very well with this requirement. It deploys surface meshes only, which are easy to generate even for complex industrial geometries. The assembly of the large dense BEM-matrix, as well as the iterative solution process of the resulting system and the evaluation along the field-lines all require us to carry out the same type of calculation many times. Graphical processing units (GPUs) promise to be more efficient than standard processors (CPUs) in these situations. In this paper we investigate if GPU acceleration can speed up the established CPU-parallel BEM solver.

math.NA↗

Fast large-scale boundary element algorithms

Boundary element methods (BEM) reduce a partial differential equation in a domain to an integral equation on the domain's boundary. They are particularly attractive for solving problems on unbounded domains, but handling the dense matrices corresponding to the integral operators requires efficient algorithms. This article describes two approaches that allow us to solve boundary element equations on surface meshes consisting of several millions of triangles while preserving the optimal convergence rates of the Galerkin discretization.

math.NA↗

Algorithmic Counting of Zero-Dimensional Finite Topological Spaces With Respect to the Covering Dimension

Taking the covering dimension dim as notion for the dimension of a topological space, we first specify thenumber zdim_{T_0}(n) of zero-dimensional T_0-spaces on {1,...,n}$ and the number zdim(n) of zero-dimensional arbitrary topological spaces on {1,\ldots,n} by means oftwo mappings po and P that yieldthe number po(n) of partial orders on {1,...,n} and the set P(n) of partitions of {1,...,n}, respectively. Algorithms for both mappings exist. Assuming one for po to be at hand, we use our specification of zdim_{T_0}(n) and modify one for P in such a way that it computes zdim_{T_0}(n) instead of P(n). The specification of zdim(n) then allows to compute this number from zdim_{T_0}(1) to zdim_{T_0}(n) and the Stirling numbers of the second kind S(n,1) to S(n,n). The resulting algorithms have been implemented in C and we also present results of practical experiments with them. To considerably reduce the running times for computing zdim_{T_0}(n), we also describe a backtracking approach and its parallel implementation in C using the OpenMP library.

cs.DM↗

Semi-Automatic Task Graph Construction for $\mathcal{H}$-Matrix Arithmetic

A new method to construct task graphs for \mcH-matrix arithmetic is introduced, which uses the information associated with all tasks of the standard recursive \mcH-matrix algorithms, e.g., the block index set of the matrix blocks involved in the computation. Task refinement, i.e., the replacement of tasks by sub-computations, is then used to proceed in the \mcH-matrix hierarchy until the matrix blocks containing the actual matrix data are reached. This process is a natural extension of the classical, recursive way in which \mcH-matrix arithmetic is defined and thereby simplifies the efficient usage of many-core systems. Examples for standard and accumulator based \mcH-arithmetic are shown for model problems with different block structures.

cs.MS↗

Hierarchical matrix arithmetic with accumulated updates

Hierarchical matrices can be used to construct efficient preconditioners for partial differential and integral equations by taking advantage of low-rank structures in triangular factorizations and inverses of the corresponding stiffness matrices. The setup phase of these preconditioners relies heavily on low-rank updates that are responsible for a large part of the algorithm's total run-time, particularly for matrices resulting from three-dimensional problems. This article presents a new algorithm that significantly reduces the number of low-rank updates and can reduce the setup time by 50 percent or more.

math.NA↗

Exploiting nested task-parallelism in the $\mathcal{H}-LU$ factorization

We address the parallelization of the LU factorization of hierarchical matrices ($\mathcal{H}$-matrices) arising from boundary element methods. Our approach exploits task-parallelism via the OmpSs programming model and runtime, which discovers the data-flow parallelism intrinsic to the operation at execution time, via the analysis of data dependencies based on the memory addresses of the tasks' operands. This is especially challenging for $\mathcal{H}$-matrices, as the structures containing the data vary in dimension during the execution. We tackle this issue by decoupling the data structure from that used to detect dependencies. Furthermore, we leverage the support for weak operands and early release of dependencies, recently introduced in OmpSs-2, to accelerate the execution of parallel codes with nested task-parallelism and fine-grain tasks.

cs.MS↗

Complexity estimates for triangular hierarchical matrix algorithms

Triangular factorizations are an important tool for solving integral equations and partial differential equations with hierarchical matrices ($\mathcal{H}$-matrices). Experiments show that using an $\mathcal{H}$-matrix LR factorization to solve a system of linear questions is superior to direct inversion both with respect to accuracy and efficiency, but so far theoretical estimates quantifying these advantages were missing. Due to a lack of symmetry in $\mathcal{H}$-matrix algorithms, we cannot hope to prove that the LR factorization takes one third of the operations of the inversion or the matrix multiplication, as in standard linear algebra. We can, however, prove that the LR factorization together with two other operations of similar complexity, i.e., the inversion and multiplication of triangular matrices, requires not more operations than the matrix multiplication. We can complete the estimates by proving an improved upper bound for the complexity of the matrix multiplication, designed for recently introduced variants of classical $\mathcal{H}$-matrices.

math.NA↗

Variable Order, Directional H2-Matrices for Helmholtz Problems with Complex Frequency

The sparse approximation of high-frequency Helmholtz-type integral operators has many important physical applications such as problems in wave propagation and wave scattering. The discrete system matrices are huge and densely populated; hence their sparse approximation is of outstanding importance. In our paper we will generalize the directional $\mathcal{H}^{2}$-matrix techniques from the \textquotedblleft pure\textquotedblright\ Helmholtz operator $\mathcal{L}u=-Δu+ζ^{2}u$ with $ζ=-\operatorname*{i}k$, $k\in\mathbb{R}$, to general complex frequencies $ζ\in\mathbb{C}$ with $\operatorname{Re}ζ>0$. In this case, the fundamental solution decreases exponentially for large arguments. We will develop a new admissibility condition which contains $\operatorname{Re}ζ$ in an explicit way and introduce the approximation of the integral kernel function on admissible blocks in terms of frequency-dependent \textit{directional expansion functions}. We develop an error analysis which is explicit with respect to the expansion order and with respect to $\operatorname{Re}ζ$ and $\operatorname{Im}ζ$. This allows to choose the \textit{variable }expansion order in a quasi-optimal way depending on $\operatorname{Re}ζ$ but independent of, possibly large, $\operatorname{Im}ζ$. The complexity analysis is explicit with respect to $\operatorname{Re}ζ$ and $\operatorname{Im}ζ$ and shows how higher values of $\operatorname{Re}% ζ$ reduce the complexity. In certain cases, it even turns out that the discrete matrix can be replaced by its nearfield part. Numerical experiments illustrate the sharpness of the derived estimates and the efficiency of our sparse approximation.

math.NA↗

Large-scale magnetostatic field calculation in finite element micromagnetics with H2-matrices

Magnetostatic field calculations in micromagnetic simulations can be numerically expensive, particularly in the case of large-scale finite element simulations. The established finite element / boundary element method (FEM/BEM) by Fredkin & Koehler involves a densely populated matrix with unacceptable numerical costs for problems involving a large number of degrees of freedom $N$. By using hierarchical matrices of $\mathcal{H}^2$ type, we show that the memory requirements for the FEM/BEM method can be reduced dramatically, effectively converting the quadratic complexity $\mathcal{O}(N^2)$ of the problem to a linear one $\mathcal{O}(N)$. We obtain matrix size reductions of nearly $99\%$ in test cases with more than $10^6$ degrees of freedom, and we test the computed magnetostatic energy values by means of comparison with analytic values. The efficiency of the $\mathcal{H}^2$-matrix compression opens the way to large-scale magnetostatic field calculations in micromagnetic modeling, all while preserving the accuracy of the established FEM/BEM formalism.

math.NA↗