SearcharxivSearch

arXiv subjects

Martin Kronbichler

Publications and source records attributed to Martin Kronbichler.

At least 19 recordsLinked to original sources

Multi-Solver Coupling for Parallel Adaptive Multi-Physics Simulations with Trixi$.$jl and deal$.$II

Many standalone frameworks for numerical solvers have been developed to tackle the simulation of specific single- or multi-physics problems. For solving coupled problems, common approaches are to extend existing solvers, to develop an entirely new solver or to couple two existing solvers by using the interface provided by a coupling library or framework. However, to the best of our knowledge, there is not yet a framework that allows the easy and direct development of coupled solvers. In this work, we prototype a portable reproducible cross-language framework for the development of coupled parallel adaptive solvers using Trixi$.$jl and deal$.$II for the numerical simulation of coupled multi-physics problems. Currently, this is tightly entangled with an example coupled solver. We show its usability by developing a partitioned strongly-coupled multi-physics solver for the dynamics of Newtonian self-gravitational gases. For the coupled solver, we validate the expected order of convergence, physical sensibility of the results and mesh adaptivity. Finally, we investigate its parallel scaling. A publicly accessible reproducibility repository for the numerical results and code is available.

math.NA

Solving the (Navier-)Stokes equations with space and time adaptivity using deal.II

In this article, we solve the Stokes and Navier-Stokes equations with the deal$.$II finite-element library. In particular, we use its multigrid, adaptive-mesh, and matrix-free infrastructures to design efficient linear and nonlinear iterative solvers, respectively. We solve the stationary Stokes equations on hp-adaptive meshes with a hp-multigrid approach, the transient Stokes equations with space-time finite elements and space-time multigrid, and, finally, the stabilized incompressible Navier-Stokes equations on locally refined meshes with a monolithic multigrid solver. The selected examples underline the flexibility and modularity of the multigrid infrastructure of deal$.$II.

math.NA

A Discontinuous Galerkin Consistent Splitting Method for the Incompressible Navier-Stokes Equations

This work presents the discontinuous Galerkin discretization of the consistent splitting scheme proposed by Liu [J. Liu, J. Comp. Phys., 228(19), 2009]. The method enforces the divergence-free constraint implicitly, removing velocity--pressure compatibility conditions and eliminating pressure boundary layers. Consistent boundary conditions are imposed, also for settings with open and traction boundaries. Hence, accuracy in time is no longer limited by a splitting error. The symmetric interior penalty Galerkin method is used for second spatial derivatives. The convective term is treated in a semi-implicit manner, which relaxes the CFL restriction of explicit schemes while avoiding the need to solve nonlinear systems required by fully implicit formulations. For improved mass conservation, Leray projection is combined with divergence and normal continuity penalty terms. By selecting appropriate fluxes for both the divergence of the velocity field and the divergence of the convective operator, the consistent pressure boundary condition can be shown to reduce to contributions arising solely from the acceleration and the viscous term for the $L^2$ discretization. Per time step, the decoupled nature of the scheme with respect to the velocity and pressure fields leads to a single pressure Poisson equation followed by a single vector-valued convection-diffusion-reaction equation. We verify optimal convergence rates of the method in both space and time and demonstrate compatibility with higher-order time integration schemes. A series of numerical experiments, including the two-dimensional flow around a cylinder benchmark and the three-dimensional Taylor--Green vortex problem, verify the applicability to practically relevant flow problems.

math.NA

Matrix-Free Evaluation Strategies for Continuous and Discontinuous Galerkin Discretizations on Unstructured Tetrahedral Grids

This study presents novel strategies for improving the node-level performance of matrix-free evaluation of continuous and discontinuous Galerkin spatial discretizations on unstructured tetrahedral grids. In our approach the underlying integrals of a generic finite-element operator are computed cell-by-cell through numerical quadrature using tabulated dense local matrices of shape functions, achieving high throughput for low to moderate-order polynomial degrees. By employing dense matrix-matrix products instead of matrix-vector products for the cell-wise interpolation, the method reaches over $60\%$ of peak performance. The optimization strategies exploit explicit data parallelism to enhance computational efficiency, complemented by a hierarchical mesh reordering algorithm that improves data locality. The matrix-free implementation achieves up to a $6\times$ speedup compared to a global sparse matrix-based approach at a polynomial degree of three. The effectiveness of the method is demonstrated through numerical experiments on the Poisson and Navier--Stokes equations. The Poisson operator is preconditioned by a hybrid multigrid scheme that combines auxiliary continuous finite-element spaces, polynomial and geometric coarsening where possible while employing algebraic multigrid on the coarse mesh. Within the preconditioner, the implementation transitions between the matrix-free and matrix-based strategies for optimal efficiency. Finally, we analyze the strong scaling behavior of the Poisson and Helmholtz operators, demonstrating the method's potential to solve large real-world problems.

math.NA

A conservative invariant-domain preserving projection technique for hyperbolic systems under adaptive mesh refinement

We propose a rigorous, conservative invariant-domain preserving (IDP) projection technique for hierarchical discretizations that enforces membership in physics-implied convex sets when mapping between solution spaces. When coupled with suitable refinement indicators, the proposed scheme enables a provably IDP adaptive numerical method for hyperbolic systems where preservation of physical properties is essential. In addition to proofs of these characteristics, we supply a detailed construction of the method in the context of a high-performance finite element code. To illustrate our proposed scheme, we study a suite of computationally challenging benchmark problems, demonstrating enhanced accuracy and efficiency properties while entirely avoiding \emph{ad hoc} corrections to preserve physical invariants.

math.NA

Higher-Order Discontinuous Galerkin Splitting Schemes for Fluids with Variable Viscosity

This article investigates matrix-free higher-order discontinuous Galerkin discretizations of the Navier--Stokes equations for incompressible flows with variable viscosity. The viscosity field may be prescribed analytically or governed by a rheological law, as often found in biomedical or industrial applications. The DG discretization of the adapted second-order viscous terms is carried out via the symmetric interior penalty Galerkin method, obviating auxiliary variables. Based on this spatial discretization, we compare several linearized variants of saddle point block systems and projection-based splitting time integration schemes in terms of their computational performance. Compared to the velocity-pressure block-system for the former, the splitting scheme allows solving a sequence of simple problems such as mass, convection-diffusion and Poisson equations. We investigate under which conditions the improved temporal stability of fully implicit schemes and resulting expensive nonlinear solves outperform the splitting schemes and linearized variants that are stable under hyperbolic time step restrictions. The key aspects of this work are i) a higher-order DG discretization for incompressible flows with variable viscosity, ii) accelerated nonlinear solver variants and suitable linearizations adopting a matrix-free $hp$-multigrid solver, and iii) a detailed comparison of the monolithic and projection-based solvers in terms of their (non-)linear solver performance. The presented schemes are evaluated in a series of numerical examples verifying their spatial and temporal accuracy, and the preconditioner performance under increasing viscosity contrasts, while their efficiency is showcased in the backward-facing step benchmark.

cs.CE

Matrix-Free Methods for Finite-Strain Elasticity: Automatic Code Generation with No Performance Overhead

This study explores matrix-free tangent evaluations in finite-strain elasticity with the use of automatically-generated code for the quadrature-point level calculations. The code generation is done via automatic differentiation (AD) with AceGen. We compare hand-written and AD-generated codes under two computing strategies: on-the-fly evaluation and caching intermediate results. The comparison reveals that the AD-generated code achieves superior performance in matrix-free computations.

math.NA

A consistent diffuse-interface finite element approach to rapid melt--vapor dynamics with application to metal additive manufacturing

Metal additive manufacturing via laser-based powder bed fusion (PBF-LB/M) faces performance-critical challenges due to complex melt pool and vapor dynamics, often oversimplified by computational models that neglect crucial aspects, such as vapor jet formation. To address this limitation, we propose a consistent computational multi-physics mesoscale model to study melt pool dynamics, laser-induced evaporation, and vapor flow. In addition to the evaporation-induced pressure jump, we also resolve the evaporation-induced volume expansion and the resulting velocity jump at the liquid--vapor interface. We use an anisothermal incompressible Navier--Stokes solver extended by a conservative diffuse level-set framework and integrate it into a matrix-free adaptive finite element framework. To ensure accurate physical solutions despite extreme density, pressure and velocity gradients across the diffuse liquid--vapor interface, we employ consistent interface source term formulations developed in our previous work. These formulations consider projection operations to extend solution variables from the sharp liquid--vapor interface into the computational domain. Benchmark examples, including film boiling, confirm the accuracy and versatility of the model. As a key result, we demonstrate the model's ability to capture the strong coupling between melt and vapor flow dynamics in PBF-LB/M based on simulations of stationary laser illumination on a metal plate. Additionally, we show the derivation of the well-known Anisimov model and extend it to a new hybrid model. This hybrid model, together with consistent interface source term formulations, especially for the level-set transport velocity, enables PBF-LB/M simulations that combine accurate physical results with the robustness of an incompressible, diffuse-interface computational modeling framework.

cs.CE

Matrix-free implementation of the non-nested multigrid method

Traditionally, the geometric multigrid method is used with nested levels. However, the construction of a suitable hierarchy for very fine and unstructured grids is, in general, highly non-trivial. In this scenario, the non-nested multigrid method could be exploited in order to handle the burden of hierarchy generation, allowing some flexibility on the choice of the levels. We present a parallel, matrix-free, implementation of the non-nested multigrid method for continuous Lagrange finite elements, where each level may consist of independently partitioned triangulations. Our algorithm has been added to the multigrid framework of the C++ finite-element library deal.II. Several 2D and 3D numerical experiments are presented, ranging from Poisson problems to linear elasticity. We test the robustness and performance of the proposed implementation with different polynomial degrees and geometries.

math.NA

Matrix-Free Higher-Order Finite Element Methods for Hyperelasticity

This work presents a matrix-free finite element solver for finite-strain elasticity adopting an $hp$-multigrid preconditioner. Compared to classical algorithms relying on a global sparse matrix, matrix-free solution strategies significantly reduce memory traffic by repeated evaluation of the finite element integrals. Following this approach in the context of finite-strain elasticity, the precise statement of the final weak form is crucial for performance, and it is not clear a priori whether to choose problem formulations in the material or spatial domain. With a focus on hyperelastic solids in biomechanics, the arithmetic costs to evaluate the material law at each quadrature point might favor an evaluation strategy where some quantities are precomputed in each Newton iteration and reused in the Krylov solver for the linearized problem. Hence, we discuss storage strategies to balance the compute load against memory access in compressible and incompressible neo-Hookean models and an anisotropic tissue model. Additionally, numerical stability becomes increasingly important using lower/mixed-precision ingredients and approximate preconditioners to better utilize modern hardware architectures. Application of the presented method to a patient-specific geometry of an iliac bifurcation shows significant speed-ups, especially for higher polynomial degrees, when compared to alternative approaches with matrix-based geometric or black-box algebraic multigrid preconditioners.

cs.CE

Smoothers with localized residual computations for geometric multigrid methods

We improve the performance of multigrid solvers on many-core architectures with cache hierarchies by reorganizing operations in the smoothing step to minimize memory transfers. We focus on patch smoothers, which offer robust convergence rates with respect to the finite element degree for various equations, in the setting of multiplicative subspace correction for numerical efficiency. By combining the computation of local residuals with local solvers, we increase the locality of the problem and thus reduce data transfers. The thread-parallel implementation of this algorithm is based on coloring, which contradicts cache efficiency. We improve data locality by rearranging the loop into batches so that more data can be reused. The organization of consecutive batches prioritizes data locality.

math.NA

High-performance matrix-free unfitted finite element operator evaluation

Unfitted finite element methods, like CutFEM, have traditionally been implemented in a matrix-based fashion, where a sparse matrix is assembled and later applied to vectors while solving the resulting linear system. With the goal of increasing performance and enabling algorithms with polynomial spaces of higher degrees, this contribution chooses a more abstract approach by matrix-free evaluation of the operator action on vectors instead. The proposed method loops over cells and locally evaluates the cell, face, and interface integrals, including the contributions from cut cells and the different means of stabilization. The main challenge is the efficient numerical evaluation of terms in the weak form with unstructured quadrature points arising from the unfitted discretization in cells cut by the interface. We present design choices and performance optimizations for tensor-product elements and demonstrate the performance by means of benchmarks and application examples. We demonstrate a speedup of more than one order of magnitude for the operator evaluation of a discontinuous Galerkin discretization with polynomial degree three compared to a sparse matrix-vector product and develop performance models to quantify the performance properties over a wide range of polynomial degrees.

math.NA

A highly efficient computational approach for part-scale microstructure predictions in Ti-6Al-4V additive manufacturing

Fast and efficient simulations of metal additive manufacturing (AM) processes are highly relevant to exploring the full potential of this promising manufacturing technique. The microstructure composition plays an important role in characterizing the part quality and deriving mechanical properties. When complete parts are simulated, one often needs to resort to strong simplifications such as layer-wise heating due to the large number of simulated time steps compared to the small time step sizes. This article proposes a scan-resolved approach to the coupled thermo-microstructural problem. Building on a highly efficient thermal model, we discuss the implementation of a phenomenological microstructure model for the evolution of the three main constituents of Ti-6Al-4V: stable $α_s$-phase, martensite $α_m$-phase and $β$-phase. The implementation is tailored to modern hardware features using vectorization and fast approximations of transcendental functions. A performance model and numerical examples verify the high degree of optimization. We demonstrate the applicability and predictive power of the approach and the influence of scan strategy and geometry. Depending on the specific example, results can be obtained with moderate computational resources in a few hours to days. The numerical examples include a prediction of the microstructure on the full NIST AM Benchmark cantilever specimen.

cs.CE

Graph-based methods for hyperbolic systems of conservation laws using discontinuous space discretizations

We present a graph-based numerical method for solving hyperbolic systems of conservation laws using discontinuous finite elements. This work fills important gaps in the theory as well as practice of graph-based schemes. In particular, four building blocks required for the implementation of flux-limited graph-based methods are developed and tested: a first-order method with mathematical guarantees of robustness; a high-order method based on the entropy viscosity technique; a procedure to compute local bounds; and a convex limiting scheme. Two important features of the current work are the fact that (i) boundary conditions are incorporated into the mathematical theory as well as the implementation of the scheme. For instance, the first-order version of the scheme satisfies pointwise entropy inequalities including boundary effects for any boundary data that is admissible; (ii) sub-cell limiting is built into the convex limiting framework. This is in contrast to the majority of the existing methodologies that consider a single limiter per cell providing no sub-cell limiting capabilities. From a practical point of view, the implementation of graph-based methods is algebraic, meaning that they operate directly on the stencil of the spatial discretization. In principle, these methods do not need to use or invoke loops on cells or faces of the mesh. Finally, we verify convergence rates on various well-known test problems with differing regularity. We propose a simple test in order to verify the implementation of boundary conditions and their convergence rates.

math.NA

Improved accuracy of continuum surface flux models for metal additive manufacturing melt pool simulations

Computational modeling of the melt pool dynamics in laser-based powder bed fusion metal additive manufacturing (PBF-LB/M) promises to shed light on fundamental mechanisms of defect generation. These processes are accompanied by rapid evaporation so that the evaporation-induced recoil pressure and cooling arise as major driving forces for fluid dynamics and temperature evolution. The magnitude of these interface fluxes depends exponentially on the melt pool surface temperature, which, therefore, has to be predicted with high accuracy. The present work utilizes a diffuse interface finite element model based on a continuum surface flux (CSF) description of interface fluxes to study dimensionally reduced thermal two-phase problems representative for PBF-LB/M in a finite element framework. It is demonstrated that the extreme temperature gradients combined with the high ratios of material properties between metal and ambient gas lead to significant errors in the interface temperatures and fluxes when classical CSF approaches, along with typical interface thicknesses and discretizations, are applied. It is expected that this finding is also relevant for other types of diffuse interface PBF-LB/M melt pool models. A novel parameter-scaled CSF approach is proposed, which is constructed to yield a smoother temperature field in the diffuse interface region, significantly increasing the solution accuracy. The interface thickness required to predict the temperature field with a given level of accuracy is less restrictive by at least one order of magnitude for the proposed parameter-scaled approach compared to classical CSF, drastically reducing computational costs. Finally, we showcase the general applicability of the parameter-scaled CSF to a 3D simulation of stationary laser melting of PBF-LB/M considering the fully coupled thermo-hydrodynamic multi-phase problem, including phase change.

cs.CE

A consistent diffuse-interface model for two-phase flow problems with rapid evaporation

We present accurate and mathematically consistent formulations of a diffuse-interface model for two-phase flow problems involving rapid evaporation. The model addresses challenges including discontinuities in the density field by several orders of magnitude, leading to high velocity and pressure jumps across the liquid-vapor interface, along with dynamically changing interface topologies. To this end, we integrate an incompressible Navier-Stokes solver combined with a conservative level-set formulation and a regularized, i.e., diffuse, representation of discontinuities into a matrix-free adaptive finite element framework. The achievements are three-fold: First, we propose mathematically consistent definitions for the level-set transport velocity in the diffuse interface region by extrapolating the velocity from the liquid or gas phase. They exhibit superior prediction accuracy for the evaporated mass and the resulting interface dynamics compared to a local velocity evaluation, especially for strongly curved interfaces. Second, we show that accurate prediction of the evaporation-induced pressure jump requires a consistent, namely a reciprocal, density interpolation across the interface, which satisfies local mass conservation. Third, the combination of diffuse interface models for evaporation with standard Stokes-type constitutive relations for viscous flows leads to significant pressure artifacts in the diffuse interface region. To mitigate these, we propose to introduce a correction term for such constitutive model types. Through selected analytical and numerical examples, the aforementioned properties are validated. The presented model promises new insights in simulation-based prediction of melt-vapor interactions in thermal multiphase flows such as in laser-based powder bed fusion of metals.

cs.CE

A highly efficient computational framework for fast scan-resolved simulations of metal additive manufacturing processes on the scale of real parts

This article proposes a novel high-performance computing approach for the prediction of the temperature field in powder bed fusion (PBF) additive manufacturing processes. In contrast to many existing approaches to part-scale simulations, the underlying computational model consistently resolves physical scan tracks without additional heat source scaling, agglomeration strategies or any other heuristic modeling assumptions. A growing, adaptively refined mesh accurately captures all details of the laser beam motion. Critically, the fine spatial resolution required for resolved scan tracks in combination with the high scan velocities underlying these processes mandates the use of comparatively small time steps to resolve the underlying physics. Explicit time integration schemes are well-suited for this setting, while unconditionally stable implicit time integration schemes are employed for the interlayer cool down phase governed by significantly larger time scales. These two schemes are combined and implemented in an efficient fast operator evaluation framework providing significant performance gains and optimization opportunities. The capabilities of the novel framework are demonstrated through realistic AM examples on the centimeter scale including the first scan-resolved simulation of the entire NIST AM Benchmark cantilever specimen, with a computation time of less than one day. Apart from physical insights gained through these simulation examples, also numerical aspects are thoroughly studied on basis of weak and strong parallel scaling tests. As potential applications, the proposed thermal PBF simulation framework can serve as a basis for microstructure and thermo-mechanical predictions on the part-scale, but also to assess the influence of scan pattern and part geometry on melt pool shape and temperature, which are important indicators for well-known process instabilities.

cs.CE

Stage-parallel fully implicit Runge-Kutta implementations with optimal multilevel preconditioners at the scaling limit

We present an implementation of a fully stage-parallel preconditioner for Radau IIA type fully implicit Runge--Kutta methods, which approximates the inverse of $A_Q$ from the Butcher tableau by the lower triangular matrix resulting from an LU decomposition and diagonalizes the system with as many blocks as stages. For the transformed system, we employ a block preconditioner where each block is distributed and solved by a subgroup of processes in parallel. For combination of partial results, we either use a communication pattern resembling Cannon's algorithm or shared memory. A performance model and a large set of performance studies (including strong scaling runs with up to 150k processes on 3k compute nodes) conducted for a time-dependent heat problem, using matrix-free finite element methods, indicate that the stage-parallel implementation can reach higher throughputs when the block solvers operate at lower parallel efficiencies, which occurs near the scaling limit. Achievable speedup increases linearly with number of stages and are bounded by the number of stages. Furthermore, we show that the presented stage-parallel concepts are also applicable to the case that $A_Q$ is directly diagonalized, which requires complex arithmetic or the solution of two-by-two blocks and sequentializes parts of the algorithm. Alternatively to distributing stages and assigning them to distinct processes, we discuss the possibility of batching operations from different stages together.

math.NA