SearcharxivSearch

arXiv subjects

Mark S. Shephard

Publications and source records attributed to Mark S. Shephard.

9 recordsLinked to original sources

Curved Mesh Adaptation for High-Order Finite Element Simulations

The ability to take advantage of computationally efficient high-order finite element methods to perform adaptive finite element analysis of complex engineering problems over general 3D domains requires the ability to adapt meshes with curved elements that maintain the level of geometric approximation of the domain boundary required. This paper presents a conforming curved mesh adaptation procedure aimed at effectively supporting automated adaptive analysis of problems for which the domain geometry is defined in a CAD system. The local mesh modification procedures interact with the CAD geometry to curve the mesh edges and faces representing those boundaries to the order of approximation of the high-order elements being used. The curved mesh edges and faces representing the domain boundaries are based on B\'ezier approximation geometry, which provides more accurate evaluations of surface-related quantities of interest than the commonly used interpolation methods. To attain computational efficiency during mesh adaptation, the interior mesh entities have their geometric order kept as low as possible while still maintaining control of element shapes. The order of curved mesh entities is limited to no higher than cubic, which allows the definition of an effective procedure to define the shape of the limited number of interior mesh entities that must be curved. The procedures, which are fully parallelized, are demonstrated on the adaptive radio-frequency analysis of a magnetically confined fusion system containing a fully represented radio-frequency antenna.

cs.CE

Unstructured Mesh Tools for Fusion Energy System Design

The execution of accurate simulations of fusion energy systems requires the appropriate representation of critical component geometries as well as the coupling of complex fusion physics codes with one another and with engineering analysis tools. This paper examines the challenges of creating simulation workflows that fully leverage existing fusion research codes while integrating them with commercial computer-aided engineering (CAE) software. Key areas addressed include: (a) the construction and meshing of analysis geometries taking full advantage of available geometric modeling and meshing technologies; (b) the effective coupling of fusion physics and engineering analysis codes; and (c) the support for simulation workflows that couple particle and continuum modeling methods.

cs.CE

PCMS: Parallel Coupler For Multimodel Simulations

This paper presents the Parallel Coupler for Multimodel Simulations (PCMS), a new GPU accelerated generalized coupling framework for coupling simulation codes on leadership class supercomputers. PCMS includes distributed control and field mapping methods for up to five dimensions. For field mapping PCMS can utilize discretization and field information to accommodate physics constraints. PCMS is demonstrated with a coupling of the gyrokinetic microturbulence code XGC with a Monte Carlo neutral transport code DEGAS2 and with a 5D distribution function coupling of an energetic particle transport code (GNET) to a gyrokinetic microturbulence code (GTC). Weak scaling is also demonstrated on up to 2,080 GPUs of Frontier with a weak scaling efficiency of 85%.

cs.DC

GPU Acceleration of Monte Carlo Tallies on Unstructured Meshes in OpenMC with PUMI-Tally

Unstructured mesh tallies are a bottleneck in Monte Carlo neutral particle transport simulations of fusion reactors. This paper introduces the PUMI-Tally library that takes advantage of mesh adjacency information to accelerate these tallies on CPUs and GPUs. For a fixed source simulation using track-length tallies, we achieved a speed-up of 19.7X on an NVIDIA A100, and 9.2X using OpenMP on 128 threads of two AMD EPYC 7763 CPUs on NERSC Perlmutter. On the Empire AI alpha system, we achieved a speed-up of 20X using an NVIDIA H100 and 96 threads of an Intel Xenon 8568Y+. Our method showed better scaling with number of particles and number of elements. Additionally, we observed a 199X reduction in the number of allocations during initialization and the first three iterations, with a similar overall memory consumption. And, our hybrid CPU/GPU method demonstrated a 6.69X improvement in the energy consumption over the current approach.

cs.DC

A new open source framework for multiscale modeling of fibrous materials on heterogeneous supercomputers

This article presents MuMFiM, an open source application for multiscale modeling of fibrous materials on massively parallel computers. MuMFiM uses two scales to represent fibrous materials such as biological network materials (extracellular matrix, connective tissue, etc.). It is designed to make use of multiple levels of parallelism, including distributed parallelism of the macro and microscales as well as GPU accelerated data-parallelism of the microscale. Scaling results of the GPU accelerated microscale show that solving microscale problems concurrently on the GPU can lead to a 1000x speedup over the solution of a single RVE on the GPU. In addition, we show nearly optimal strong and weak scaling results of MuMFiM on up to 128 nodes of AiMOS (Rensselaer Polytechnic Institute) which is composed of IBM AC922 nodes with 6 Volta V100 GPU and 2 20 core Power 9 CPUs each. We also show how MuMFiM can be used to solve problems of interest to the broader engineering community, in particular providing an example of the facet capsule ligament (FCL) of the human spine undergoing uniaxial extension.

cs.DC

Using Hierarchical Parallelism to Accelerate the Solution of Many Small Partial Differential Equations

This paper presents efforts to improve the hierarchical parallelism of a two scale simulation code. Two methods to improve the GPU parallel performance were developed and compared. The first used the NVIDIA Multi-Process Service and the second moved the entire sub-problem loop into a single kernel using Kokkos hierarchical parallelism and a PackedView data structure. Both approaches improved parallel performance with the second method providing the greatest improvements.

cs.DC

Efficient Exascale Discretizations: High-Order Finite Element Methods

Efficient exploitation of exascale architectures requires rethinking of the numerical algorithms used in many large-scale applications. These architectures favor algorithms that expose ultra fine-grain parallelism and maximize the ratio of floating point operations to energy intensive data movement. One of the few viable approaches to achieve high efficiency in the area of PDE discretizations on unstructured grids is to use matrix-free/partially-assembled high-order finite element methods, since these methods can increase the accuracy and/or lower the computational time due to reduced data motion. In this paper we provide an overview of the research and development activities in the Center for Efficient Exascale Discretizations (CEED), a co-design center in the Exascale Computing Project that is focused on the development of next-generation discretization software and algorithms to enable a wide range of finite element applications to run efficiently on future hardware. CEED is a research partnership involving more than 30 computational scientists from two US national labs and five universities, including members of the Nek5000, MFEM, MAGMA and PETSc projects. We discuss the CEED co-design activities based on targeted benchmarks, miniapps and discretization libraries and our work on performance optimizations for large-scale GPU architectures. We also provide a broad overview of research and development activities in areas such as unstructured adaptive mesh refinement algorithms, matrix-free linear solvers, high-order data visualization, and list examples of collaborations with several ECP and external applications.

cs.DC

Boundary Layer Adaptivity For Incompressible Turbulent Flows

Boundary layers in turbulent flows require fine grid spacings near the walls which depend on the choice of turbulence model. To satisfy these requirements a semi-structured mesh is generally used in this area with orthogonal and layered elements. Adaptation of such a mesh needs to take into account the flow physics along with the standard error indicator approach. In this paper a novel methodology which combines Hessian based error indicators with flow physics to drive mesh adaptation is illustrated. Particular focus is on the thickness adaptation of the layered mesh. The technique is applied to two turbulent incompressible flow cases and its effectiveness is studied.

physics.flu-dyn

Anisotropic Boundary Layer Adaptivity of Multi-Element Wings

Multi-element wings are popular in the aerospace community due to their high lift performance. Turbulent flow simulations of these configurations require very fine mesh spacings especially near the walls, thereby making use of a boundary layer mesh necessary. However, it is difficult to accurately determine the required mesh resolution a priori to the simulations. In this paper we use an anisotropic adaptive meshing approach including adaptive control of elements in the boundary layers and study its effectiveness for two multi-element wing configurations. The results are compared with experimental data as well as nested refinements to show the efficiency of adaptivity driven by error indicators, where superior resolution in wakes and near the tip region through adaptivity are highlighted.

physics.flu-dyn