SearcharxivSearch

arXiv subjects

Spencer Starr

Publications and source records attributed to Spencer Starr.

4 recordsLinked to original sources

APU-Accelerated Large Eddy Simulation with the Discontinuous Galerkin Solver GAL{\AE}XI

The exascale computing era, driven by heterogeneous GPU architectures, requires a fundamental redesign of traditional CFD solvers to fully leverage those heterogeneous systems. The discontinuous Galerkin spectral element method (DGSEM) provides an ideal foundation for this transition due to its high-order accuracy and local computational stencil. This work presents recent advances in the development and application of the architecture-agnostic DGSEM framework GAL{\AE}XI by linking hardware optimization, software implementation, and physical validation. The performance of GAL{\AE}XI on the AMD MI300A Accelerated Processing Units (APUs) featured on the Hunter supercomputer is analyzed. Specifically, evaluations of the strong and weak scaling performance and the impact of the compute partitioning modes available on the AMD MI300As are performed. Second, the strategy used to integrate the algorithms necessary for wall-modeled large eddy simulations into the GPU-accelerated framework is outlined. Validation of those algorithms is presented in the form of a plane turbulent channel testcase. Finally, the solver is applied to a demanding flow problem in the form of a wall-resolved large eddy simulation of a transonic compressor cascade. The results from this investigation demonstrate the capabilities of GAL{\AE}XI to accurately capture complex shock-wave/turbulent boundary-layer interactions.

physics.flu-dyn

An Architecture-Agnostic High-Order Discontinuous Galerkin Framework for Compressible Flows

With the recent proliferation of heterogeneous, GPU-accelerated supercomputers, high-order computational fluid dynamics (CFD) simulations of complex, turbulent flows are more accessible than ever. To leverage the computing power of these machines, CFD software must adapt. However, complicating the situation is the emerging need to support hardware from multiple GPU vendors. Addressing this need is the GPU-accelerated, discontinuous Galerkin spectral element method (DGSEM) framework GAL{\AE}XI, a high-order, open source, architecture-agnostic toolchain for the study of complex, compressible, turbulent flows on unstructured, hexahedral grids. GPU-accelerated computations with GAL{\AE}XI are possible on GPU hardware by interfacing Fortran source code to the vendor models CUDA C++ for NVIDIA and HIP C++ for AMD. The DGSEM implementation in GAL{\AE}XI was verified using the method of manufactured solutions to rigorously confirm the expected order of convergence. Simulations of a compressible Taylor-Green-Vortex also demonstrated excellent agreement with reference solutions across all supported architectures. GAL{\AE}XI achieved near ideal strong and weak scaling on GPU hardware from both NVIDIA and AMD. In the largest case, GAL{\AE}XI performed a simulation with 67.1 billion degrees of freedom on 65,536 AMD MI250X graphics compute devices with a parallel efficiency of 82.6%. Comparing node-to-node performance, GPU simulations offered speedups between 7.75x and 8.08x over CPU computations in time-to-solution while consuming less than half the energy. To demonstrate GAL{\AE}XI's effectiveness for production-scale simulations, wall-resolved large eddy simulations of the transonic flow past a NACA 64A-110 airfoil and an ONERA OAT15A airfoil under shock buffet conditions were computed.

physics.flu-dyn

GAL{\AE}XI: Solving complex compressible flows with high-order discontinuous Galerkin methods on accelerator-based systems

This work presents GALAEXI as a novel, energy-efficient flow solver for the simulation of compressible flows on unstructured meshes leveraging the parallel computing power of modern Graphics Processing Units (GPUs). GALAEXI implements the high-order Discontinuous Galerkin Spectral Element Method (DGSEM) using shock capturing with a finite-volume subcell approach to ensure the stability of the high-order scheme near shocks. This work provides details on the general code design, the parallelization strategy, and the implementation approach for the compute kernels with a focus on the element local mappings between volume and surface data due to the unstructured mesh. GALAEXI exhibits excellent strong scaling properties up to 1024 GPUs if each GPU is assigned a minimum of one million degrees of freedom degrees of freedom. To verify its implementation, a convergence study is performed that recovers the theoretical order of convergence of the implemented numerical schemes. Moreover, the solver is validated using both the incompressible and compressible formulation of the Taylor-Green-Vortex at a Mach number of 0.1 and 1.25, respectively. A mesh convergence study shows that the results converge to the high-fidelity reference solution and that the results match the original CPU implementation. Finally, GALAEXI is applied to a large-scale wall-resolved large eddy simulation of a linear cascade of the NASA Rotor 37. Here, the supersonic region and shocks at the leading edge are captured accurately and robustly by the implemented shock-capturing approach. It is demonstrated that GALAEXI requires less than half of the energy to carry out this simulation in comparison to the reference CPU implementation. This renders GALAEXI as a potent tool for accurate and efficient simulations of compressible flows in the realm of exascale computing and the associated new HPC architectures.

cs.MS

Supercurrent Diode Effect, Spin Torques, and Robust Zero-Energy Peak in Planar Half-Metallic Trilayers

We consider a Josephson junction with ${\rm F_1 F_2 F_3}$ ferromagnetic trilayers in the ballistic regime, where the magnetization in each ferromagnet ${\rm F}_i (i=1,2,3)$, can have arbitrary orientations and magnetization strengths. The trilayers are sandwiched between two $s$-wave superconductors with a macroscopic phase difference $Δφ$. A broad range of magnetization strengths of the central $\rm F_2$ layer are considered, from an unpolarized normal metal (N) to a half-metallic phase, supporting only one spin species. Our results reveal that when the magnetization configuration in ${\rm F_1 F_2 F_3}$ has three orthogonal components, a supercurrent can flow at $Δφ=0$, and a strong second harmonic in the current-phase relation appears. Upon increasing the magnetization strength in the central ferromagnet layer up to the half-metallic limit, the self-biased current and second harmonic component become dramatically enhanced, and the critical supercurrent reaches its maximum value. The higher harmonics in the current-phase relations can be controlled by the relative magnetization orientations, with negligible current damping compared to the corresponding ${\rm F_1 N F_3}$ counterparts. For a broad range of exchange field strengths in the central ferromagnet ${\rm F}_2$, the ground state of the system can be tuned to an arbitrary phase difference $φ_0$ by rotating the magnetization in the outer ferromagnet $\rm F_3$. For intermediate exchange field strengths in ${\rm F}_2$, a $φ_0$ state can arise that creates a superconducting diode effect, whereby $Δφ$ can be tuned to create a one-way dissipationless current flow. The density of states demonstrates the emergence of zero energy peaks for the mutually orthogonal magnetization configurations, which is strongest in the half-metallic phase.

cond-mat.supr-con