SearcharxivSearch

arXiv subjects

Olaf Willocx

Publications and source records attributed to Olaf Willocx.

2 recordsLinked to original sources

Astrophysics on GPUs: introducing AGILE 1.0

We present AGILE, a GPU-enabled adaptive mesh refinement (AMR) framework for the solution of (near-) conservation laws which occur in astro- and solar-physical applications. AGILE is written in modern fortran 2003, inherits a part of its modules and mesh handling from MPI-AMRVAC, and achieves excellent GPU performance via OpenACC offloading. We here discuss the design decisions which enable AGILE to perform cost-efficient and scalable deeply nested AMR simulations with moderate block sizes of e.g. $16^3$ cells. AGILE currently implements several physics modules, ie. hydrodynamics, frozen-field hydrodynamics, magnetohydrodynamics and special-relativistic hydrodynamics and can easily be extended further through its modular design. Besides strong scaling tests to up to 2048 GPUs and standard benchmarks which show consistent performance across a large range of devices and problem sizes, we demonstrate AGILE's capabilities by means of state-of-the art science applications with all currently available physics modules.

astro-ph.IM

foap4: Adaptive mesh refinement with OpenACC, MPI, and p4est

GPUs and other accelerators are increasingly used for scientific computing. In the future, we want to add GPU support to parallel adaptive mesh refinement (AMR) codes written in Fortran. To understand which changes are necessary to obtain good performance we have developed foap4, an AMR framework implemented in Fortran that uses OpenACC, MPI, and the p4est library. We discuss the design and implementation of the framework. Several benchmark problems are considered, in which Euler's equations of gas dynamics are solved using explicit time integration. These benchmarks are performed in both 2D and 3D, using static and adaptive meshes, for varying problem sizes on different hardware. Our results show that AMR simulations can be carried out efficiently on GPUs with OpenACC and MPI, even when using relatively small grid blocks of $8^3$ or $16^3$ cells.

physics.comp-ph