SearcharxivSearch

arXiv subjects

Shiquan Su

Publications and source records attributed to Shiquan Su.

5 recordsLinked to original sources

Portability of Fortran's 'do concurrent' on GPUs II

There continues to be growing interest in using standard language constructs for parallel and accelerated HPC computing, avoiding the need for (sometimes vendor-specific) external APIs. For Fortran applications, language features such as 'do concurrent' loops open the door for compilers to implement multi-threaded, GPU-accelerated, and even distributed multi-node code with only the standard language. Here, we explore the current status of using 'do concurrent' for GPU-accelerated Fortran applications across three major GPU vendors (NVIDIA, AMD, and Intel). Using a production application, we test their current capabilities, showing where the standard language alone can be used, and where augmenting the code with a directive-based API (e.g., OpenMP) is still desirable or required. Multi-GPU tests are performed with GPU-aware MPI libraries. We find that the three GPU vendors can now GPU-accelerate pure Fortran (zero directives), but that manual data movement directives can help with performance and compatibility. The results show that there is rapid advancement towards making GPU-accelerated scientific HPC code performance portable using the Fortran standard language.

cs.PL

Portability of Fortran's `do concurrent' on GPUs

There is a continuing interest in using standard language constructs for accelerated computing in order to avoid (sometimes vendor-specific) external APIs. For Fortran codes, the {\tt do concurrent} (DC) loop has been successfully demonstrated on the NVIDIA platform. However, support for DC on other platforms has taken longer to implement. Recently, Intel has added DC GPU offload support to its compiler, as has HPE for AMD GPUs. In this paper, we explore the current portability of using DC across GPU vendors using the in-production solar surface flux evolution code, HipFT. We discuss implementation and compilation details, including when/where using directive APIs for data movement is needed/desired compared to using a unified memory system. The performance achieved on both data center and consumer platforms is shown.

cs.PL

Refactoring the MPS/University of Chicago Radiative MHD(MURaM) Model for GPU/CPU Performance Portability Using OpenACC Directives

The MURaM (Max Planck University of Chicago Radiative MHD) code is a solar atmosphere radiative MHD model that has been broadly applied to solar phenomena ranging from quiet to active sun, including eruptive events such as flares and coronal mass ejections. The treatment of physics is sufficiently realistic to allow for the synthesis of emission from visible light to extreme UV and X-rays, which is critical for a detailed comparison with available and future multi-wavelength observations. This component relies critically on the radiation transport solver (RTS) of MURaM; the most computationally intensive component of the code. The benefits of accelerating RTS are multiple fold: A faster RTS allows for the regular use of the more expensive multi-band radiation transport needed for comparison with observations, and this will pave the way for the acceleration of ongoing improvements in RTS that are critical for simulations of the solar chromosphere. We present challenges and strategies to accelerate a multi-physics, multi-band MURaM using a directive-based programming model, OpenACC in order to maintain a single source code across CPUs and GPUs. Results for a $288^3$ test problem show that MURaM with the optimized RTS routine achieves 1.73x speedup using a single NVIDIA V100 GPU over a fully subscribed 40-core Intel Skylake CPU node and with respect to the number of simulation points (in millions) per second, a single NVIDIA V100 GPU is equivalent to 69 Skylake cores. We also measure parallel performance on up to 96 GPUs and present weak and strong scaling results.

physics.space-ph

Quenched Charmed Meson Spectra using Tadpole Improved Quark Action on Anisotropic Lattices

Charmed meson charmonium spectra are studied with improved quark actions on anisotropic lattices. We measured the pseudo-scalar and vector meson dispersion relations for 4 lowest lattice momentum modes with quark mass values ranging from the strange quark to charm quark with 3 different values of gauge coupling $β$ and 4 different values of bare speed of light $ν$. With the bare speed of light parameter $ν$ tuned in a mass-dependent way, we study the mass spectra of $D$, $D_s$, $η_c$, $D^{\ast}$, $D_s^{\ast}$ and $J/ψ$ mesons. The results extrapolated to the continuum limit are compared with the experiment and qualitative agreement is found.

hep-lat

A Numerical Study of Improved Quark Actions on Anisotropic Lattices

Tadpole improved Wilson quark actions with clover terms on anisotropic lattices are studied numerically. Using asymmetric lattice volumes, the pseudo-scalar meson dispersion relations are measured for 8 lowest lattice momentum modes with quark mass values ranging from the strange to the charm quark with various values of the gauge coupling $β$ and 3 different values of the bare speed of light parameter $ν$. These results can be utilized to extrapolate or interpolate to obtain the optimal value for the bare speed of light parameter $ν_{opt}(m)$ at a given gauge coupling for all bare quark mass values $m$. In particular, the optimal values of $ν$ at the physical strange and charm quark mass are given for various gauge couplings. The lattice action with these optimized parameters can then be used to study physical properties of hadrons involving either light or heavy quarks.

hep-lat