SearcharxivSearch

arXiv subjects

Massimiliano Fatica

Publications and source records attributed to Massimiliano Fatica.

14 recordsLinked to original sources

GPU-Accelerated Simulations of Moving Boundary Problems and Fluid-Structure Interaction at Extreme Scales

Computational fluid dynamics and fluid-structure interaction simulations involving moving and deforming bodies is extremely hard. In this work, we present a graphical processing unit (GPU) optimized implementation of the sharp-interface immersed boundary method. The method allows performing simulation around complex stationary as well as moving bodies on a Cartesian grid. We base our implementation on the ViCar3D framework and make use of OpenACC, CUDA, NCCL and MPI. We test the implementation across grid sizes ranging from O(10million) to O(1billion) points and achieved a 20X speedup compared to existing CPU implementation. We next present our multi-GPU implementation by utilizing CUDA streams and NCCL communicators. This enables us to obtain a >90% strong and weak scaling efficiencies. Next we demonstrate the capability of the developed software to simulate a turbulent fluid flow and coupled fluid-structure interaction in flapping bat wing in flight at Re=5000.

physics.comp-ph

A GPU-Accelerated Sharp Interface Immersed Boundary Solver for Large Scale Flow Simulations

Immersed boundary methods (IBMs) facilitate the simulation of flows around stationary, moving, and deforming bodies on Cartesian grids. However, extending these simulations to the large grid sizes required for realistic flow problems remains a significant computational challenge. In this work, we present the implementation and acceleration of ViCar3D, a sharp-interface immersed boundary solver, on graphical processing units (GPUs). We utilize OpenACC, CUDA Fortran and MPI to reprogram \emph{ViCar3D}, a sharp-interface immersed boundary solver, on multi-GPU architectures. Verification and scalability studies are performed for two benchmark cases: two-dimensional flow past a circular cylinder and direct numerical simulation (DNS) of flow past a finite rectangular wing. For the latter, we observe an approximately 20X speedup (node-to-node comparison) relative to the CPU-based implementation. The GPU-accelerated solver is capable of simulating complex 3D flows with up to 200 million mesh points on a single node equipped with four GPUs. Strong and weak scaling tests demonstrate maximum scaling efficiencies of 92\% and 93\%, respectively, on multi-GPU systems. We further test the code to simulate fluid flow past complex-shaped single-body and multi-body cases.

physics.flu-dyn

A fast incompressible Navier-Stokes solver for non-uniform grids

We present a scalable incompressible Navier--Stokes solver for three-dimensional Cartesian grids with non-uniform spacing. A direct tensor-product--Thomas method solves the constant-coefficient Poisson and Helmholtz equations arising from pressure projection and implicit diffusion, without approximate factorization. On uniform grids, the method recovers the classical eigenfunction-expansion method evaluated with fast Fourier transforms (FFTs). On stretched grids, diagonal scaling symmetrizes the one-dimensional Laplace operators, and the resulting numerical eigenbasis transforms are evaluated as general matrix--matrix multiplications (GEMMs). FFTs and GEMMs can be selected independently in each diagonalized direction while retaining the pencil decomposition, collective transposes, and tridiagonal machinery of an established FFT-based solver. The elliptic solver is verified to round-off accuracy, and the complete flow solver is validated against benchmark flows. Against geometric multigrid and block cyclic reduction with FFT diagonalization, the present method achieves the lowest time-to-solution among the tested approaches; at fixed grid dimensions, its cost is insensitive to grid stretching. CPU and multi-GPU tests show that GEMM-rich variants attain higher strong-scaling efficiency by better amortizing communication. On a single GPU, the fully GEMM-based variant increases the Poisson cost by $2.8\times$ but the complete Navier--Stokes step cost by only $1.8\times$. Weak scaling exposes the trade-off: dense-transform costs grow with the global transform dimension, whereas non-uniform meshes can reduce the required number of grid points. The resulting open-source solver, \texttt{CaNS-EIGEN}, extends an FFT-based Navier--Stokes solver to large-scale simulations on grids stretched in multiple directions.

physics.comp-ph

Universal exotic dynamics in critical mesoscopic systems: Simulating the square root of Avogadro's number of spins

We explicitly demonstrate the universality of critical dynamics through unprecedented large-scale GPU-based simulations of two out-of-equilibrium processes, comparing the behavior of spin-$1/2$ Ising and spin-$1$ Blume-Capel models on a square lattice. In the first protocol, a completely disordered system is instantaneously brought into contact with a thermal bath at the critical temperature, allowing it to evolve until the coherence length exceeds $10^{3}$ lattice spacings. Finite-size effects are negligible due to the mesoscopic scale of the lattice sizes studied, with linear dimensions up to $L=2^{22}$ and $2^{19}$ for the Ising and Blume-Capel models, respectively. Our numerical data, and the subsequent analysis, demonstrate a strong dynamic universality between the two models and provide the most precise estimate to date of the dynamic critical exponent for this universality class, $z = 2.1676(1)$. In the second protocol, we corroborate the role of the universal ratio of dynamic and static length scales in achieving an exponential acceleration in the approach to equilibrium just \emph{above} the critical temperature, through a time-dependent variation of the thermal bath temperature. The results presented in this work leverage our CUDA-based numerical code, breaking the world record for the simulation speed of the Ising model.

cond-mat.stat-mech

Massive-Scale Simulations of 2D Ising and Blume-Capel Models on Rack-Scale Multi-GPU Systems

We present high-performance implementations of the two-dimensional Ising and Blume-Capel models for large-scale, multi-GPU simulations. Our approach takes full advantage of the NVIDIA GB200 NVL72 system, which features up to $72$ GPUs interconnected via high-bandwidth NVLink, enabling direct GPU-to-GPU memory access across multiple nodes. By utilizing Fabric Memory and an optimized Monte Carlo kernel for the Ising model, our implementation supports simulations of systems with linear sizes up to $L=2^{23}$, corresponding to approximately $70$ trillion spins. This allows for a peak processing rate of nearly $1.15 \times 10^5$ lattice updates per nanosecond-setting a new performance benchmark for Ising model simulations. Additionally, we introduce a custom protocol for computing correlation functions, which strikes an optimal balance between computational efficiency and statistical accuracy. This protocol enables large-scale simulations without incurring prohibitive runtime costs. Benchmark results show near-perfect strong and weak scaling up to $64$ GPUs, demonstrating the effectiveness of our approach for large-scale statistical physics simulations.

physics.comp-ph

Improving computation efficiency using input and architecture features for a virtual screening application

Virtual screening is an early stage of the drug discovery process that selects the most promising candidates. In the urgent computing scenario it is critical to find a solution in a short time frame. In this paper, we focus on a real-world virtual screening application to evaluate out-of-kernel optimizations, that consider input and architecture features to improve the computation efficiency on GPU. Experiment results on a modern supercomputer node show that we can almost double the performance. Moreover, we implemented the optimization using SYCL and it provides a consistent benefit with the CUDA optimization. A virtual screening campaign can use this gain in performance to increase the number of evaluated candidates, improving the probability of finding a drug.

cs.CE

GPU-optimized Approaches to Molecular Docking-based Virtual Screening in Drug Discovery: A Comparative Analysis

COVID-19 has shown the importance of having a fast response against pandemics. Finding a novel drug is a very long and complex procedure, and it is possible to accelerate the preliminary phases by using computer simulations. In particular, virtual screening is an in-silico phase that is needed to filter a large set of possible drug candidates to a manageable number. This paper presents the implementations and a comparative analysis of two GPU-optimized implementations of a virtual screening algorithm targeting novel GPU architectures. The first adopts a traditional approach that spreads the computation required to evaluate a single molecule across the entire GPU. The second uses a batched approach that exploits the parallel architecture of the GPU to evaluate more molecules in parallel, without considering the latency to process a single molecule. The paper describes the advantages and disadvantages of the proposed solutions, highlighting implementation details that impact the performance. Experimental results highlight the different performance of the two methods on several target molecule databases while running on NVIDIA A100 GPUs. The two implementations have a strong dependency with respect to the data to be processed. For both cases, the performance is improving while reducing the dimension of the target molecules (number of atoms and rotatable bonds). The two methods demonstrated a different behavior with respect to the size of the molecule database to be screened. While the latency one reaches sooner (with fewer molecules) the performance plateau in terms of throughput, the batched one requires a larger set of molecules. However, the performances after the initial transient period are much higher (up to 5x speed-up). Finally, to check the efficiency of both implementations we deeply analyzed their workload characteristics using the instruction roof-line methodology.

cs.DC

One-point statistics for turbulent pipe flow up to $Re_τ \approx 6000$

We study turbulent flows in a smooth straight pipe of circular cross--section up to $Re_τ \approx 6000$ using direct--numerical-simulation (DNS) of the Navier--Stokes equations. The DNS results highlight systematic deviations from Prandtl friction law, amounting to about $2\%$, which would extrapolate to about $4\%$ at extreme Reynolds numbers. Data fitting of the DNS friction coefficient yields an estimated von Kármán constant $k \approx 0.387$, which nicely fits the mean velocity profile, and which supports universality of canonical wall-bounded flows. The same constant also applies to the pipe centerline velocity, thus providing support for the claim that the asymptotic state of pipe flow at extreme Reynolds numbers should be plug flow. At the Reynolds numbers under scrutiny, no evidence for saturation of the logarithmic growth of the inner peak of the axial velocity variance is found. Although no outer peak of the velocity variance directly emerges in our DNS, we provide strong evidence that it should appear at $Re_τ \gtrsim 10^4$, as a result of turbulence production exceeding dissipation over a large part of the outer wall layer, thus invalidating the classical equilibrium hypothesis.

physics.flu-dyn

FSEI-GPU: GPU accelerated simulations of the fluid-structure-electrophysiology interaction in the left heart

The reliability of cardiovascular computational models depends on the accurate solution of the hemodynamics, the realistic characterization of the hyperelastic and electric properties of the tissues along with the correct description of their interaction. The resulting fluid-structure-electrophysiology interaction (FSEI) thus requires an immense computational power, usually available in large supercomputing centers, and requires long time to obtain results even if multi-CPU processors are used (MPI acceleration). In recent years, graphics processing units (GPUs) have emerged as a convenient platform for high performance computing, as they allow for considerable reductions of the time-to-solution. This approach is particularly appealing if the tool has to support medical decisions that require solutions within reduced times and possibly obtained by local computational resources. Accordingly, our multi-physics solver has been ported to GPU architectures using CUDA Fortran to tackle fast and accurate hemodynamics simulations of the human heart without resorting to large-scale supercomputers. This work describes the use of CUDA to accelerate the FSEI on heterogeneous clusters, where both the CPUs and GPUs are used in synergistically with minor modifications of the original source code. The resulting GPU accelerated code solves a single heartbeat within a few hours (from three to ten depending on the grid resolution) running on premises computing facility made of few GPU cards, which can be easily installed in a medical laboratory or in a hospital, thus opening towards a systematic computational fluid dynamics (CFD) aided diagnostic.

physics.flu-dyn

GPU acceleration of CaNS for massively-parallel direct numerical simulations of canonical fluid flows

This work presents the GPU acceleration of the open-source code CaNS for very fast massively-parallel simulations of canonical fluid flows. The distinct feature of the many-CPU Navier-Stokes solver in CaNS is its fast direct solver for the second-order finite-difference Poisson equation, based on the method of eigenfunction expansions. The solver implements all the boundary conditions valid for this type of problems in a unified framework. Here, we extend the solver for GPU-accelerated clusters using CUDA Fortran. The porting makes extensive use of CUF kernels and has been greatly simplified by the unified memory feature of CUDA Fortran, which handles the data migration between host (CPU) and device (GPU) without defining new arrays in the source code. The overall implementation has been validated against benchmark data for turbulent channel flow and its performance assessed on a NVIDIA DGX-2 system (16 Tesla V100 32Gb, connected with NVLink via NVSwitch). The wall-clock time per time step of the GPU-accelerated implementation is impressively small when compared to its CPU implementation on state-of-the-art many-CPU clusters, as long as the domain partitioning is sufficiently small that the data resides mostly on the GPUs. The implementation has been made freely available and open-source under the terms of an MIT license.

physics.flu-dyn

A Performance Study of the 2D Ising Model on GPUs

The simulation of the two-dimensional Ising model is used as a benchmark to show the computational capabilities of Graphic Processing Units (GPUs). The rich programming environment now available on GPUs and flexible hardware capabilities allowed us to quickly experiment with several implementation ideas: a simple stencil-based algorithm, recasting the stencil operations into matrix multiplies to take advantage of Tensor Cores available on NVIDIA GPUs, and a highly optimized multi-spin coding approach. Using the managed memory API available in CUDA allows for simple and efficient distribution of these implementations across a multi-GPU NVIDIA DGX-2 server. We show that even a basic GPU implementation can outperform current results published on TPUs and that the optimized multi-GPU implementation can simulate very large lattices faster than custom FPGA solutions.

cs.DC

Exascale Deep Learning for Climate Analytics

We extract pixel-level masks of extreme weather patterns using variants of Tiramisu and DeepLabv3+ neural networks. We describe improvements to the software frameworks, input pipeline, and the network training algorithms necessary to efficiently scale deep learning on the Piz Daint and Summit systems. The Tiramisu network scales to 5300 P100 GPUs with a sustained throughput of 21.0 PF/s and parallel efficiency of 79.0%. DeepLabv3+ scales up to 27360 V100 GPUs with a sustained throughput of 325.8 PF/s and a parallel efficiency of 90.7% in single precision. By taking advantage of the FP16 Tensor Cores, a half-precision version of the DeepLabv3+ network achieves a peak and sustained throughput of 1.13 EF/s and 999.0 PF/s respectively.

cs.DC

AFiD-GPU: a versatile Navier-Stokes Solver for Wall-Bounded Turbulent Flows on GPU Clusters

The AFiD code, an open source solver for the incompressible Navier-Stokes equations ({\color{blue}\burl{http://www.afid.eu}}), has been ported to GPU clusters to tackle large-scale wall-bounded turbulent flow simulations. The GPU porting has been carried out in CUDA Fortran with the extensive use of kernel loop directives (CUF kernels) in order to have a source code as close as possible to the original CPU version; just a few routines have been manually rewritten. A new transpose scheme, which is not limited to the GPU version only and can be generally applied to any CFD code that uses pencil distributed parallelization, has been devised to improve the scaling of the Poisson solver, the main bottleneck of incompressible solvers. The GPU version can reduce the wall clock time by an order of magnitude compared to the CPU version for large meshes. Due to the increased performance and efficient use of memory, the GPU version of AFiD can perform simulations in parameter ranges that are unprecedented in thermally-driven wall-bounded turbulence. To verify the accuracy of the code, turbulent Rayleigh-Bénard convection and plane Couette flow are simulated and the results are in good agreement with the experimental and computational data that published in previous literatures.

physics.flu-dyn

GPU peer-to-peer techniques applied to a cluster interconnect

Modern GPUs support special protocols to exchange data directly across the PCI Express bus. While these protocols could be used to reduce GPU data transmission times, basically by avoiding staging to host memory, they require specific hardware features which are not available on current generation network adapters. In this paper we describe the architectural modifications required to implement peer-to-peer access to NVIDIA Fermi- and Kepler-class GPUs on an FPGA-based cluster interconnect. Besides, the current software implementation, which integrates this feature by minimally extending the RDMA programming model, is discussed, as well as some issues raised while employing it in a higher level API like MPI. Finally, the current limits of the technique are studied by analyzing the performance improvements on low-level benchmarks and on two GPU-accelerated applications, showing when and how they seem to benefit from the GPU peer-to-peer method.

physics.comp-ph