SearcharxivSearch

arXiv subjects

Benedict D. Rogers

Publications and source records attributed to Benedict D. Rogers.

10 recordsLinked to original sources

Memory Layouts for GPU-Data Transfer Buffering in SPH

The rise in GPU compute speed has outpaced improvements in host-to-device memory transfer speeds, despite the advent of shared-memory superchips. Consequently, memory transfer times now constitute an increasingly large fraction of total time-to-solution, compelling developers to compress GPU kernel input and output data into compact, minimal formats prior to GPU-offloading. This complements existing work on GPU- and compute-friendly data arrangements. We study a Smoothed Particle Hydrodynamics solver and propose memory layout strategies for host-side particle data that are particularly well-suited to GPU-offloading. Specifically, we advocate splitting classic array-of-struct data structures into a split array-of-struct arrangement, in which each logical struct decomposes into substructs determined by kernel read/write access patterns and attribute types. Splitting a monolithic particle struct into several bespoke, finer-grained structs can reduce the time required to pack data to and from buffers by ~20% - 40%, lowering total time spent on GPU-offloading by ~12% - 25%.

cs.DC

A high-order Fourier Continuation (FC)-based spectral incompressible Smoothed Particle Hydrodynamics (ISPH) scheme for general boundary conditions in wall-bounded domains

In this paper, a high-order Fourier Continuation (FC) algorithm is introduced into the spectral smoothed particle hydrodynamics (SPH) scheme to simulate the wall-bounded incompressible flows. This work aims to extend the spectral ISPH scheme towards the high-order simulation of flows with non-periodic wall boundary conditions. Herein, a polynomial-based Fourier continuation technique is applied to the velocity and pressure to make the domain both periodic and Cp smooth. The spatial SPH discretisation is performed subsequently in the frequency space on the FC-extended domain by building upon the convolution theorem using fast Fourier transform (FFT). The incorporation of Neumann boundary conditions is straightforward, and more generally, the FC method enforces periodicity across the domain regardless of the boundary condition type. The convergence order, additional computational cost, and implementation technique of the FC method are also discussed. Combined with a projection-based time integration scheme and a spectral PPE solver, the FC-based spectral ISPH framework is validated against several classical CFD benchmarks. The principal finding of this work is that the incorporation of FC techniques enables the spectral ISPH scheme to simulate wall-bounded flows with high-order convergence, and accurately capturing complex vortex dynamics. This work therefore represents a step towards a fully high-order spectral Lagrangian SPH solver with complex geometries

physics.flu-dyn

Task-parallelism in SWIFT for heterogeneous compute architectures

This paper highlights first steps towards enabling graphics processing unit (GPU) acceleration of the task-parallel smoothed particle hydrodynamics (SPH) solver SWIFT. Novel combinations of algorithms are presented, enabling SWIFT to function as a truly heterogeneous software leveraging task-parallelism on CPUs for memory-bound computations concurrently with GPUs for compute-bound computations while minimising the effects of CPU-GPU communication latency. The proposed algorithms are validated in extensive testing. The GPU acceleration methodology is shown to deliver up to 3.5 and 7.5 speedups for the offloaded computations when including and excluding the time required to prepare and post-process data transfers on the CPU side, respectively. The overall performance of the GPU-accelerated hydrodynamic solver for a full simulation on a single Grace-Hopper superchip is 1.8 times faster compared to the superchips fully parallelised CPU capabilities. This constitutes an improvement from 8 million particle updates/s for the full CPU-only baseline (115,000 updates per CPU core) to 15 million updates/s for the GPU-accelerated SPH solver. Moreover, it displays near-perfect strong scaling on 4 Grace-Hopper nodes. The GPU-acceleration is also demonstrated to give a 29 percent improvement in energy efficiency in comparison to CPU-only baselines. Finally, inter-influential bottlenecks in the prototype solver presented in this work are identified: A significant amount of time (up to 80 percent) of a GPU-offloading cycle is spent on preparing and post-processing particle data on the CPU for the transfer to and from the GPU, respectively. Approaches are suggested to minimise their effects and maximise the solver's performance in our future work.

cs.PF

Dynamics of jet formation and collapse for axisymmetric surface gravity waves: coupled 3D potential flow and SPH simulations

Axisymmetric waves occur across a wide range of scales. This study analyses large-scale gravity-dominated axisymmetric waves, with jet heights of up to 6 m, for which surface-tension effects are negligible. The Bond number is O(10^5) and the Weber number ranges from O(10^4) to O(10^6). Our aim is to clarify the dynamics of highly nonlinear axisymmetric jet formation, cavity collapse and the consequent generation of secondary jets. The newly developed three-dimensional framework OceanSPHysics3D, combining unsteady potential flow with smoothed particle hydrodynamics, enables full simulation of jet initiation and collapse. The computed free-surface elevations and jet evolution agree well with the experiments of McAllister et al. (Journal of Fluid Mechanics, 2022) and with an analytical jet-tip-angle formulation by Longuet-Higgins (Journal of Fluid Mechanics, 1983). The simulations elucidate how the falling primary jet induces a secondary jet. The mechanisms forming the pre-jet trough and the post-jet cavity are fundamentally different. The pre-jet trough arises geometrically from directional focusing of the constituent waves, yielding a self-similar shape when appropriately scaled. In contrast, the post-jet cavity is formed inertially by the falling continuous jet and lacks both spatial and temporal self-similarity. Its collapse also differs: the cavity pinches off at the neck to generate upward and downward secondary jets, with local accelerations reaching approximately 150 times gravity. The primary jet scale governs the ensuing secondary-jet dynamics, including vortex-ring formation and strong vertical mixing. These findings illustrate the complexity of axisymmetric jet dynamics and demonstrate the ability of the present framework to reproduce the key coupled processes in such extreme free-surface events.

physics.flu-dyn

Interaction of water surface waves with periodic and quasiperiodic cylinder arrays

Inspired by transformation optics and photonic crystals, this paper presents a computational investigation into the interaction between water surface waves and array waveguides of cylinders with multiple previously unexplored lattice geometries, including, for the first time, quasiperiodic geometries. Extending beyond conventional square and hexagonal periodic arrays, transformation optics has opened up entirely new opportunities to investigate water wave propagation through arrays based on quasiperiodic lattices, and quasiperiodically arranged vacancy defects. Using the linear potential flow open-source code Capytaine, missing element and $τ$-scaled Fibonacci square lattices, the Penrose lattice, hexagonal $H_{00}$ lattice and Amman-Beenker lattice are investigated. The existence of band gaps for all arrays is observed. An hexagonal lattice with vacancy defects transmits the least energy. Bragg diffraction consistent with rotational symmetry is observed from all arrays. Waves will distort significantly to achieve resonance with arrays, supporting transformation-based waveguides. The possible uses include adaptation to more versatile waveguides with applications such as offshore renewable energy and coastal defence.

physics.flu-dyn

Divergence cleaning for weakly compressible smoothed particle hydrodynamics

This paper presents a divergence cleaning formulation for the velocity in the weakly compressible smoothed particle hydrodynamics (SPH) scheme. The proposed hyperbolic/parabolic divergence cleaning, ensures that the velocity divergence, $div(\mathbf{u})$, is minimised throughout the simulation. The divergence equation is coupled with the momentum conservation equation through a scalar field $ψ$. A parabolic term is added to the time-evolving divergence equation, resulting in a hyperbolic/parabolic form, dissipating acoustic waves with a speed of sound proportional to the local Mach number in order to maximise dissipation of the velocity divergence, preventing unwanted diffusion of the pressure field. The $div(\mathbf{u})$-SPH algorithm is implemented in the open-source weakly compressible SPH solver DualSPHysics. The new formulation is validated against a range of challenging 2-D test cases including the Taylor-Green vortices, patch impact test, jet impinging on a surface, and wave impact in a sloshing tank. The results show that the new formulation reduces the divergence in the velocity field by at least one order of magnitude which prevents spurious numerical noise and the formation of unphysical voids. The temporal evolution of the impact pressures shows that the $div(\mathbf{u})$-SPH formulation virtually eliminates unwanted acoustic pressure oscillations. Investigation of particle resolution confirms that the new $div(\mathbf{u})$-SPH formulation does not reduce the spatial convergence rate.

physics.flu-dyn

Large Eddy Simulations of bubbly flows and breaking waves with Smoothed Particle Hydrodynamics

For turbulent bubbly flows, multi-phase simulations resolving both the liquid and bubbles are prohibitively expensive in the context of different natural phenomena. One example is breaking waves, where bubbles strongly influence wave impact loads, acoustic emissions, and atmospheric-ocean transfer, but detailed simulations in all but the simplest settings are infeasible. An alternative approach is to resolve only large scales, and model small scale bubbles adopting sub-resolution closures. Here we introduce a large eddy simulation (LES) Smoothed Particle Hydrodynamics (SPH) scheme for simulations of bubbly flows. The continuous liquid phase is resolved with a semi-implicit isothermally compressible SPH framework. This is coupled with a discrete Lagrangian bubble model. Bubbles and liquid interact via exchanges of volume and momentum, through turbulent closures, bubble breakup and entrainment, and free-surface interaction models. By representing bubbles as individual particles, they can be tracked over their lifetimes, allowing closure models for sub-resolution fluctuations, bubble deformation, breakup and free-surface interaction in integral form, accounting for the finite timescales over which these events occur. We investigate two flows: bubble plumes, and breaking waves, and find close quantitative agreement with published experimental and numerical data. In particular, for plunging breaking waves, our framework accurately predicts the Hinze scale, bubble size distribution, and growth rate of the entrained bubble population. This is the first coupling of an SPH framework with a discrete bubble model, with potential for cost effective simulations of wave-structure interactions and more accurate predictions of wave impact loads.

physics.flu-dyn

State-of-the-art SPH solver DualSPHysics: from fluid dynamics to multiphysics problems

DualSPHysics is a weakly compressible smoothed particle hydrodynamics (SPH) Navier-Stokes solver initially conceived to deal with coastal engineering problems, especially those related to wave impact with coastal structures. Since the first release back in 2011, DualSPHysics has shown to be robust and accurate for simulating extreme wave events along with a continuous improvement in efficiency thanks to the exploitation of hardware such as graphics processing units (GPUs) for scientific computing or the coupling with wave propagating models such as SWASH and OceanWave3D. Numerous additional functionalities have also been included in the DualSPHysics package over the last few years which allow the simulation of fluid-driven objects. The use of the discrete element method (DEM) has allowed the solver to simulate the interaction among different bodies (sliding rocks, for example), which provides a unique tool to analyse debris flows. In addition, the recent coupling with other solvers like Project Chrono or MoorDyn has been a milestone in the development of the solver. Project Chrono allows the simulation of articulated structures with joints, hinges, sliders and springs and MoorDyn allows simulating moored structures. Both functionalities make DualSPHysics one of the meshless model world leaders in the simulation of offshore energy harvesting devices. Lately, the present state of maturity of the solver goes beyond single phase simulations, allowing multi-phase simulations with gas-liquid and a combination of Newtonian and non-Newtonian models expanding further the capabilities and range of applications for the DualSPHysics solver. These advances and functionalities make DualSPHysics a state-of-the-art meshless solver with emphasis on free-surface flow modelling.

physics.flu-dyn

Investigation of wall bounded flows using SPH and the unified semi-analytical wall boundary conditions

The semi-analytical wall boundary conditions present a mathematically rigorous framework to prescribe the influence of solid walls in SPH for fluid flows. In this paper they are investigated with respect to the skew-adjoint property which implies exact energy conservation. It will be shown that this property holds only in the limit of the continuous SPH approximation, whereas in the discrete SPH formulation it is only approximately true, leading to numerical noise. This noise, interpreted as form of "turbulence", is treated using an additional volume diffusion term in the continuity equation which we show is equivalent to an approximate Riemann solver. Subsequently two extensions to the boundary conditions are presented. The first dealing with a variable driving force when imposing a volume flux in a periodic flow and the second showing a generalization of the wall boundary condition to Robin type and arbitrary-order interpolation. Two modifications for free-surface flows are presented for the volume diffusion term as well as the wall boundary condition. In order to validate the theoretical constructs numerical experiments are performed showing that the present volume flux term yields results with an error 5 orders of magnitude smaller then previous methods while the Robin boundary conditions are imposed correctly with an error depending on the order of the approximation. Furthermore, the proposed modifications for free-surface flows improve the behaviour at the intersection of free surface and wall as well as prevent free-surface detachment when using the volume diffusion term. Finally, this paper is concluded by a simulation of a dam break over a wedge demonstrating the improvements proposed in this paper.

physics.flu-dyn

Towards accelerating Smoothed Particle Hydrodynamics simulations for free-surface flows on multi-GPU clusters

Starting from the single graphics processing unit (GPU) version of the Smoothed Particle Hydrodynamics (SPH) code DualSPHysics, a multi-GPU SPH program is developed for free-surface flows. The approach is based on a spatial decomposition technique, whereby different portions (sub-domains) of the physical system under study are assigned to different GPUs. Communication between devices is achieved with the use of Message Passing Interface (MPI) application programming interface (API) routines. The use of the sorting algorithm radix sort for inter-GPU particle migration and sub-domain halo building (which enables interaction between SPH particles of different subdomains) is described in detail. With the resulting scheme it is possible, on the one hand, to carry out simulations that could also be performed on a single GPU, but they can now be performed even faster than on one of these devices alone. On the other hand, accelerated simulations can be performed with up to 32 million particles on the current architecture, which is beyond the limitations of a single GPU due to memory constraints. A study of weak and strong scaling behaviour, speedups and efficiency of the resulting programis presented including an investigation to elucidate the computational bottlenecks. Last, possibilities for reduction of the effects of overhead on computational efficiency in future versions of our scheme are discussed.

physics.comp-ph