Searcharxiv⌕ Search

arXiv subjects

Guillaume Latu

Publications and source records attributed to Guillaume Latu.

8 recordsLinked to original sources

Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive Laws

Many applications require solving large numbers of independent linear systems on GPUs. While this need is well addressed for small to large systems, tiny ones, understood here as systems of dimension below 32, remain challenging. This is especially relevant in constitutive law evaluation, where millions of integration points are handled independently, and where each constitutive update generally relies on a Newton iterative method. Each iteration then requires the double-precision solution of a tiny general square linear system using LU factorization with partial pivoting (LUpp). The present study was conducted within a closed-source prototype, which serves as a demonstrator for porting to NVIDIA GPUs constitutive law evaluations currently provided on CPUs by TFEL/MFront, an open-source code generation tool for material knowledge. We compare several double-precision LUpp solvers, including implementations from GPU linear algebra libraries as well as custom-designed CUDA kernels. The comparison is performed first on large batches of standalone linear systems, and then within the full constitutive-law evaluation workflow, where each integration point requires a sequence of distinct linear systems, one per iteration of its own Newton loop. The study shows that the best LUpp solving strategy strongly depends on several factors including system size and application context. We discuss several key aspects, including register pressure, occupancy, the ability to invoke the LUpp solver directly from device code, and whether assigning several threads to each system is the most efficient strategy. Experiments on an NVIDIA H100 GPU show that the specialized LUpp solvers proposed in this work can outperform existing state-of-the-art approaches for this class of workloads, with speedups of up to 6.5x over cuSolverDx, and up to 17.7x over MAGMA.

cs.DC↗

Parallel simulation and adaptive mesh refinement for 3D elastostatic contact mechanics problems between deformable bodies

Parallel implementation of numerical adaptive mesh refinement (AMR)strategies for solving 3D elastostatic contact mechanics problems is an essential step toward complex simulations that exceed current performance levels. This paper introduces a scalable, robust, and efficient algorithm to deal with 2D and 3D elastostatics contact problems between deformable bodies in a finite element framework. The proposed solution combines a treatment of the contact problem by a node-to-node pairing algorithm with a penalization technique and a non-conforming h-adaptive refinement of quadrilateral/hexahedral meshes based on an estimate-mark-refine approach in a parallel framework. One of the special features of our parallel strategy is that contact paired nodes are hosted by the same MPI tasks, which reduces the number of exchanges between processes for building the contact operator. The mesh partitioning introduced in this paper respects this rule and is based on an equidistribution of elements over processes, without any other constraints. In order to preserve the domain curvature while hierarchical mesh refinement, super-parametric elements are used. This functionality enables the contact zone to be well detected during the AMR process, even for an initial coarse mesh and low-order discretization schemes. The efficiency of our contact-AMR-HPC strategy is assessed on 2D and 3D Hertzian contact problems. Different AMR detection criteria are considered. Various convergence analyses are conducted. Parallel performances up to 1024 cores are illustrated. Furthermore, memory footprint and preconditionners performance are analyzed.

math.NA↗

Targeting realistic geometry in Tokamak code Gysela

In magnetically confined plasmas used in Tokamak, turbulence is responsible for specific transport that limits the performance of this kind of reactors. Gyrokinetic simulations are able to capture ion and electron turbulence that give rise to heat losses, but require also state-of-the-art HPC techniques to handle computation costs. Such simulations are a major tool to establish good operating regime in Tokamak such as ITER, which is currently being built. Some of the key issues to address more realistic gyrokinetic simulations are: efficient and robust numerical schemes, accurate geometric description, good parallelization algorithms. The framework of this work is the Semi-Lagrangian setting for solving the gyrokinetic Vlasov equation and the Gyseka code. In this paper, a new variant for the interpolation method is proposed that can handle the mesh singularity in the poloidal plane at r=0 (polar system is used for the moment in Gysela). A non-uniform meshing of the poloidal plane is proposed instead of uniform one in order to save memory and computations. The interpolation method, the gyroaverage operator, and the Poisson solver are revised in order to cope with non-uniform meshes. A mapping that establish a bijection from polar coordinates to more realistic plasma shape is used to improve realism. Convergence studies are provided to establish the validity and robustness of our new approach.

physics.plasm-ph↗

Dynamic Load Balancing for PIC code using Eulerian/Lagrangian partitioning

This document presents an analysis of different load balance strategies for a Plasma physics code that models high energy particle beams with PIC method. A comparison of different load balancing algorithms is given: static or dynamic ones. Lagrangian and Eulerian partitioning techniques have been investigated.

cs.DC↗

Some numerical aspects of the conservative PSM scheme in a 4D drift-kinetic code

The purpose of this work is simulation of magnetised plasmas in the ITER project framework. In this context, kinetic Vlasov-Poisson like models are used to simulate core turbulence in the tokamak in a toroidal geometry. This leads to heavy simulations because a 6D dimensional problem has to be solved, even if reduced to a 5D in so called gyrokinetic models. Accurate schemes, parallel algorithms need to be designed to bear these simulations. This paper describes the numerical studies to improve robustness of the conservative PSM scheme in the context of its development in the GYSELA code. In this paper, we only consider the 4D drift-kinetic model which is the backbone of the 5D gyrokinetic models and relevant to build a robust and accurate numerical method.

math.NA↗

Non regression testing for the JOREK code

Non Regression Testing (NRT) aims to check if software modifications result in undesired behaviour. Suppose the behaviour of the application previously known, this kind of test makes it possible to identify an eventual regression, a bug. Improving and tuning a parallel code can be a time-consuming and difficult task, especially whenever people from different scientific fields interact closely. The JOREK code aims at investing Magnetohydrodynamic (MHD) instabilities in a Tokamak plasma. This paper describes the NRT procedure that has been tuned for this simulation code. Automation of the NRT is one keypoint to keeping the code healthy in a source code repository.

cs.DC↗

Accuracy of unperturbed motion of particles in a gyrokinetic semi-Lagrangian code

Inaccurate description of the equilibrium can yield to spurious effects in gyrokinetic turbulence simulations. Also, the Vlasov solver and time integration schemes impact the conservation of physical quantities, especially in long-term simulations. Equilibrium and Vlasov solver have to be tuned in order to preserve constant states (equilibrium) and to provide good conservation property along time (mass to begin with). Several illustrative simple test cases are given to show typical spurious effects that one can observes for poor settings. We explain why Forward Semi-Lagrangian scheme bring us some benefits. Some toroidal and cylindrical GYSELA runs are shown that use FSL.

math.NA↗

Test of some numerical limiters for the conservative PSM scheme for 4D Drift-Kinetic simulations

The purpose of this work is simulation of magnetised plasmas in the ITER project framework. In this context, Vlasov-Poisson like models are used to simulate core turbulence in the tokamak in a toroidal geometry. This leads to heavy simulation because a 6D dimensional problem has to be solved, 3D in space and 3D in velocity. The model is reduced to a 5D gyrokinetic model, taking advantage of the particular motion of particles due to the presence of a strong magnetic field. However, accurate schemes, parallel algorithms need to be designed to bear these simulations. This paper describes a Hermite formulation of the conservative PSM scheme which is very generic and allows to implement different semi-Lagrangian schemes. We also test and propose numerical limiters which should improve the robustness of the simulations by diminishing spurious oscillations. We only consider here the 4D drift-kinetic model which is the backbone of the 5D gyrokinetic models and relevant to build a robust and accurate numerical method.

math.NA↗