Searcharxiv⌕ Search

arXiv subjects

Francis X. Giraldo

Publications and source records attributed to Francis X. Giraldo.

16 recordsLinked to original sources

GPU Performance of an Entropy-Stable Discontinuous Galerkin Euler Solver with Non-Conservative Terms

The entropy-stable discontinuous Galerkin method for compressible Euler equations with buoyancy is implemented on graphics processing unit (GPU) hardware. We measure the performance of the solver on three-dimensional problems: the rising thermal bubble and the baroclinic instability in a channel. On NVIDIA A100 hardware, the solver achieves nearly 70\% of 64-bit floating-point peak performance for the most computationally expensive kernel (volume terms) and significantly reduces the computational overhead typically incurred by two point entropy-stable fluxes in the volume terms. We also present impressive strong and weak scaling performance of the solver and compare to a highly-optimized central processing unit (CPU) code showing that the GPU kernels are a factor of $10\times$ faster and better than $13\times$ more energy efficient than the CPU code. We also show that the solver achieves the expected $2\times$ speedup when run at 32-bit floating-point peak performance. We discuss the different modifications that we implemented to reach the final form of the GPU implementation and measure the performance gain of each of the implementation strategies ranging from reduction in complex operations and memory traffic as well as load balancing. We also extend symmetry-based flux savings to the non-symmetric gravity term, preserving nearly the full factor-of-two speedup achieved for the symmetric flux.

math.NA↗

A GPU-accelerated simulation of rapid intensification of a tropical cyclone with observed heating

This paper presents a limited-area atmospheric simulation of a tropical cyclone accelerated using GPUs. The OpenACC directive-based programming model is used to port the atmospheric model to the GPU. The GPU implementation of the main functions and kernels is discussed. The GPU-accelerated code produces high-fidelity simulations of a realistic tropical cyclone forced by observational latent heating. Performance tests show that the GPU-accelerated code yields energy-efficient simulations and scales well in both the strong and weak limit.

physics.comp-ph↗

A Nonhydrostatic Mass-Conserving Dynamical Core for Deep Atmospheres of Variable Composition

This paper develops and tests a deep-atmosphere, nonhydrostatic dynamical core (DyCore) targeted towards ground-thermosphere atmospheric prediction using the spectral element method (SEM) with Implicit-Explicit (IMEX) and Horizontally Explicit Vertically Implicit (HEVI) time-integration. Two versions of the DyCore are presented and tested, each based on a different formulation of the specific internal energy and continuity equations, which, unlike standard potential temperature formulations, are valid for variable composition atmospheres. The first version, which uses a product-rule (PR) forms of the continuity and specific internal energy equation, contains an additional pressure dilation term and does not conserve mass. The second version, which does not use the product-rule (no-PR) in the continuity and specific internal energy, contains two terms to represent pressure dilation and conserves mass to machine precision regardless of time truncation error. The pressure gradient and gravitational forces in the momentum balance equation are reformulated to reduce numerical errors at high altitudes. These new equation sets were implemented in two SEM-based atmospheric models: the Nonhydrostatic Unified Model of the Atmosphere (NUMA) and the Navy Environmental Prediction sysTem Using a Nonhydrostatic Engine (NEPTUNE). Numerical results using both a deep-atmosphere and shallow-atmosphere baroclinic instability, a balanced zonal flow, and a high-altitude orographic gravity wave verify the fidelity of the dynamics at low and high altitudes and for constant and variable composition atmospheres. These results are compared to existing deep-atmosphere dynamical cores and a Fourier-ray code, indicating that the proposed discretized equation sets are viable DyCore candidates for next-generation ground-to-thermosphere atmospheric models.

physics.ao-ph↗

Accelerating Simulations of Tropical Cyclones using Adaptive Mesh Refinement

Tropical cyclones (TCs) are powerful, natural phenomena that can severely impact populations and infrastructure. Enhancing our understanding of the mechanisms driving their intensification is crucial for mitigating these impacts. To this end, researchers are pushing the boundaries of TC simulation resolution down to scales of just a few meters. However, higher resolution simulations come with significant computational challenges, increasing both time and energy costs. Adaptive mesh refinement (AMR) is a technique widely used in computational fluid dynamics but has seen limited application in atmospheric simulations. This study explores the use of h-adaptive grids using the spectral element discretization technique to accelerate TC simulations while allowing very high resolutions in certain parts of the domain. By applying AMR to a rapidly intensifying TC test case, we demonstrate that AMR can replicate the results of uniform grid simulations in terms of mean and local wind speed maxima while dramatically reducing computational costs. We show that AMR can speed up dry TC simulations by a factor of 2 - 13 for the set of tested refinement criteria. Additionally, we show that TC intensity changes as resolution is increased and that AMR can deliver high-resolution simulations at the cost of coarser static simulations. Our findings indicate that AMR and spectral element methods are promising tools for enhancing TC simulations.

physics.ao-ph↗

Multiscale Modeling Framework using Element-based Galerkin Methods for Moist Atmospheric Limited-Area Simulations

This paper presents a multiscale modeling framework (MMF) to model moist atmospheric limited-area weather. The MMF resolves large-scale convection using a coarse grid while simultaneously resolving local features through numerous fine local grids and coupling them seamlessly. Both large- and small-scale processes are modeled using the compressible Navier-Stokes equations within the Nonhydrostatic Unified Model of the Atmosphere (NUMA), and they are discretized using a continuous element-based Galerkin method (spectral elements) with high-order basis functions. Consequently, the large-scale and small-scale models share the same dynamical core but have the flexibility to be adjusted individually. The proposed MMF method is tested in 2D and 3D idealized limited-area weather problems involving storm clouds produced by squall line and supercell simulations. The MMF numerical results showed enhanced representation of cloud processes compared to the coarse model.

math.NA↗

Comparison of adaptive mesh refinement techniques for numerical weather prediction

This paper examines the application of adaptive mesh refinement (AMR) in the field of numerical weather prediction (NWP). We implement and assess two distinct AMR approaches and evaluate their performance through standard NWP benchmarks. In both cases, we solve the fully compressible Euler equations, fundamental to many non-hydrostatic weather models. The first approach utilizes oct-tree cell-based mesh refinement coupled with a high-order discontinuous Galerkin method for spatial discretization. In the second approach, we employ level-based AMR with the finite difference method. Our study provides insights into the accuracy and benefits of employing these AMR methodologies for the multi-scale problem of NWP. Additionally, we explore essential properties including their impact on mass and energy conservation. Moreover, we present and evaluate an AMR solution transfer strategy for the tree-based AMR approach that is simple to implement, memory-efficient, and ensures conservation for both flow in the box and sphere. Furthermore, we discuss scalability, performance portability, and the practical utility of the AMR methodology within an NWP framework -- crucial considerations in selecting an AMR approach. The current de facto standard for mesh refinement in NWP employs a relatively simplistic approach of static nested grids, either within a general circulation model or a separately operated regional model with loose one-way synchronization. It is our hope that this study will stimulate further interest in the adoption of AMR frameworks like AMReX in NWP. These frameworks offer a triple advantage: a robust dynamic AMR for tracking localized and consequential features such as tropical cyclones, extreme scalability, and performance portability.

math.NA↗

A Performance Study of Horizontally Explicit Vertically Implicit (HEVI) Time-Integrators for Non-Hydrostatic Atmospheric Models

We conduct a thorough study of different forms of horizontally explicit and vertically implicit (HEVI) time-integration strategies for the compressible Euler equations on spherical domains typical of nonhydrostatic global atmospheric applications. We compare the computational time and complexity of two nonlinear variants (NHEVI-GMRES and NHEVI-LU) and a linear variant (LHEVI). We report on the performance of these three variants for a number of additive Runge-Kutta Methods ranging in order of accuracy from second through fifth, and confirm the expected order of accuracy of the HEVI methods for each time-integrator. To gauge the maximum usable time-step of each HEVI method, we run simulations of a nonhydrostatic baroclinic instability for 100 days and then use this time-step to compare the time-to-solution of each method. The results show that NHEVI-LU is 2x faster than NHEVI-GMRES, and LHEVI is 5x faster than NHEVI-LU, for the idealized cases tested. The baroclinic instability and inertia-gravity wave simulations indicate that the optimal choice of time-integrator is LHEVI with either second or third order schemes, as both schemes yield similar time to solution and relative L2 error at their maximum usable time-steps. In the future, we will report on whether these results hold for more complex problems using, e.g., real atmospheric data and/or a higher model top typical of space weather applications.

math.NA↗

IMplicit-EXplicit Formulations for Discontinuous Galerkin Non-Hydrostatic Atmospheric Models

This work presents IMplicit-EXplicit (IMEX) formulations for discontinuous Galerkin (DG) discretizations of the compressible Euler equations governing non-hydrostatic atmospheric flows. In particular, we show two different IMEX formulations that not only treat the stiffness due to the governing dynamics but also the domain discretization. We present these formulations for two different equation sets typically employed in atmospheric modeling. For both equation sets, efficient Schur complements are derived and the challenges and remedies for deriving them are discussed. The performance of these IMEX formulations of different orders are investigated on both 2D (box) and 3D (sphere) test problems and shown to achieve their theoretical rates of convergence and their efficiency with respect to both mesoscale and global applications are presented.

math.NA↗

Entropy Stable Discontinuous Galerkin Methods for Balance Laws in Non-Conservative Form: Applications to the Euler Equations with Gravity

In this work a non-conservative balance law formulation is considered that encompasses the rotating, compressible Euler equations for dry atmospheric flows. We develop a semi-discretely entropy stable discontinuous Galerkin method on curvilinear meshes using a generalization of flux differencing for numerical fluxes in fluctuation form. The method uses the skew-hybridized formulation of the element operators to ensure that, even in the presence of under-integration on curvilinear meshes, the resulting discretization is entropy stable. Several atmospheric flow test cases in one, two, and three dimensions confirm the theoretical entropy stability results as well as show the high-order accuracy and robustness of the method.

math.NA↗

Comparison of Sub-Grid Scale Models for Large-Eddy Simulation using a High-Order Spectral Element Approximation of the Compressible Navier-Stokes Equations at Low Mach Number

This study aims to identify the properties, advantages, and drawbacks of some common (and some less common) sub-grid scale (SGS) models for large eddy simulation of low Mach compressible flows using high order spectral elements. The models investigated are the classical constant coefficient Smagorinsky-Lilly, the model by Vreman and two variants of a dynamic SGS (DSGS) model designed to stabilize finite and spectral elements for transport dominated problems. In particular, we compare one variant of DSGS that is based on a time-dependent residual version (R-DSGS) in contrast to a time-independent residual based scheme (T-DSGS). The SGS models are compared against the reference model by Smagorinsky and Lilly for their ability to: (i) stabilize the numerical solution, (ii) minimize undershoots and overshoots, (iii) capture/preserve discontinuities, and (iv) transfer energy across different length scales. These abilities are investigated on problems for: (1) passively advected tracers, (2) coupled, nonlinear system of equations exhibiting discontinuities, (3) gravity-driven flows in a stratified atmosphere, and (4) homogenous, isotropic turbulence. All models were able to preserve sharp discontinuities. Vreman and the R-DSGS models also reduce the undershoots and overshoots in the solution of linear and non-linear advection with sharp gradients. Our analysis shows that the R-DSGS and T-DSGS models are more robust than Vreman and Smagorinsky-Lilly for numerical stabilization of high-order spectral methods. The Smagorinsky and Vreman models are better able to resolve the finer flow structures in shear flows, while the nodal R-DSGS model shows better energy conservation. Overall, the nodal implementation of R-DSGS (in contrast to its element-based counterpart) is shown to outperform the other SGS models in most metrics listed above, and on par with respect to the remaining ones.

physics.flu-dyn↗

Large-eddy simulations with ClimateMachine: a new open-source code for atmospheric simulations on GPUs and CPUs

We introduce ClimateMachine, a new open-source atmosphere modeling framework using the Julia language to be performance portable on central processing units (CPUs) and graphics processing units (GPUs). ClimateMachine uses a common framework both for coarser-resolution global simulations and for high-resolution, limited-area large-eddy simulations (LES). Here, we demonstrate the LES configuration of the atmosphere model in canonical benchmark cases and atmospheric flows, using an energy-conserving nodal discontinuous-Galerkin (DG) discretization of the governing equations. Resolution dependence, conservation characteristics and scaling metrics are examined in comparison with existing LES codes. They demonstrate the utility of ClimateMachine as a modelling tool for limited-area LES flow configurations.

physics.ao-ph↗

IMEX HDG-DG: a coupled implicit hybridized discontinuous Galerkin (HDG) and explicit discontinuous Galerkin (DG) approach for shallow water systems

We propose IMEX HDG-DG schemes for planar and spherical shallow water systems. Of interest is subcritical flow, where the speed of the gravity wave is faster than that of nonlinear advection. In order to simulate these flows efficiently, we split the governing system into a stiff part describing the gravity wave and a non-stiff part associated with nonlinear advection. The former is discretized implicitly with the HDG method while an explicit Runge-Kutta DG discretization is employed for the latter. The proposed IMEX HDG-DG framework: 1) facilitates high-order solutions both in time and space; 2) avoids overly small time-step sizes; 3) requires only one linear system solve per time stage; 4) relative to DG generates smaller and sparser linear systems while promoting further parallelism. Numerical results of various test cases demonstrate that our methods are comparable to explicit Runge-Kutta DG schemes in terms of accuracy while allowing for much larger time step sizes.

cs.CE↗

Acceleration of the Implicit-Explicit Non-hydrostatic Unified Model of the Atmosphere (NUMA) on Manycore Processors

We present the acceleration of an IMplicit-EXplicit (IMEX) non-hydrostatic atmospheric model on manycore processors such as GPUs and Intel's MIC architecture. IMEX time integration methods sidestep the constraint imposed by the Courant-Friedrichs-Lewy condition on explicit methods through corrective implicit solves within each time step. In this work, we implement and evaluate the performance of IMEX on manycore processors relative to explicit methods. Using 3D-IMEX at Courant number C=15 , we obtained a speedup of about 4X relative to an explicit time stepping method run with the maximum allowable C=1. In addition, we demonstrate a much larger speedup of 100X at C=150 using 1D-IMEX due to the unconditional stability of the method in the vertical direction. Several improvements on the IMEX procedure were necessary in order to outperform our results with explicit methods: a) reducing the number of degrees of freedom of the IMEX formulation by forming the Schur complement; b) formulating a horizontally-explicit vertically-implicit (HEVI) 1D-IMEX scheme that has a lower workload and potentially better scalability than 3D-IMEX; c) using high-order polynomial preconditioners to reduce the condition number of the resulting system; d) using a direct solver for the 1D-IMEX method by performing and storing LU factorizations once to obtain a constant cost for any Courant number. Without all of these improvements, explicit time integration methods turned out to be difficult to beat. We discuss in detail the IMEX infrastructure required for formulating and implementing efficient methods on manycore processors. Finally, we validate our results with standard benchmark problems in NWP and evaluate the performance and scalability of the IMEX method using up to 4192 GPUs and 16 Knights Landing processors.

math.NA↗

Strong Scaling for Numerical Weather Prediction at Petascale with the Atmospheric Model NUMA

Numerical weather prediction (NWP) has proven to be computationally challenging due to its inherent multiscale nature. Currently, the highest resolution NWP models use a horizontal resolution of about 10km. In order to increase the resolution of NWP models highly scalable atmospheric models are needed. The Non-hydrostatic Unified Model of the Atmosphere (NUMA), developed by the authors at the Naval Postgraduate School, was designed to achieve this purpose. NUMA is used by the Naval Research Laboratory, Monterey as the engine inside its next generation weather prediction system NEPTUNE. NUMA solves the fully compressible Navier-Stokes equations by means of high-order Galerkin methods (both spectral element as well as discontinuous Galerkin methods can be used). Mesh generation is done using the p4est library. NUMA is capable of running middle and upper atmosphere simulations since it does not make use of the shallow-atmosphere approximation. This paper presents the performance analysis and optimization of the spectral element version of NUMA. The performance at different optimization stages is analyzed using a theoretical performance model as well as measurements via hardware counters. Machine independent optimization is compared to machine specific optimization using BG/Q vector intrinsics. By using vector intrinsics the main computations reach 1.2 PFlops on the entire machine Mira (12% of the theoretical peak performance). The paper also presents scalability studies for two idealized test cases that are relevant for NWP applications. The atmospheric model NUMA delivers an excellent strong scaling efficiency of 99% on the entire supercomputer Mira using a mesh with 1.8 billion grid points. This allows to run a global forecast of a baroclinic wave test case at 3km uniform horizontal resolution and double precision within the time frame required for operational weather prediction.

cs.DC↗

A continuous/discontinuous Galerkin solution of the shallow water equations with dynamic viscosity, high-order wetting and drying, and implicit time integration

The high-order numerical solution of the non-linear shallow water equations (and of hyperbolic systems in general) is susceptible to unphysical Gibbs oscillations that form in the proximity of strong gradients. The solution to this problem is still an active field of research as no general cure has been found yet. In this paper, we tackle this issue by presenting a dynamically adaptive viscosity based on a residual-based sub-grid scale model that has the following properties: $(i)$ it removes the spurious oscillations in the proximity of strong wave fronts while preserving the overall accuracy and sharpness of the solution. This is possible because of the residual-based definition of the dynamic diffusion coefficient. $(ii)$ For coarse grids, it prevents energy from building up at small wave-numbers. $(iii)$ The model has no tunable parameter. Our interest in the shallow water equations is tied to the simulation of coastal inundation, where a careful handling of the transition between dry and wet surfaces is particularly challenging for high-order Galerkin approximations. In this paper, we extend to a unified continuous/discontinuous Galerkin (CG/DG) framework a very simple, yet effective wetting and drying algorithm originally designed for DG [Xing, Zhang, Shu (2010)]. We show its effectiveness for problems in one and two dimensions on domains of increasing characteristic lengths varying from centimeters to kilometers. Finally, to overcome the time-step restriction incurred by the high-order Galerkin approximation, we advance the equations forward in time via a three stage, second order explicit-first-stage, singly diagonally implicit Runge-Kutta (ESDIRK) time integration scheme. Via ESDIRK, we are able to preserve numerical stability for an advective CFL number 10 times larger than its explicit counterpart.

math.NA↗

Partial Observability and its Consistency for PDEs

In this paper, a quantitative measure of partial observability is defined for PDEs. The quantity is proved to be consistent if the PDE is approximated using well-posed approximation schemes. A first order approximation of an unobservability index using an empirical Gramian is introduced. Several examples are presented to illustrate the concept of partial observability, including Burgers' equation and a one-dimensional nonlinear shallow water equation.

math.OC↗