SearcharxivSearch

arXiv subjects

Ann Almgren

Publications and source records attributed to Ann Almgren.

At least 19 recordsLinked to original sources

ERF: Energy Research and Forecasting Model

High performance computing (HPC) architectures have undergone rapid development in recent years. As a result, established software suites face an ever increasing challenge to remain performant on and portable across modern systems. Many of the widely adopted atmospheric modeling codes cannot fully (or in some cases, at all) leverage the acceleration provided by General-Purpose Graphics Processing Units (GPGPUs), leaving users of those codes constrained to increasingly limited HPC resources. Energy Research and Forecasting (ERF) is a regional atmospheric modeling code that leverages the latest HPC architectures, whether composed of only Central Processing Units (CPUs) or incorporating GPUs. ERF contains many of the standard discretizations and basic features needed to model general atmospheric dynamics as well as flows relevant to renewable energy. The modular design of ERF provides a flexible platform for exploring different physics parameterizations and numerical strategies. ERF is built on a state-of-the-art, well-supported, software framework (AMReX) that provides a performance portable interface and ensures ERF's long-term sustainability on next generation computing systems. This paper details the numerical methodology of ERF and presents results for a series of verification and validation cases.

physics.ao-ph

AMReX and pyAMReX: Looking Beyond ECP

AMReX is a software framework for the development of block-structured mesh applications with adaptive mesh refinement (AMR). AMReX was initially developed and supported by the AMReX Co-Design Center as part of the U.S. DOE Exascale Computing Project, and is continuing to grow post-ECP. In addition to adding new functionality and performance improvements to the core AMReX framework, we have also developed a Python binding, pyAMReX, that provides a bridge between AMReX-based application codes and the data science ecosystem. pyAMReX provides zero-copy application GPU data access for AI/ML, in situ analysis and application coupling, and enables rapid, massively parallel prototyping. In this paper we review the overall functionality of AMReX and pyAMReX, focusing on new developments, new functionality, and optimizations of key operations. We also summarize capabilities of ECP projects that used AMReX and provide an overview of new, non-ECP applications.

cs.DC

AMReX: Block-Structured Adaptive Mesh Refinement for Multiphysics Applications

Block-structured adaptive mesh refinement (AMR) provides the basis for the temporal and spatial discretization strategy for a number of ECP applications in the areas of accelerator design, additive manufacturing, astrophysics, combustion, cosmology, multiphase flow, and wind plant modelling. AMReX is a software framework that provides a unified infrastructure with the functionality needed for these and other AMR applications to be able to effectively and efficiently utilize machines from laptops to exascale architectures. AMR reduces the computational cost and memory footprint compared to a uniform mesh while preserving accurate descriptions of different physical processes in complex multi-physics algorithms. AMReX supports algorithms that solve systems of partial differential equations (PDEs) in simple or complex geometries, and those that use particles and/or particle-mesh operations to represent component physical processes. In this paper, we will discuss the core elements of the AMReX framework such as data containers and iterators as well as several specialized operations to meet the needs of the application projects. In addition we will highlight the strategy that the AMReX team is pursuing to achieve highly performant code across a range of accelerator-based architectures for a variety of different applications.

cs.MS

A Hybrid Algorithm for Systems of Non-interacting Particles with an External Potential

Our focus is on simulating the dynamics of non-interacting particles including the effects of an external potential, which, under certain assumptions, can be formally described by the Dean-Kawasaki equation. The Dean-Kawasaki equation can be solved numerically using standard finite volume methods. However, the numerical approximation implicitly requires a sufficiently large number of particles to ensure the positivity of the solution and accurate approximation of the stochastic flux. To address this challenge, we extend hybrid algorithms for particle systems to scenarios where the density is low. The aim is to create a hybrid algorithm that switches from a finite volume discretization to a particle-based method when the particle density falls below a certain threshold. We develop criteria for determining this threshold by comparing higher-order statistics obtained from the finite volume method with particle simulations. We then demonstrate the use of the resulting criteria for dynamic adaptation in both two- and three-dimensional spatial settings in the absence of an external potential. Finally we consider the dynamics when an external potential is included.

math.NA

Comparison of adaptive mesh refinement techniques for numerical weather prediction

This paper examines the application of adaptive mesh refinement (AMR) in the field of numerical weather prediction (NWP). We implement and assess two distinct AMR approaches and evaluate their performance through standard NWP benchmarks. In both cases, we solve the fully compressible Euler equations, fundamental to many non-hydrostatic weather models. The first approach utilizes oct-tree cell-based mesh refinement coupled with a high-order discontinuous Galerkin method for spatial discretization. In the second approach, we employ level-based AMR with the finite difference method. Our study provides insights into the accuracy and benefits of employing these AMR methodologies for the multi-scale problem of NWP. Additionally, we explore essential properties including their impact on mass and energy conservation. Moreover, we present and evaluate an AMR solution transfer strategy for the tree-based AMR approach that is simple to implement, memory-efficient, and ensures conservation for both flow in the box and sphere. Furthermore, we discuss scalability, performance portability, and the practical utility of the AMR methodology within an NWP framework -- crucial considerations in selecting an AMR approach. The current de facto standard for mesh refinement in NWP employs a relatively simplistic approach of static nested grids, either within a general circulation model or a separately operated regional model with loose one-way synchronization. It is our hope that this study will stimulate further interest in the adoption of AMR frameworks like AMReX in NWP. These frameworks offer a triple advantage: a robust dynamic AMR for tracking localized and consequential features such as tropical cyclones, extreme scalability, and performance portability.

math.NA

A cast of thousands: How the IDEAS Productivity project has advanced software productivity and sustainability

Computational and data-enabled science and engineering are revolutionizing advances throughout science and society, at all scales of computing. For example, teams in the U.S. DOE Exascale Computing Project have been tackling new frontiers in modeling, simulation, and analysis by exploiting unprecedented exascale computing capabilities-building an advanced software ecosystem that supports next-generation applications and addresses disruptive changes in computer architectures. However, concerns are growing about the productivity of the developers of scientific software, its sustainability, and the trustworthiness of the results that it produces. Members of the IDEAS project serve as catalysts to address these challenges through fostering software communities, incubating and curating methodologies and resources, and disseminating knowledge to advance developer productivity and software sustainability. This paper discusses how these synergistic activities are advancing scientific discovery-mitigating technical risks by building a firmer foundation for reproducible, sustainable science at all scales of computing, from laptops to clusters to exascale and beyond.

cs.CY

From Compact Plasma Particle Sources to Advanced Accelerators with Modeling at Exascale

Developing complex, reliable advanced accelerators requires a coordinated, extensible, and comprehensive approach in modeling, from source to the end of beam lifetime. We present highlights in Exascale Computing to scale accelerator modeling software to the requirements set for contemporary science drivers. In particular, we present the first laser-plasma modeling on an exaflop supercomputer using the US DOE Exascale Computing Project WarpX. Leveraging developments for Exascale, the new DOE SCIDAC-5 Consortium for Advanced Modeling of Particle Accelerators (CAMPA) will advance numerical algorithms and accelerate community modeling codes in a cohesive manner: from beam source, over energy boost, transport, injection, storage, to application or interaction. Such start-to-end modeling will enable the exploration of hybrid accelerators, with conventional and advanced elements, as the next step for advanced accelerator modeling. Following open community standards, we seed an open ecosystem of codes that can be readily combined with each other and machine learning frameworks. These will cover ultrafast to ultraprecise modeling for future hybrid accelerator design, even enabling virtual test stands and twins of accelerators that can be used in operations.

physics.acc-ph

In Situ Data Summaries for Flexible Feature Analysis in Large-Scale Multiphase Flow Simulations

The study of multiphase flow is essential for understanding the complex interactions of various materials. In particular, when designing chemical reactors such as fluidized bed reactors (FBR), a detailed understanding of the hydrodynamics is critical for optimizing reactor performance and stability. An FBR allows experts to conduct different types of chemical reactions involving multiphase materials, especially interaction between gas and solids. During such complex chemical processes, formation of void regions in the reactor, generally termed as bubbles, is an important phenomenon. Study of these bubbles has a deep implication in predicting the reactor's overall efficiency. But physical experiments needed to understand bubble dynamics are costly and non-trivial. Therefore, to study such chemical processes and bubble dynamics, a state-of-the-art massively parallel computational fluid dynamics discrete element model (CFD-DEM), MFIX-Exa is being developed for simulating multiphase flows. Despite the proven accuracy of MFIX-Exa in modeling bubbling phenomena, the very-large size of the output data prohibits the use of traditional post hoc analysis capabilities in both storage and I/O time. To address these issues and allow the application scientists to explore the bubble dynamics in an efficient and timely manner, we have developed an end-to-end visual analytics pipeline that enables in situ detection of bubbles using statistical techniques, followed by a flexible and interactive visual exploration of bubble dynamics in the post hoc analysis phase. Positive feedback from the experts has indicated the efficacy of the proposed approach for exploring bubble dynamics in very-large scale multiphase flow simulations.

cs.HC

Massively parallel finite difference elasticity using block-structured adaptive mesh refinement with a geometric multigrid solver

Computationally solving the equations of elasticity is a key component in many materials science and mechanics simulations. Phenomena such as deformation-induced microstructure evolution, microfracture, and microvoid nucleation are examples of applications for which accurate stress and strain fields are required. A characteristic feature of these simulations is that the problem domain is simple (typically a rectilinear representative volume element (RVE)), but the evolution of internal topological features is extremely complex. Traditionally, the finite element method (FEM) is used for elasticity calculations; FEM is nearly ubiquituous due to (1) its ability to handle meshes of complex geometry using isoparametric elements, and (2) the weak formulation which eschews the need for computation of second derivatives. However, variable topology problems (e.g. microstructure evolution) require either remeshing, or adaptive mesh refinement (AMR) - both of which can cause extensive overhead and limited scaling. Block-structured AMR (BSAMR) is a method for adaptive mesh refinement that exhibits good scaling and is well-suited for many problems in materials science. Here, it is shown that the equations of elasticity can be efficiently solved using BSAMR using the finite difference method. The boundary operator method is used to treat different types of boundary conditions, and the "reflux-free" method is introduced to efficiently and easily treat the coarse-fine boundaries that arise in BSAMR. Examples are presented that demonstrate the use of this method in a variety of cases relevant to materials science: Eshelby inclusions, fracture, and microstructure evolution. Reasonable scaling is demonstrated up to $\sim$4000 processors with tens of millions of grid points, and good AMR efficiency is observed.

cs.CE

Preparing Nuclear Astrophysics for Exascale

Astrophysical explosions such as supernovae are fascinating events that require sophisticated algorithms and substantial computational power to model. Castro and MAESTROeX are nuclear astrophysics codes that simulate thermonuclear fusion in the context of supernovae and X-ray bursts. Examining these nuclear burning processes using high resolution simulations is critical for understanding how these astrophysical explosions occur. In this paper we describe the changes that have been made to these codes to transform them from standard MPI + OpenMP codes targeted at petascale CPU-based systems into a form compatible with the pre-exascale systems now online and the exascale systems coming soon. We then discuss what new science is possible to run on systems such as Summit and Perlmutter that could not have been achieved on the previous generation of supercomputers.

astro-ph.IM

An embedded boundary approach for efficient simulations of viscoplastic fluids in three dimensions

We present a methodology for simulating three-dimensional flow of incompressible viscoplastic fluids modelled by generalised Newtonian rheological equations. It is implemented in a highly efficient framework for massively parallelisable computations on block-structured grids. In this context, geometric features are handled by the embedded boundary approach, which requires specialised treatment only in cells intersecting or adjacent to the boundary. This constitutes the first published implementation of an embedded boundary algorithm for simulating flow of viscoplastic fluids on structured grids. The underlying algorithm employs a two-stage Runge-Kutta method for temporal discretisation, in which viscous terms are treated semi-implicitly and projection methods are utilised to enforce the incompressibility constraint. We augment the embedded boundary algorithm to deal with the variable apparent viscosity of the fluids. Since the viscosity depends strongly on the strain rate tensor, special care has been taken to approximate the components of the velocity gradients robustly near boundary cells, both for viscous wall fluxes in cut cells and for updates of apparent viscosity in cells adjacent to them. After performing convergence analysis and validating the code against standard test cases, we present the first ever fully three-dimensional simulations of creeping flow of Bingham plastics around translating objects. Our results shed new light on the flow fields around these objects.

physics.flu-dyn

Highly parallelisable simulations of time-dependent viscoplastic fluid flow simulations with structured adaptive mesh refinement

We present the extension of an efficient and highly parallelisable framework for incompressible fluid flow simulations to viscoplastic fluids. The system is governed by incompressible conservation of mass, the Cauchy momentum equation and a generalised Newtonian constitutive law. In order to simulate a wide range of viscoplastic fluids, we employ the Herschel-Bulkley model for yield-stress fluids with nonlinear stress-strain dependency above the yield limit. We utilise Papanastasiou regularisation in our algorithm to deal with the singularity in apparent viscosity. The resulting system of partial differential equations is solved using the IAMR code (Incompressible Adaptive Mesh Refinement), which uses second-order Godunov methodology for the advective terms and semi-implicit diffusion in the context of an approximate projection method to solve on adaptively refined meshes. By augmenting the IAMR code with the ability to simulate regularised Herschel-Bulkley fluids, we obtain efficient numerical software for time-dependent viscoplastic flow in three dimensions, which can be used to investigate systems not considered previously due to computational expense. We validate results from simulations using this new capability against previously published data for Bingham plastics and power-law fluids in the two-dimensional lid-driven cavity. In doing so, we expand the range of Bingham and Reynolds numbers which have been considered in the benchmark tests. Moreover, extensions to time-dependent flow of Herschel-Bulkley fluids and three spatial dimensions offer new insights into the flow of viscoplastic fluids in this test case, and we provide missing benchmark results for these extensions.

physics.flu-dyn

A Hybrid Adaptive Low-Mach-Number/Compressible Method: Euler Equations

Flows in which the primary features of interest do not rely on high-frequency acoustic effects, but in which long-wavelength acoustics play a nontrivial role, present a computational challenge. Integrating the entire domain with low-Mach-number methods would remove all acoustic wave propagation, while integrating the entire domain with the fully compressible equations can in some cases be prohibitively expensive due to the CFL time step constraint. For example, simulation of thermoacoustic instabilities might require fine resolution of the fluid/chemistry interaction but not require fine resolution of acoustic effects, yet one does not want to neglect the long-wavelength wave propagation and its interaction with the larger domain. The present paper introduces a new multi-level hybrid algorithm to address these types of phenomena. In this new approach, the fully compressible Euler equations are solved on the entire domain, potentially with local refinement, while their low-Mach-number counterparts are solved on subregions of the domain with higher spatial resolution. The finest of the compressible levels communicates inhomogeneous divergence constraints to the coarsest of the low-Mach-number levels, allowing the low-Mach-number levels to retain the long-wavelength acoustics. The performance of the hybrid method is shown for a series of test cases, including results from a simulation of the aeroacoustic propagation generated from a Kelvin-Helmholtz instability in low-Mach-number mixing layers. It is demonstrated that compared to a purely compressible approach, the hybrid method allows time-steps two orders of magnitude larger at the finest level, leading to an overall reduction of the computational time by a factor of 8.

math.NA

A Survey of High Level Frameworks in Block-Structured Adaptive Mesh Refinement Packages

Over the last decade block-structured adaptive mesh refinement (SAMR) has found increasing use in large, publicly available codes and frameworks. SAMR frameworks have evolved along different paths. Some have stayed focused on specific domain areas, others have pursued a more general functionality, providing the building blocks for a larger variety of applications. In this survey paper we examine a representative set of SAMR packages and SAMR-based codes that have been in existence for half a decade or more, have a reasonably sized and active user base outside of their home institutions, and are publicly available. The set consists of a mix of SAMR packages and application codes that cover a broad range of scientific domains. We look at their high-level frameworks, and their approach to dealing with the advent of radical changes in hardware architecture. The codes included in this survey are BoxLib, Cactus, Chombo, Enzo, FLASH, and Uintah.

cs.DC

BoxLib with Tiling: An AMR Software Framework

In this paper we introduce a block-structured adaptive mesh refinement (AMR) software framework that incorporates tiling, a well-known loop transformation. Because the multiscale, multiphysics codes built in BoxLib are designed to solve complex systems at high resolution, performance on current and next generation architectures is essential. With the expectation of many more cores per node on next generation architectures, the ability to effectively utilize threads within a node is essential, and the current model for parallelization will not be sufficient. We describe a new version of BoxLib in which the tiling constructs are embedded so that BoxLib-based applications can easily realize expected performance gains without extra effort on the part of the application developer. We also discuss a path forward to enable future versions of BoxLib to take advantage of NUMA-aware optimizations using the TiDA portable library.

cs.MS

The Lyman-$α$ forest in optically-thin hydrodynamical simulations

We study the statistics of the Lyman-$α$ forest in a flat LCDM cosmology with the N-body + Eulerian hydrodynamics code Nyx. We produce a suite of simulations, covering the observationally relevant redshift range $2 \leq z \leq 4$. We find that a grid resolution of 20 kpc/h is required to produce one percent convergence of Lyman-$α$ flux statistics, up to k = 10 h/Mpc. In addition to establishing resolution requirements, we study the effects of missing modes in these simulations, and find that box sizes of L > 40 Mpc/h are needed to suppress numerical errors to a sub-percent level. Our optically-thin simulations with the ionizing background prescription of Haardt & Madau (2012) reproduce an IGM equation of state with $T_0 \approx 10^4 K$ and $γ\approx 1.55$ at z=2, with a mean transmitted flux close to the observed values. When using the ionizing background prescription of Faucher-Giguere et al. (2009), the mean flux is 10-15 per cent below observed values at z=2, and a factor of 2 too small at z = 4. We show the effects of the common practice of rescaling optical depths to the observed mean flux and how it affects convergence rates. We also investigate the common practice of `splicing' results from a number of different simulations to estimate the 1D flux power spectrum and show it is accurate at the 10 percent level. Finally, we find that collisional heating of the gas from dark matter particles is negligible in modern cosmological simulations.

astro-ph.CO

Pair-Instability Supernovae of Non-Zero Metallicity Stars

Observational evidence suggests that some very massive stars in the local Universe may die as pair-instability supernovae. We present 2D simulations of the pair-instability supernova of a non-zero metallicity star. We find that very little mixing occurs in this explosion because metals in the stellar envelope drive strong winds that strip the hydrogen envelope from the star prior to death. Consequently, a reverse shock cannot form and trigger fluid instabilities during the supernova. Only weak mixing driven by nuclear burning occurs in the earliest stages of the supernova, and it is too weak to affect the observational signatures of the explosion.

astro-ph.HE

A Low Mach Number Model for Moist Atmospheric Flows

We introduce a low Mach number model for moist atmospheric flows that accurately incorporates reversible moist processes in flows whose features of interest occur on advective rather than acoustic time scales. Total water is used as a prognostic variable, so that water vapor and liquid water are diagnostically recovered as needed from an exact Clausius--Clapeyron formula for moist thermodynamics. Low Mach number models can be computationally more efficient than a fully compressible model, but the low Mach number formulation introduces additional mathematical and computational complexity because of the divergence constraint imposed on the velocity field. Here, latent heat release is accounted for in the source term of the constraint by estimating the rate of phase change based on the time variation of saturated water vapor subject to the thermodynamic equilibrium constraint. We numerically assess the validity of the low Mach number approximation for moist atmospheric flows by contrasting the low Mach number solution to reference solutions computed with a fully compressible formulation for a variety of test problems.

physics.ao-ph