SearcharxivSearch

arXiv subjects

Delyan Z. Kalchev

Publications and source records attributed to Delyan Z. Kalchev.

8 recordsLinked to original sources

Beyond Exascale: Dataflow Domain Translation on a Cerebras Cluster

Simulation of physical systems is essential across scientific and engineering domains. Commonly used domain decomposition methods are unable to simultaneously deliver both high simulation rate and high utilization in network computing environments. In particular, Exascale systems deliver only a small fraction their peak performance for these workloads. This paper introduces the novel Domain Translation algorithm, designed to overcome these limitations. On a cluster of 64 Cerebras CS-3 systems, we use this method to demonstrate unprecedented cluster performance across a range of metrics: we show simulations running in excess of 1.6 million time steps per second; we also demonstrate perfect weak scaling at 88% of peak performance. At this cluster scale, our implementation provides 112 PFLOP/s in a power-unconstrained environment, and 57 GFLOP/J in a power-limited environment. We illustrate the method by applying the shallow-water equations to model a tsunami following an asteroid impact at 460m-resolution on a planetary scale.

cs.DC

Breaking the mold: overcoming the time constraints of molecular dynamics on general-purpose hardware

The evolution of molecular dynamics (MD) simulations has been intimately linked to that of computing hardware. For decades following the creation of MD, simulations have improved with computing power along the three principal dimensions of accuracy, atom count (spatial scale), and duration (temporal scale). Since the mid-2000s, computer platforms have however failed to provide strong scaling for MD as scale-out CPU and GPU platforms that provide substantial increases to spatial scale do not lead to proportional increases in temporal scale. Important scientific problems therefore remained inaccessible to direct simulation, prompting the development of increasingly sophisticated algorithms that present significant complexity, accuracy, and efficiency challenges. While bespoke MD-only hardware solutions have provided a path to longer timescales for specific physical systems, their impact on the broader community has been mitigated by their limited adaptability to new methods and potentials. In this work, we show that a novel computing architecture, the Cerebras Wafer Scale Engine, completely alters the scaling path by delivering unprecedentedly high simulation rates up to 1.144M steps/second for 200,000 atoms whose interactions are described by an Embedded Atom Method potential. This enables direct simulations of the evolution of materials using general-purpose programmable hardware over millisecond timescales, dramatically increasing the space of direct MD simulations that can be carried out.

cs.DC

Scalable Multilevel Monte Carlo Methods Exploiting Parallel Redistribution on Coarse Levels

We study an element agglomeration coarsening strategy that requires data redistribution at coarse levels when the number of coarse elements becomes smaller than the used computational units (cores). The overall procedure generates coarse elements (general unstructured unions of fine grid elements) within the framework of element-based algebraic multigrid methods (or AMGe) studied previously. The AMGe generated coarse spaces have the ability to exhibit approximation properties of the same order as the fine-level ones since by construction they contain the piecewise polynomials of the same order as the fine level ones. These approximation properties are key for the successful use of AMGe in multilevel solvers for nonlinear partial differential equations as well as for multilevel Monte Carlo (MLMC) simulations. The ability to coarsen without being constrained by the number of available cores, as described in the present paper, allows to improve the scalability of these solvers as well as in the overall MLMC method. The paper illustrates this latter fact with detailed scalability study of MLMC simulations applied to model Darcy equations with a stochastic log-normal permeability field.

math.NA

Parallel Element-based Algebraic Multigrid for H(curl) and H(div) Problems Using the ParELAG Library

This paper presents the use of element-based algebraic multigrid (AMGe) hierarchies, implemented in the ParELAG (Parallel Element Agglomeration Algebraic Multigrid Upscaling and Solvers) library, to produce multilevel preconditioners and solvers for H(curl) and H(div) formulations. ParELAG constructs hierarchies of compatible nested spaces, forming an exact de Rham sequence on each level. This allows the application of hybrid smoothers on all levels and AMS (Auxiliary-space Maxwell Solver) or ADS (Auxiliary-space Divergence Solver) on the coarsest levels, obtaining complete multigrid cycles. Numerical results are presented, showing the parallel performance of the proposed methods. As a part of the exposition, this paper demonstrates some of the capabilities of ParELAG and outlines some of the components and procedures within the library.

math.NA

A Condensed Constrained Nonconforming Mortar-based Approach for Preconditioning Finite Element Discretization Problems

This paper presents and studies an approach for constructing auxiliary space preconditioners for finite element problems using a constrained nonconforming reformulation, that is based on a proposed modified version of the mortar method. The well-known mortar finite element discretization method is modified to admit a local structure, providing an element-by-element or subdomain-by-subdomain assembly property. This is achieved via the introduction of additional trace finite element spaces and degrees of freedom (unknowns) associated with the interfaces between adjacent elements or subdomains. The resulting nonconforming formulation and a reduced via static condensation Schur complement form on the interfaces are used in the construction of auxiliary space preconditioners for a given conforming finite element discretization problem. The properties of these preconditioners are studied and their performance is illustrated on model second order scalar elliptic problems utilizing high order elements.

math.NA

Auxiliary Space Preconditioning of Finite Element Equations Using a Nonconforming Interior Penalty Reformulation and Static Condensation

We modify the well-known interior penalty finite element discretization method so that it allows for element-by-element assembly. This is possible due to the introduction of additional unknowns associated with the interfaces between neighboring elements. The resulting bilinear form, and a Schur complement (reduced) version of it, are utilized in a number of auxiliary space preconditioners for the original conforming finite element discretization problem. These preconditioners are analyzed on the fine scale and their performance is illustrated on model second order scalar elliptic problems discretized with high order elements.

math.NA

Mixed $(\mathcal{L}\mathcal{L}^*)^{-1}$ and $\mathcal{L}\mathcal{L}^*$ least-squares finite element methods with application to linear hyperbolic problems

In this paper, a few dual least-squares finite element methods and their application to scalar linear hyperbolic problems are studied. The purpose is to obtain $L^2$-norm approximations on finite element spaces of the exact solutions to hyperbolic partial differential equations of interest. This is approached by approximating the generally infeasible quadratic minimization, that defines the $L^2$-orthogonal projection of the exact solution, by feasible least-squares principles using the ideas of the original $\mathcal{L}\mathcal{L}^*$ method proposed in the context of elliptic equations. All methods in this paper are founded upon and extend the $\mathcal{L}\mathcal{L}^*$ approach which is rather general and applicable beyond the setting of elliptic problems. Error bounds are shown that point to the factors affecting the convergence and provide conditions that guarantee optimal rates. Furthermore, the preconditioning of the resulting linear systems is discussed. Numerical results are provided to illustrate the behavior of the methods on common finite element spaces.

math.NA

A Least-Squares Finite Element Method Based on the Helmholtz Decomposition for Hyperbolic Balance Laws

In this paper, a least-squares finite element method for scalar nonlinear hyperbolic balance laws is proposed and studied. The approach is based on a formulation that utilizes an appropriate Helmholtz decomposition of the flux vector and is related to the standard notion of a weak solution. This relationship, together with a corresponding connection to negative-norm least-squares, is described in detail. As a consequence, an important numerical conservation theorem is obtained, similar to the famous Lax-Wendroff theorem. The numerical conservation properties of the method in this paper do not fall precisely in the framework introduced by Lax and Wendroff, but they are similar in spirit as they guarantee that when $L^2$ convergence holds, the resulting approximations approach a weak solution to the hyperbolic problem. The least-squares functional is continuous and coercive in an $H^{-1}$-type norm, but not $L^2$-coercive. Nevertheless, the $L^2$ convergence properties of the method are discussed. Convergence can be obtained either by an explicit regularization of the functional, that provides control of the $L^2$ norm, or by properly choosing the finite element spaces, providing implicit control of the $L^2$ norm. Numerical results for the inviscid Burgers equation with discontinuous source terms are shown, demonstrating the $L^2$ convergence of the obtained approximations to the physically admissible solution. The numerical method utilizes a least-squares functional, minimized on finite element spaces, and a Gauss-Newton technique with nested iteration. We believe that the linear systems encountered with this formulation are amenable to multigrid techniques and combining the method with adaptive mesh refinement would make this approach an efficient tool for solving balance laws (this is the focus of a future study).

math.NA