Searcharxiv⌕ Search

arXiv subjects

Miroslav Stoyanov

Publications and source records attributed to Miroslav Stoyanov.

9 recordsLinked to original sources

Accelerating phase-field simulations on exascale computing systems for faster-than-real-time precipitate aging predictions

Three-dimensional phase-field simulations are a gold-standard for microstructure prediction for materials, but their computational cost often limits application-relevant calculations to modest domain sizes and timescales. Here, we present a holistic approach for accelerating Fourier pseudospectral phase-field simulations by combining performance-portable GPU computing, large-scale distributed-memory parallelism, and high-order implicit-explicit time integration within the MEUMAPPS C++ framework. We demonstrate this approach with a 1.5-billion-grid-point simulation of the growth and coarsening of 1,920 gamma'' precipitates in a Ni-Nb-Fe alloy. The simulated seven-hour heat treatment is completed in 5.5 hours, making the calculation faster than real time and an estimated 217-501x faster than a CPU-only, first-order baseline. This performance makes three-dimensional simulations of thousands of interacting precipitates tractable, enabling quantitative studies ocollective microstructural phenomena, large simulation ensembles, and real-time integration into controlling and interpreting experiments. Benchmarking shows single-node GPU speedups of up to 17.5x relative to comparable CPU resources, near-ideal strong scaling to 4096 GPUs for a 12-billion-grid-point problem, and a further 2.6-2.9x acceleration from fourth-order time integration at scientifically relevant error tolerances. These results establish a performance-portable strategy for exploiting leadership-scale GPU systems that is applicable to a broad class of Fourier pseudospectral simulations beyond phase-field models such as fluid dynamics and crystal plasticity simulations.

cond-mat.mtrl-sci↗

Adaptive Sparse-grid Discontinuous Galerkin Approximations the Bhatnagar--Gross--Krook Model

This work studies adaptive sparse-grid discontinuous Galerkin (DG) discretizations for the Bhatnagar--Gross--Krook (BGK) model, a kinetic equation posed in four- and six-dimensional phase-space. Standard DG methods are rendered impractical for the BGK model by the curse of dimensionality, motivating compressed representations that adapt to the solution in time. Using the adaptive sparse-grid DG method, we quantify accuracy and compression by comparing the adaptive degrees of freedom to full-grid DG methods and by assessing the resulting kinetic and fluid quantities in both fluid and rarefied regimes. Test cases include a relaxation problem, a multidimensional Sod shock tube, and shear/expansion flows used in prior low-rank BGK studies. To build an efficient Maxwellian evaluation without violating conservation, a central obstacle for structure-perserving BGK simulations, we introduce a hybrid interpolation strategy that exploits velocity separability to recover the correct discrete collision invariants and prove conservation of the resulting discrete collision operator on adaptive sparse grids. Our results show that the adaptive sparse-grid strategy can recover accurate and physically relevant solutions with sharp gradients, and the method reduces the active degrees of freedom by factors ranging from several-fold to several orders of magnitude, with the largest reductions occurring in the six-dimensional examples. All computations are performed with the open-source ASGarD adaptive sparse-grid DG library.

math.NA↗

Sparse-grid Discontinuous Galerkin Methods for the Vlasov-Poisson-Lenard-Bernstein Model

Sparse-grid methods have recently gained interest in reducing the computational cost of solving high-dimensional kinetic equations. In this paper, we construct adaptive and hybrid sparse-grid methods for the Vlasov-Poisson-Lenard-Bernstein (VPLB) model. This model has applications to plasma physics and is simulated in two reduced geometries: a 0x3v space homogeneous geometry and a 1x3v slab geometry. We use the discontinuous Galerkin (DG) method as a base discretization due to its high-order accuracy and ability to preserve important structural properties of partial differential equations. We utilize a multiwavelet basis expansion to determine the sparse-grid basis and the adaptive mesh criteria. We analyze the proposed sparse-grid methods on a suite of three test problems by computing the savings afforded by sparse-grids in comparison to standard solutions of the DG method. The results are obtained using the adaptive sparse-grid discretization library ASGarD.

math.NA↗

Scaling and performance portability of the particle-in-cell scheme for plasma physics applications through mini-apps targeting exascale architectures

We perform a scaling and performance portability study of the particle-in-cell scheme for plasma physics applications through a set of mini-apps we name "Alpine", which can make use of exascale computing capabilities. The mini-apps are based on Independent Parallel Particle Layer, a framework that is designed around performance portable and dimension independent particles and fields. We benchmark the simulations with varying parameters such as grid resolutions ($512^3$ to $2048^3$) and number of simulation particles ($10^9$ to $10^{11}$) with the following mini-apps: weak and strong Landau damping, bump-on-tail and two-stream instabilities, and the dynamics of an electron bunch in a charge-neutral Penning trap. We show strong and weak scaling and analyze the performance of different components on several pre-exascale architectures such as Piz-Daint, Cori, Summit and Perlmutter. While the scaling and portability study helps identify the performance critical components of the particle-in-cell scheme in the current state-of-the-art computing architectures, the mini-apps by themselves can be used to develop new algorithms and optimize their high performance implementations targeting exascale architectures.

physics.comp-ph↗

Stable Anderson Acceleration for Deep Learning

Anderson acceleration (AA) is an extrapolation technique designed to speed-up fixed-point iterations like those arising from the iterative training of DL models. Training DL models requires large datasets processed in randomly sampled batches that tend to introduce in the fixed-point iteration stochastic oscillations of amplitude roughly inversely proportional to the size of the batch. These oscillations reduce and occasionally eliminate the positive effect of AA. To restore AA's advantage, we combine it with an adaptive moving average procedure that smoothes the oscillations and results in a more regular sequence of gradient descent updates. By monitoring the relative standard deviation between consecutive iterations, we also introduce a criterion to automatically assess whether the moving average is needed. We applied the method to the following DL instantiations: (i) multi-layer perceptrons (MLPs) trained on the open-source graduate admissions dataset for regression, (ii) physics informed neural networks (PINNs) trained on source data to solve 2d and 100d Burgers' partial differential equations (PDEs), and (iii) ResNet50 trained on the open-source ImageNet1k dataset for image classification. Numerical results obtained using up to 1,536 NVIDIA V100 GPUs on the OLCF supercomputer Summit showed the stabilizing effect of the moving average on AA for all the problems above.

cs.LG↗

A Feynman-Kac based numerical method for the exit time probability of a class of transport problems

The exit time probability, which gives the likelihood that an initial condition leaves a prescribed region of the phase space of a dynamical system at, or before, a given time, is arguably one of the most natural and important transport problems. Here we present an accurate and efficient numerical method for computing this probability for systems described by non-autonomous (time-dependent) stochastic differential equations (SDEs) or their equivalent Fokker-Planck partial differential equations. The method is based on the direct approximation of the Feynman-Kac formula that establishes a link between the adjoint Fokker-Planck equation and the forward SDE. The Feynman-Kac formula is approximated using the Gauss-Hermite quadrature rules and piecewise cubic Hermite interpolating polynomials, and a GPU accelerated matrix representation is used to compute the entire time evolution of the exit time probability using a single pass of the algorithm. The method is unconditionally stable, exhibits second-order convergence in space, first-order convergence in time, and is straightforward to parallelize. Applications are presented to the advection-diffusion of a passive tracer in a fluid flow exhibiting chaotic advection, and to the runaway acceleration of electrons in a plasma in the presence of an electric field, collisions, and radiation damping. Benchmarks against analytical solutions as well as comparisons with explicit and implicit finite difference standard methods for the adjoint Fokker-Planck equation are presented.

physics.comp-ph↗

A Method for Dimensionally Adaptive Sparse Trigonometric Interpolation of Periodic Functions

We present a method for dimensionally adaptive sparse trigonometric interpolation of multidimensional periodic functions belonging to a smoothness class of finite order. This method targets applications where periodicity must be preserved and the precise anisotropy is not known a priori. To the authors' knowledge, this is the first instance of a dimensionally adaptive sparse interpolation algorithm that uses a trigonometric interpolation basis. The motivating application behind this work is the adaptive approximation of a multi-input model for a molecular potential energy surface (PES) where each input represents an angle of rotation. Our method is based on an anisotropic quasi-optimal estimate for the decay rate of the Fourier coefficients of the model; a least-squares fit to the coefficients of the interpolant is used to estimate the anisotropy. Thus, our adaptive approximation strategy begins with a coarse isotropic interpolant, which is gradually refined using the estimated anisotropic rates. The procedure takes several iterations where ever-more accurate interpolants are used to generate ever-improving anisotropy rates. We present several numerical examples of our algorithm where the adaptive procedure successfully recovers the theoretical "best" convergence rate, including an application to a periodic PES approximation. An open-source implementation of our algorithm resides in the Tasmanian UQ library developed at Oak Ridge National Laboratory.

math.NA↗

A sparse-grid probabilistic scheme for approximation of the runaway probability of electrons in fusion tokamak simulation

Runaway electrons (RE) generated during magnetic disruptions present a major threat to the safe operation of plasma nuclear fusion reactors. A critical aspect of understanding RE dynamics is to calculate the runaway probability, i.e., the probability that an electron in the phase space will runaway on, or before, a prescribed time. Such probability can be obtained by solving the adjoint equation of the underlying Fokker-Planck equation that controls the electron dynamics. In this effort, we present a sparse-grid probabilistic scheme for computing the runaway probability. The key ingredient of our approach is to represent the solution of the adjoint equation as a conditional expectation, such that discretizing the differential operator reduces to the approximation of a set of integrals. Adaptive sparse grid interpolation is utilized to approximate the map from the phase space to the runaway probability. The main novelties of this effort are the integration of the sparse-grid method into the probabilistic numerical scheme for computing escape probability, as well as the demonstration in computing RE probabilities. Two numerical examples are given to illustrate that the proposed method can achieve $\mathcal{O}(Δt)$ convergence, as well as the adaptive refinement strategy can effectively handle the sharp transition layer between the runaway and non-runaway regions.

physics.comp-ph↗

Power-Law Noises over General Spatial Domains and on Non-Standard Meshes

Power-law noises abound in nature and have been observed extensively in both time series and spatially varying environmental parameters. Although, recent years have seen the extension of traditional stochastic partial differential equations to include systems driven by fractional Brownian motion, spatially distributed scale-invariance has received comparatively little attention, especially for parameters defined over non-standard spatial domains. This paper discusses the generalization of power-law noises to general spatial domains by outlining their theoretical underpinnings as well as addressing their numerical simulation on arbitrary meshes. Three computational algorithms are presented for efficiently generating their sample paths, accompanied by numerous numerical illustrations.

stat.CO↗