SearcharxivSearch

arXiv subjects

Jack Deslippe

Publications and source records attributed to Jack Deslippe.

At least 19 recordsLinked to original sources

StochasticGW-GPU: rapid quasi-particle energies for molecules beyond 10000 atoms

$\mathtt{StochasticGW}$ is a code for computing accurate Quasi-Particle (QP) energies of molecules and material systems in the GW approximation. $\mathtt{StochasticGW}$ utilizes the stochastic Resolution of the Identity (sROI) technique to enable a massively-parallel implementation with computational costs that scale semi-linearly with system size, allowing the method to access systems with tens of thousands of electrons. We introduce a new implementation, $\mathtt{StochasticGW-GPU}$, for which the main bottleneck steps have been ported to GPUs and which gives substantial performance improvements over previous versions of the code. We showcase the new code by computing band gaps of hydrogenated silicon clusters ($\textrm{S}\textrm{i}_{\textrm{x}}\textrm{H}_{\textrm{y}}$) containing up to 10001 atoms and 35144 electrons, and we obtain individual QP energies with a statistical precision of better than $\pm0.03$ eV with times-to-solution on the order of minutes.

physics.chem-ph

GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN

The numerical solution of the Kadanoff-Baym nonlinear integro-differential equations, which yields the non-equilibrium Green's functions (NEGFs) of quantum many-body systems, poses significant computational challenges due to its high computational complexity. In this work, we present efficient implementations of a numerical method for solving these equations on distributed-memory architectures, including many-core CPUs and multi-GPU systems. For CPU-based platforms, we adopt a hybrid MPI/OpenMP programming model to exploit both inter-node and intra-node parallelism. On GPU-accelerated systems, we implement the method using two distinct approaches: MPI/OpenACC and MPI/CUDA FORTRAN. Several optimization strategies are employed to enhance GPU performance, including techniques to maximize computational resource utilization and minimize the overhead associated with kernel launches and memory management. Although OpenACC is easy to use, CUDA FORTRAN provides more advanced features for configuring and managing multiple levels of concurrency, while also simplifying memory allocation and data movement between host and device. This flexibility translates into significant performance improvements. We compare the performance of the three implementations and demonstrate that the GPU-based approaches achieve substantial speedups over CPU-based implementations. Furthermore, both CPU and GPU versions exhibit excellent strong and weak scaling, confirming the scalability and efficiency of our approach for large-scale NEGF computations.

cs.DC

Fully relativistic $GW$/Bethe-Salpeter calculations in BerkeleyGW: implementation, symmetries, benchmarking, and performance

Computing the $GW$ quasiparticle bandstructure and Bethe-Salpeter Equation (BSE) absorption spectra for materials with spin-orbit coupling has commonly been done by treating $GW$ corrections and spin-orbit coupling as separate perturbations to density-functional theory. However, accurate treatment of materials with strong spin-orbit coupling often requires a fully relativistic approach using spinor wavefunctions in the Kohn-Sham equation and $GW$/BSE. Such calculations have only recently become available, in particular for the BSE. We have implemented this approach in the plane-wave pseudopotential $GW$/BSE code BerkeleyGW, which is highly parallelized and widely used in the electronic-structure community. We present reference results for quasiparticle bandstructures and optical absorption spectra of solids with different strengths of spin-orbit coupling, including Si, Ge, GaAs, GaSb, CdSe, Au, and Bi$_2$Se$_3$. The calculated quasiparticle band gaps of these systems are found to agree with experiment to within a few tens of meV. The absorption spectrum of GaSb calculated with the fully-relativistic $GW$-BSE captures the large spin-orbit splitting of peaks in the spectrum. For Bi$_2$Se$_3$, we find a drastic change in the low-energy bandstructure compared to that of DFT, with the fully-relativistic treatment of the $GW$ approximation correctly capturing the parabolic nature of the valence and conduction bands after including off-diagonal self-energy matrix elements. We present the detailed methodology, approach to spatial symmetries for spinors, comparison against other codes, and performance compared to spinless $GW$/BSE calculations and perturbative approaches to SOC. This work aims to spur further development of spinor $GW$/BSE methodology in excited-state research software.

cond-mat.mtrl-sci

In-Situ Assessment of Device-Side Compute Work for Dynamic Load Balancing in a GPU-Accelerated PIC Code

Maintaining computational load balance is important to the performant behavior of codes which operate under a distributed computing model. This is especially true for GPU architectures, which can suffer from memory oversubscription if improperly load balanced. We present enhancements to traditional load balancing approaches and explicitly target GPU architectures, exploring the resulting performance. A key component of our enhancements is the introduction of several GPU-amenable strategies for assessing compute work. These strategies are implemented and benchmarked to find the most optimal data collection methodology for in-situ assessment of GPU compute work. For the fully kinetic particle-in-cell code WarpX, which supports MPI+CUDA parallelism, we investigate the performance of the improved dynamic load balancing via a strong scaling-based performance model and show that, for a laser-ion acceleration test problem run with up to 6144 GPUs on Summit, the enhanced dynamic load balancing achieves from 62%--74% (88% when running on 6 GPUs) of the theoretically predicted maximum speedup; for the 96-GPU case, we find that dynamic load balancing improves performance relative to baselines without load balancing (3.8x speedup) and with static load balancing (1.2x speedup). Our results provide important insights into dynamic load balancing and performance assessment, and are particularly relevant in the context of distributed memory applications ran on GPUs.

cs.DC

Rapid Exploration of Optimization Strategies on Advanced Architectures using TestSNAP and LAMMPS

The exascale race is at an end with the announcement of the Aurora and Frontier machines. This next generation of supercomputers utilize diverse hardware architectures to achieve their compute performance, providing an added onus on the performance portability of applications. An expanding fragmentation of programming models would provide a compounding optimization challenge were it not for the evolution of performance-portable frameworks, providing unified models for mapping abstract hierarchies of parallelism to diverse architectures. A solution to this challenge is the evolution of performance-portable frameworks, providing unified models for mapping abstract hierarchies of parallelism to diverse architectures. Kokkos is one such performance portable programming model for C++ applications, providing back-end implementations for each major HPC platform. Even with a performance portable framework, restructuring algorithms to expose higher degrees of parallelism is non-trivial. The Spectral Neighbor Analysis Potential (SNAP) is a machine-learned inter-atomic potential utilized in cutting-edge molecular dynamics simulations. Previous implementations of the SNAP calculation showed a downward trend in their performance relative to peak on newer-generation CPUs and low performance on GPUs. In this paper we describe the restructuring and optimization of SNAP as implemented in the Kokkos CUDA backend of the LAMMPS molecular dynamics package, benchmarked on NVIDIA GPUs. We identify novel patterns of hierarchical parallelism, facilitating a minimization of memory access overheads and pushing the implementation into a compute-saturated regime. Our implementation via Kokkos enables recompile-and-run efficiency on upcoming architectures. We find a $\sim$22x time-to-solution improvement relative to an existing implementation as measured on an NVIDIA Tesla V100-16GB for an important benchmark.

cs.DC

A Molecular-MNIST Dataset for Machine Learning Study on Diffraction Imaging and Microscopy

An image dataset of 10 different size molecules, where each molecule has 2,000 structural variants, is generated from the 2D cross-sectional projection of Molecular Dynamics trajectories. The purpose of this dataset is to provide a benchmark dataset for the increasing need of machine learning, deep learning and image processing on the study of scattering, imaging and microscopy.

eess.IV

Reproducibility in $G_0W_0$ Calculations for Solids

Ab initio many-body perturbation theory within the $GW$ approximation is a Green's function formalism widely used in the calculation of quasiparticle excitation energies of solids. In what has become an increasingly standard approach, Kohn-Sham eigenenergies, generated from a DFT calculation with a strategically-chosen exchange correlation functional ``starting point'', are used to construct $G$ and $W$, and then perturbatively corrected by the resultant $GW$ self-energy. In practice, there are several ways to construct the $GW$ self-energy, and these can lead to variations in predicted quasiparticle energies. For example, for ZnO and TiO$_2$, reported $GW$ fundamental gaps can vary by more than 1 eV. In this work, we address the convergence and key approximations in contemporary $G_0W_0$ calculations, including frequency-integration schemes and the treatment of the Coulomb divergence in the exact-exchange term. We study several systems,and compare three different $GW$ codes: BerkeleyGW, Abinit and Yambo. We demonstrate, for the first time, that the same quasiparticle energies for systems in the condensed phase can be obtained with different codes, and we provide a comprehensive assessment of implementations of the $GW$ approximation.

cond-mat.mtrl-sci

Exascale Deep Learning for Climate Analytics

We extract pixel-level masks of extreme weather patterns using variants of Tiramisu and DeepLabv3+ neural networks. We describe improvements to the software frameworks, input pipeline, and the network training algorithms necessary to efficiently scale deep learning on the Piz Daint and Summit systems. The Tiramisu network scales to 5300 P100 GPUs with a sustained throughput of 21.0 PF/s and parallel efficiency of 79.0%. DeepLabv3+ scales up to 27360 V100 GPUs with a sustained throughput of 325.8 PF/s and a parallel efficiency of 90.7% in single precision. By taking advantage of the FP16 Tensor Cores, a half-precision version of the DeepLabv3+ network achieves a peak and sustained throughput of 1.13 EF/s and 999.0 PF/s respectively.

cs.DC

Galactos: Computing the Anisotropic 3-Point Correlation Function for 2 Billion Galaxies

The nature of dark energy and the complete theory of gravity are two central questions currently facing cosmology. A vital tool for addressing them is the 3-point correlation function (3PCF), which probes deviations from a spatially random distribution of galaxies. However, the 3PCF's formidable computational expense has prevented its application to astronomical surveys comprising millions to billions of galaxies. We present Galactos, a high-performance implementation of a novel, O(N^2) algorithm that uses a load-balanced k-d tree and spherical harmonic expansions to compute the anisotropic 3PCF. Our implementation is optimized for the Intel Xeon Phi architecture, exploiting SIMD parallelism, instruction and thread concurrency, and significant L1 and L2 cache reuse, reaching 39% of peak performance on a single node. Galactos scales to the full Cori system, achieving 9.8PF (peak) and 5.06PF (sustained) across 9636 nodes, making the 3PCF easily computable for all galaxies in the observable universe.

astro-ph.CO

Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data

This paper presents the first, 15-PetaFLOP Deep Learning system for solving scientific pattern classification problems on contemporary HPC architectures. We develop supervised convolutional architectures for discriminating signals in high-energy physics data as well as semi-supervised architectures for localizing and classifying extreme weather in climate data. Our Intelcaffe-based implementation obtains $\sim$2TFLOP/s on a single Cori Phase-II Xeon-Phi node. We use a hybrid strategy employing synchronous node-groups, while using asynchronous communication across groups. We use this strategy to scale training of a single model to $\sim$9600 Xeon-Phi nodes; obtaining peak performance of 11.73-15.07 PFLOP/s and sustained performance of 11.41-13.27 PFLOP/s. At scale, our HEP architecture produces state-of-the-art classification accuracy on a dataset with 10M images, exceeding that achieved by selections on high-level physics-motivated features. Our semi-supervised architecture successfully extracts weather patterns in a 15TB climate dataset. Our results demonstrate that Deep Learning can be optimized and scaled effectively on many-core, HPC systems.

cs.PF

A structure preserving Lanczos algorithm for computing the optical absorption spectrum

We present a new structure preserving Lanczos algorithm for approximating the optical absorption spectrum in the context of solving full Bethe--Salpeter equation without Tamm--Dancoff approximation. The new algorithm is based on a structure preserving Lanczos procedure, which exploits the special block structure of Bethe--Salpeter Hamiltonian matrices. A recently developed technique of generalized averaged Gauss quadrature is incorporated to accelerate the convergence. We also establish the connection between our structure preserving Lanczos procedure with several existing Lanczos procedures developed in different contexts. Numerical examples are presented to demonstrate the effectiveness of our Lanczos algorithm.

math.NA

Low Rank Approximation in $G_0W_0$ Approximation

The single particle energies obtained in a Kohn--Sham density functional theory (DFT) calculation are generally known to be poor approximations to electron excitation energies that are measured in transport, tunneling and spectroscopic experiments such as photo-emission spectroscopy. The correction to these energies can be obtained from the poles of a single particle Green's function derived from a many-body perturbation theory. From a computational perspective, the accuracy and efficiency of such an approach depends on how a self energy term that properly accounts for dynamic screening of electrons is approximated. The $G_0W_0$ approximation is a widely used technique in which the self energy is expressed as the convolution of a non-interacting Green's function ($G_0$) and a screened Coulomb interaction ($W_0$) in the frequency domain. The computational cost associated with such a convolution is high due to the high complexity of evaluating $W_0$ at multiple frequencies. In this paper, we discuss how the cost of $G_0W_0$ calculation can be reduced by constructing a low rank approximation to the frequency dependent part of $W_0$. In particular, we examine the effect of such a low rank approximation on the accuracy of the $G_0W_0$ approximation. We also discuss how the numerical convolution of $G_0$ and $W_0$ can be evaluated efficiently and accurately by using a contour deformation technique with an appropriate choice of the contour.

math.NA

Structure Preserving Parallel Algorithms for Solving the Bethe-Salpeter Eigenvalue Problem

The Bethe-Salpeter eigenvalue problem is a dense structured eigenvalue problem arising from discretized Bethe-Salpeter equation in the context of computing exciton energies and states. A computational challenge is that at least half of the eigenvalues and the associated eigenvectors are desired in practice. We establish the equivalence between Bethe-Salpeter eigenvalue problems and real Hamiltonian eigenvalue problems. Based on theoretical analysis, structure preserving algorithms for a class of Bethe-Salpeter eigenvalue problems are proposed. We also show that for this class of problems all eigenvalues obtained from the Tamm-Dancoff approximation are overestimated. In order to solve large scale problems of practical interest, we discuss parallel implementations of our algorithms targeting distributed memory systems. Several numerical examples are presented to demonstrate the efficiency and accuracy of our algorithms.

math.NA

Effects of self-consistency and plasmon-pole models on GW calculations for closed-shell molecules

We present theoretical calculations of quasiparticle energies in closed-shell molecules using the GW method. We compare three different approaches: a full-frequency $G_0W_0$ (FF-$G_0W_0$) method with density functional theory (DFT-PBE) used as a starting mean field; a full-frequency $GW_0$ (FF-$GW_0$) method where the interacting Green's function is approximated by replacing the DFT energies with self-consistent quasiparticle energies or Hartree-Fock energies; and a $G_0W_0$ method with a Hybertsen-Louie generalized plasmon-pole model (HL GPP-$G_0W_0$). While the latter two methods lead to good agreement with experimental ionization potentials and electron affinities for methane, ozone, and beryllium oxide molecules, FF-$G_0W_0$ results can differ by more than one electron volt from experiment. We trace this failure of the FF-$G_0W_0$ method to the occurrence of incorrect self-energy poles describing shake-up processes in the vicinity of the quasiparticle energies.

cond-mat.mtrl-sci

Numerical integration for ab initio many-electron self energy calculations within the GW approximation

We present a numerical integration scheme for evaluating the convolution of a Green's function with a screened Coulomb potential on the real axis in the GW approximation of the self energy. Our scheme takes the zero broadening limit in Green's function first, replaces the numerator of the integrand with a piecewise polynomial approximation, and performs principal value integration on subintervals analytically. We give the error bound of our numerical integration scheme and show by numerical examples that it is more reliable and accurate than the standard quadrature rules such as the composite trapezoidal rule. We also discuss the benefit of using different self energy expressions to perform the numerical convolution at different frequencies.

physics.comp-ph

Tuning many-body interactions in graphene: The effects of doping on excitons and carrier lifetimes

The optical properties of graphene are strongly affected by electron-electron (e-e) and electron-hole (e-h) interactions. Here we tune these many-body interactions through varying the density of free charge carriers. Measurements from the infrared to the ultraviolet reveal significant changes in the optical conductivity of graphene for both electron and hole doping. The shift, broadening, and modification in shape of the saddle-point exciton resonance reflect strong screening of the many-body interactions by the carriers, as well as changes in quasi-particle lifetimes. Ab initio calculations by the GW Bethe-Salpeter equation (GW-BSE), which take into account modification of both the repulsive e-e and the attractive e-h interactions, provide excellent agreement with experiment. Understanding the optical properties and high-energy carrier dynamics of graphene over a wide range of doping is crucial for both fundamental graphene physics and for emerging applications of graphene in photonics.

cond-mat.mtrl-sci

BerkeleyGW: A Massively Parallel Computer Package for the Calculation of the Quasiparticle and Optical Properties of Materials and Nanostructures

BerkeleyGW is a massively parallel computational package for electron excited-state properties that is based on the many-body perturbation theory employing the ab initio GW and GW plus Bethe-Salpeter equation methodology. It can be used in conjunction with many density-functional theory codes for ground-state properties, including PARATEC, PARSEC, Quantum ESPRESSO, OCTOPUS and SIESTA. The package can be used to compute the electronic and optical properties of a wide variety of material systems from bulk semiconductors and metals to nanostructured materials and molecules. The package scales to 10,000's of CPUs and can be used to study systems containing up to 100's of atoms.

cond-mat.mtrl-sci

Coulomb-hole summations and energies for GW calculations with limited number of empty orbitals: a modified static remainder approach

Ab initio GW calculations are a standard method for computing the spectroscopic properties of many materials. The most computationally expensive part in conventional implementations of the method is the generation and summation over the large number of empty orbitals required to converge the electron self energy. We propose a scheme to reduce the summation over empty states by the use of a modified static-remainder approximation, which is simple to implement and yields accurate self energies for both bulk and molecular systems requiring a small fraction of the typical number of empty orbitals.

cond-mat.mtrl-sci