SearcharxivSearch

arXiv subjects

Ahmed E. Ismail

Publications and source records attributed to Ahmed E. Ismail.

14 recordsLinked to original sources

The Tersoff many-body potential: Sustainable performance through vectorization

Molecular dynamics models materials by simulating each individual particle's trajectory. Many-body potentials lead to a more accurate trajectory simulation, and are used in materials science and computational chemistry. We present optimization results for one multi-body potential on a range of vector instruction sets, targeting both CPUs and accelerators like the Intel Xeon Phi. Parallelization of MD simulations is well-studied; by contrast, vectorization is relatively unexplored. Given the prevalence and power of modern vector units, exploiting them is imperative for high performance software. When running on a highly parallel machine, any improvement to the scalar performance is paid back in hundreds or thousands of saved core hours. Vectorization is already commonly used in the optimization or pair potentials; multi-body potentials pose new, unique challenges. Indeed, their optimization pushes the boundaries of current compilers, forcing us to use explicit vectorization techniques for now. In this study, we add an optimized implementation of Tersoff potential to the LAMMPS molecular dynamics simulation package. To reduce the burden of explicit vectorization, we abstract from the specific vector instruction set and desired precision: From one algorithm, we get optimized implementations for many platforms, from SSE4.2 to AVX512, and the Intel Xeon Phi. We compare the kernels across different architectures, and determine suitable architecture-dependent parameters. Our optimizations benefit any architecture, but have a disproportionate effect on the Intel Xeon Phi, which beats the CPU (2xE5-2650) after optimization.

cs.CE

LAMMPS' PPPM Long-Range Solver for the Second Generation Xeon Phi

Molecular Dynamics is an important tool for computational biologists, chemists, and materials scientists, consuming a sizable amount of supercomputing resources. Many of the investigated systems contain charged particles, which can only be simulated accurately using a long-range solver, such as PPPM. We extend the popular LAMMPS molecular dynamics code with an implementation of PPPM particularly suitable for the second generation Intel Xeon Phi. Our main target is the optimization of computational kernels by means of vectorization, and we observe speedups in these kernels of up to 12x. These improvements carry over to LAMMPS users, with overall speedups ranging between 2-3x, without requiring users to retune input parameters. Furthermore, our optimizations make it easier for users to determine optimal input parameters for attaining top performance.

cs.CE

A Scalable, Linear-Time Dynamic Cutoff Algorithm for Molecular Dynamics

Recent results on supercomputers show that beyond 65K cores, the efficiency of molecular dynamics simulations of interfacial systems decreases significantly. In this paper, we introduce a dynamic cutoff method (DCM) for interfacial systems of arbitrarily large size. The idea consists in adopting a cutoff-based method in which the cutoff is cho- sen on a particle-by-particle basis, according to the distance from the interface. Computationally, the challenge is shifted from the long-range solvers to the detection of the interfaces and to the computation of the particle-interface distances. For these tasks, we present linear-time algorithms that do not rely on global communication patterns. As a result, the DCM algorithm is suited for large systems of particles and mas- sively parallel computers. To demonstrate its potential, we integrated DCM into the LAMMPS open-source molecular dynamics package, and simulated large liquid/vapor systems on two supercomputers: SuperMuc and JUQUEEN. In all cases, the accuracy of DCM is comparable to the traditional particle-particle particle-mesh (PPPM) algorithm, while the performance is considerably superior for large numbers of particles. For JUQUEEN, we provide timings for simulations running on the full system (458, 752 cores), and show nearly perfect strong and weak scaling.

physics.comp-ph

Accelerating scientific codes by performance and accuracy modeling

Scientific software is often driven by multiple parameters that affect both accuracy and performance. Since finding the optimal configuration of these parameters is a highly complex task, it extremely common that the software is used suboptimally. In a typical scenario, accuracy requirements are imposed, and attained through suboptimal performance. In this paper, we present a methodology for the automatic selection of parameters for simulation codes, and a corresponding prototype tool. To be amenable to our methodology, the target code must expose the parameters affecting accuracy and performance, and there must be formulas available for error bounds and computational complexity of the underlying methods. As a case study, we consider the particle-particle particle-mesh method (PPPM) from the LAMMPS suite for molecular dynamics, and use our tool to identify configurations of the input parameters that achieve a given accuracy in the shortest execution time. When compared with the configurations suggested by expert users, the parameters selected by our tool yield reductions in the time-to-solution ranging between 10% and 60%. In other words, for the typical scenario where a fixed number of core-hours are granted and simulations of a fixed number of timesteps are to be run, usage of our tool may allow up to twice as many simulations. While we develop our ideas using LAMMPS as computational framework and use the PPPM method for dispersion as case study, the methodology is general and valid for a range of software tools and methods.

cs.CE

The Vectorization of the Tersoff Multi-Body Potential: An Exercise in Performance Portability

Molecular dynamics simulations, an indispensable research tool in computational chemistry and materials science, consume a significant portion of the supercomputing cycles around the world. We focus on multi-body potentials and aim at achieving performance portability. Compared with well-studied pair potentials, multibody potentials deliver increased simulation accuracy but are too complex for effective compiler optimization. Because of this, achieving cross-platform performance remains an open question. By abstracting from target architecture and computing precision, we develop a vectorization scheme applicable to both CPUs and accelerators. We present results for the Tersoff potential within the molecular dynamics code LAMMPS on several architectures, demonstrating efficiency gains not only for computational kernels, but also for large-scale simulations. On a cluster of Intel Xeon Phi's, our optimized solver is between 3 and 5 times faster than the pure MPI reference.

cs.CE

A Note on Time Measurements in LAMMPS

We examine the issue of assessing the efficiency of components of a parallel program at the example of the MD package LAMMPS. In particular, we look at how LAMMPS deals with the issue and explain why the approach adopted might lead to inaccurate conclusions. The misleading nature of this approach is subsequently verified experimentally with a case study. Afterwards, we demonstrate how one should correctly determine the efficiency of the components and show what changes to the code base of LAMMPS are necessary in order to get the correct behavior.

cond-mat.mtrl-sci

Coating thickness and coverage effects on the forces between silica nanoparticles in water

The structure and interactions of coated silica nanoparticles have been studied in water using molecular dynamics simulations. For 5 nm diameter amorphous silica nanoparticles we studied the effects of varying the chain length and grafting density of polyethylene oxide (PEO) on the nanoparticle coating's shape and on nanoparticle-nanoparticle effective forces. For short ligands of length $n=6$ and $n=20$ repeat units, the coatings are radially symmetric while for longer chains ($n=100$) the coatings are highly anisotropic. This anisotropy appears to be governed primarily by chain length, with coverage playing a secondary role. For the largest chain lengths considered, the strongly anisotropic shape makes fitting to a simple radial force model impossible. For shorter ligands, where the coatings are isotropic, we found that the force between pairs of nanoparticles is purely repulsive and can be fit to the form $(R/2r_\text{core}-1)^{-b}$ where $R$ is the separation between the center of the nanoparticles, $r_\text{core}$ is the radius of the silica core, and $b$ is measured to be between 2.3 and 4.1.

cond-mat.soft

Multilevel Summation for Dispersion: A Linear-Time Algorithm for $r^{-6}$ Potentials

We have extended the multilevel summation (MLS) method, originally developed to evaluate long-range Coulombic interactions in molecular dynamics (MD) simulations [Skeel et al., J. Comput. Chem., 23, 673 (2002)], to handle dispersion interactions. While dispersion potentials are formally short-ranged, accurate calculation of forces and energies in interfacial and inhomogeneous systems require long-range methods. The MLS method offers some significant advantages compared to the particle-particle particle-mesh and smooth particle mesh Ewald methods. Unlike mesh-based Ewald methods, MLS does not use fast Fourier transforms and is thus not limited by communication and bandwidth concerns. In addition, it scales linearly in the number of particles, as compared with the $\mathcal{O}(N \log N)$ complexity of the mesh-based Ewald methods. While the structure of the MLS method is invariant for different potentials, every algorithmic step had to be adapted to accommodate the $r^{-6}$ form of the dispersion interactions. In addition, we have derived error bounds, similar to those obtained by Hardy for the electrostatic MLS [Hardy, Ph.D. thesis, University of Illinois at Urbana-Champaign (2006)]. Using a prototype implementation, we have demonstrated the linear scaling of the MLS method for dispersion, and present results establishing the accuracy and efficiency of the method.

cond-mat.mtrl-sci

Development and application of a particle-particle particle-mesh Ewald method for dispersion interactions

For inhomogeneous systems with interfaces, the inclusion of long-range dispersion interactions is necessary to achieve consistency between molecular simulation calculations and experimental results. For accurate and efficient incorporation of these contributions, we have implemented a particle-particle particle-mesh (PPPM) Ewald solver for dispersion ($r^{-6}$) interactions into the LAMMPS molecular dynamics package. We demonstrate that the solver's $\mathcal{O}(N\log N)$ scaling behavior allows its application to large-scale simulations. We carefully determine a set of parameters for the solver that provides accurate results and efficient computation. We perform a series of simulations with Lennard-Jones particles, SPC/E water, and hexane to show that with our choice of parameters the dependence of physical results on the chosen cutoff radius is removed. Physical results and computation time of these simulations are compared to results obtained using either a plain cutoff or a traditional Ewald sum for dispersion.

cond-mat.mtrl-sci

Forces between functionalized silica nanoparticles in solution

To prevent the flocculation and phase separation of nanoparticles in solution, nanoparticles are often functionalized with short chain surfactants. Here we present fully-atomistic molecular dynamics simulations which characterize how these functional coatings affect the interactions between nanoparticles and with the surrounding solvent. For 5 nm diameter silica nanoparticles coated with poly(ethylene oxide) (PEO) oligomers in water, we determined the hydrodynamic drag on two approaching nanoparticles moving through solvent and on a single nanoparticle as it approaches a planar surface. In most circumstances, acroscale fluid theory accurately predicts the drag on these nano-scale particles. Good agreement is seen with Brenner's analytical solutions for wall separations larger than the soft nanoparticle radius. For two approaching coated nanoparticles, the solvent-mediated (velocity-independent) and lubrication (velocity-dependent) forces are purely repulsive and do not exhibit force oscillations that are typical of uncoated rigid spheres.

cond-mat.mes-hall

Application of Ewald summations to long-range dispersion forces

We present results illustrating the effects of using explicit summation terms for the $r^{-6}$ dispersion term on the interfacial properties of a Lennard-Jones fluid and SPC/E water. For the Lennard-Jones fluid, we find that the use of long-range summations, even with a short ``crossover radius,'' yields results that are consistent with simulations using large cutoff radii. Simulations of SPC/E water demonstrate that the long-range dispersion forces are of secondary importance to the Coulombic forces. In both cases, we find that the ratio of box size $L_{\parallel}$ to crossover radius $r_{\rm c}^{\mathbf k}$ plays an important role in determining the magnitude of the long-range dispersion correction, although its effect is secondary when Coulombic interactions are also present.

cond-mat.mtrl-sci

Capillary waves at the liquid-vapor interface and the surface tension of water models

Capillary waves occurring at the liquid-vapor interface of water are studied using molecular dynamics simulations. In addition, the surface tension, determined thermodynamically from the difference in the normal and tangential pressure at the liquid-vapor interface, is compared for a number of standard three- and four-point water models. We study four three-point models (SPC/E, TIP3P, TIP3P-CHARMM, and TIP3P-Ew) and two four-point models (TIP4P and TIP4P-Ew). All of the models examined underestimate the surface tension; the TIP4P-Ew model comes closest to reproducing the experimental data. The surface tension can also be determined from the amplitude of capillary waves at the liquid-vapor interface by varying the surface area of the interface. The surface tensions determined from the amplitude of the logarithmic divergence of the capillary interfacial width and from the traditional thermodynamic method agree only if the density profile is fitted to an error function instead of a hyperbolic tangent function.

cond-mat.mtrl-sci

Multiresolution analysis in statistical mechanics. I. Using wavelets to calculate thermodynamic properties

The wavelet transform, a family of orthonormal bases, is introduced as a technique for performing multiresolution analysis in statistical mechanics. The wavelet transform is a hierarchical technique designed to separate data sets into sets representing local averages and local differences. Although one-to-one transformations of data sets are possible, the advantage of the wavelet transform is as an approximation scheme for the efficient calculation of thermodynamic and ensemble properties. Even under the most drastic of approximations, the resulting errors in the values obtained for average absolute magnetization, free energy, and heat capacity are on the order of 10%, with a corresponding computational efficiency gain of two orders of magnitude for a system such as a $4\times 4$ Ising lattice. In addition, the errors in the results tend toward zero in the neighborhood of fixed points, as determined by renormalization group theory.

physics.chem-ph

Multiresolution analysis in statistical mechanics. II. The wavelet transform as a basis for Monte Carlo simulations on lattices

In this paper, we extend our analysis of lattice systems using the wavelet transform to systems for which exact enumeration is impractical. For such systems, we illustrate a wavelet-accelerated Monte Carlo (WAMC) algorithm, which hierarchically coarse-grains a lattice model by computing the probability distribution for successively larger block spins. We demonstrate that although the method perturbs the system by changing its Hamiltonian and by allowing block spins to take on values not permitted for individual spins, the results obtained agree with the analytical results in the preceding paper, and ``converge'' to exact results obtained in the absence of coarse-graining. Additionally, we show that the decorrelation time for the WAMC is no worse than that of Metropolis Monte Carlo (MMC), and that scaling laws can be constructed from data performed in several short simulations to estimate the results that would be obtained from the original simulation. Although the algorithm is not asymptotically faster than traditional MMC, because of its hierarchical design, the new algorithm executes several orders of magnitude faster than a full simulation of the original problem. Consequently, the new method allows for rapid analysis of a phase diagram, allowing computational time to be focused on regions near phase transitions.

physics.chem-ph