SearcharxivSearch

arXiv subjects

Jingde Bu

Publications and source records attributed to Jingde Bu.

2 recordsLinked to original sources

Trillion-atom molecular dynamics simulations with ab initio accuracy

Material properties are fundamentally dictated by multiscale phenomena, which often reach mesoscale in size. The {\mu}m mesoscale is also the size which can be observed directly under an optical microscope, bridging the atomistic microscopic description with the continuous model macroscopic world. In this work, we report an unprecedented molecular dynamics (MD) simulation comprising 1.62 trillion atoms. Utilizing the neuroevolution potential (NEP) framework, we attained ab initio accuracy on China's New-generation Intelligent Supercomputer. Our implementation achieves a time-to-solution (s/step/atom) 100 times faster than previous state-of-the-art machine learning force field simulations, and 1,000 times faster than the Gordon Bell Prize-winning application from six years ago. Furthermore, we demonstrate an 86.9% weak scaling efficiency from a single GPGPU to 45,000 GPGPUs. These results redefine atomistic simulation boundaries, enabling direct mesoscopic modeling with quantum-level precision.

cond-mat.mtrl-sci

Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials

Universal Machine Learning Interatomic Potentials (uMLIPs), pre-trained on massively diverse datasets encompassing inorganic materials and organic molecules across the entire periodic table, serve as foundational models for quantum-accurate physical simulations. However, uMLIP training requires second-order derivatives, which lack corresponding parallel training frameworks; moreover, scaling to the billion-parameter regime causes explosive growth in computation and communication overhead, making its training a tremendous challenge. We introduce MatRIS-MoE, a billion-parameter Mixture-of-Experts model built upon invariant architecture, and {Janus}, a pioneering high-dimensional distributed training framework for uMLIPs with hardware-aware optimizations. Deployed across two Exascale supercomputers, our code attains a peak performance of 1.2/1.0 EFLOPS (24\%/{35.5\%} of theoretical peak) in single precision at over 90\% parallel efficiency, compressing the training of billion-parameter uMLIPs from weeks to hours. This work establishes a new high-water mark for AI-for-Science (AI4S) foundation models at Exascale and provides essential infrastructure for rapid scientific discovery.

cs.DC