SearcharxivSearch

arXiv subjects

Dengdong Fan

Publications and source records attributed to Dengdong Fan.

10 recordsLinked to original sources

Ascend to Science: Exploration of AI Chips for Scientific Computing

The rapid rise of AI-oriented accelerators has reshaped compute systems around low-precision tensor engines, raising a practical question for the HPC community: under what conditions can such hardware support scientific workloads that demand numerical robustness, irregular memory access, and scalability? Using the Ascend 910 NPU series as a representative tensor-centric platform, we characterize precision, execution, and memory-hierarchy bottlenecks that hinder the direct deployment of scientific codes. We then develop and evaluate workload-specific mappings across five application studies -- HPL-MxP, LRSVD, SGEMM-cube, PQSim, and SMC-X -- combining heterogeneous execution, mixed-precision numerical formulations, precision emulation, hierarchical memory orchestration, and communication--computation overlap. These studies show that AI-native NPUs can achieve numerical robustness, competitive performance, and satisfactory scalability when numerical formulation, execution placement, and data movement are addressed in a coordinated manner. Our results provide a state-of-the-practice case study of how scientific workloads can be adapted to tensor-centric architectures, while distinguishing transferable optimization principles from Ascend-specific implementation details.

cs.DC

PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent

Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of gradient first and second moments, incurring substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by steepest descent under an $\ell_p$-norm geometry, we show that applying a nonlinear transform directly to a momentum buffer yields coordinate-wise adaptivity. We prove that PowerStep converges at the optimal $O(1/\sqrt{T})$ rate for non-convex stochastic optimization. Extensive experiments on Transformer models ranging from 124M to 235B parameters demonstrate that PowerStep matches Adam's convergence speed while halving optimizer memory. Furthermore, when combined with aggressive \texttt{int8} quantization, PowerStep remains numerically stable and reduces optimizer memory by $\sim\!8\times$ compared to full-precision Adam. PowerStep thus provides a principled, scalable and resource-efficient alternative for large-scale training. Code is available at https://github.com/yaolubrain/PowerStep.

cs.LG

SMC-AI: Scaling Monte Carlo Simulation to Four Trillion Atoms with AI Accelerators

The rapid advancement of deep learning is reshaping the hardware design landscape toward AI tasks, posing fundamental challenges for HPC workloads such as atomistic simulation. Here we present SMC-AI, a general algorithmic framework that extends the SMC-X method for efficient canonical Monte Carlo simulation on AI accelerators, including GPUs and NPUs, while maintaining extreme scalability. The implementation of SMC-AI on an NPU cluster reaches unprecedented performance, achieving MC simulation of 4 trillion atoms on 4096 NPU dies. This represents the largest ML-accelerated atomistic simulation reported, delivering 32X system size and 1.3X throughput than previous records, with a relatively small computational budget. Excellent strong and weak scaling efficiency are reached for both the NPU and GPU implementation. By decoupling ML models from simulation, SMC-AI creates an abstraction that facilitates integration and porting of diverse ML models, laying a foundation for the future development of scalable scientific software.

physics.comp-ph

PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning

We present PCL-Reasoner-V1.5, a 32-billion-parameter large language model (LLM) for mathematical reasoning. The model is built upon Qwen2.5-32B and refined via supervised fine-tuning (SFT) followed by reinforcement learning (RL). A central innovation is our proposed offline RL method, which provides superior training stability and efficiency over standard online RL methods such as GRPO. Our model achieves state-of-the-art performance among models post-trained on Qwen2.5-32B, attaining average accuracies of 90.9% on AIME 2024 and 85.6% on AIME 2025. Our work demonstrates offline RL as a stable and efficient paradigm for advancing reasoning in LLMs. All experiments were conducted on Huawei Ascend 910C NPUs.

cs.LG

Revealing Nanostructures in High-Entropy Alloys via Machine-Learning Accelerated Scalable Monte Carlo Simulation

The computational cost of traditional first-principles method quickly becomes prohibitively expensive as the number of atoms increases. This challenge is further amplified by the need to evaluate finite-temperature properties with Monte Carlo (MC) simulations, which is inherently challenging to parallelize due to sequential Markov chain updates. Here, we introduce Scalable Monte Carlo (SMC), an efficient MC simulation method that overcomes the parallelization bottlenecks in conventional MC simulation, reducing the computational complexity of a MC sweep from quadratic to linear. We present a GPU implementation of the SMC method, SMC-GPU, which simultaneously harnesses the thousands of processing cores on a GPU to accelerate the computation. By adopting a data-driven workflow that surrogates the computationally expensive density functional theory (DFT) with ML models, we demonstrate that SMC-GPU is capable of simulating systems of more than one-billion atoms, while maintaining the accuracy of first-principles methods. Using this unprecedented capability, we performed billion-atom MC simulations to investigate the nanostructure evolution of two important high-entropy alloys (HEAs), FeCoNiAlTi and MoNbTaW, in which the nanostructures are believed to be responsible for their superb mechanical properties. Our results reveal a rich diversity of nanostructures, including nanoparticles (NP), 3D-connected NP, and disorder protected nanophases. We quantitatively analyze the size, composition, and morphology of the nanostructures, as well as directly simulate the atom-probe-tomography (APT) needle. The results align well with available experimental observations. This work underscores the promising potential of leveraging large-scale MC simulation to explore the largely uncharted territory of nanostructure evolution in HEAs.

cond-mat.mtrl-sci

The impacts of optimization algorithm and basis size on the accuracy and efficiency of variational quantum eigensolver

Variational quantum eigensolver (VQE) is demonstrated to be the promising methodology for quantum chemistry based on near-term quantum devices. However, many problems are yet to be investigated for this methodology, such as the influences of optimization algorithm and basis size on the accuracy and efficiency for quantum computing. To address these issues, five molecules (H2, LiH, HF, N2 and F2) are studied in this work based on the VQE method using unitary coupled cluster (UCC) ansatz. The performance of the gradient optimization L-BFGS-B is compared with that of the direct search method COBYLA. The former converges more quickly, but the accuracy of energy surface is a little lower. The basis set shows a vital influence on the accuracy and efficiency. A large basis set generally provides an accurate energy surface, but induces a significant increase in computing time. The 631g basis is generally required from the energy surface of the simplest H2 molecule. For practical applications of VQE, complete active space (CAS) is suggested based on limited quantum resources. With the same number of qubits, more occupied orbitals included in CAS gives a better accuracy for the energy surface and a smaller evaluation number in the VQE optimization. Additionally, the electronic structure, such as filling fraction of orbitals, the bond strength of a molecule and the maximum nuclear charge also influences the performance of optimization, where half occupation of orbitals generally requires a large computation cost.

physics.chem-ph

High thermoelectric performance of two-dimensional (PbTe)2 layer

The electronic, phonon and thermoelectric transport properties of (PbTe)2 layer are systematically investigated by using first-principles pseudopotential method and Boltzmann transport equation. Our calculations demonstrate that there is a valley degeneracy of six for the top valence band, which leads to larger carrier concentration and thus higher electrical conductivity without obvious reduction in the Seebeck coefficient. Moreover, the intrinsic van der Waals interactions between neighboring Pb layers induce additional phonon scattering and thus ultrasmall lattice thermal conductivity. As a consequence, a maximum p-type ZT value of 2.9 can be achieved at 1000 K. Moreover, we find almost identical n- and p-type ZT in the temperature range from 300 K to 800 K.

cond-mat.mtrl-sci

Designing graphene/hexagonal boron nitride superlattice monolayer with high thermoelectric performance

We design a hybrid graphene/hexagonal boron nitride superlattice monolayer and investigate its thermoelectric properties using density functional theory and Boltzmann transport equations with the relaxation time accurately treated by electron-phonon coupling calculations. Compared with that of pristine graphene, the lattice thermal conductivity of the superlattice structure is more than two orders of magnitude lower due to the enhanced three-phonon scattering process originated from the mixed-bond characteristics. Besides, the coexistence of light and heavy bands around the Fermi level leads to an ultrahigh power factor along the zigzag direction, where the highest ZT value of ~2.5 can be achieved for the n-type system at 1100 K. Moreover, it is noted that the carrier transport near the valance band minimum is almost entirely contributed by the graphene part of the superlattice. As a consequence, the thermoelectric performance of p-type system can be enhanced to be comparable with that of n-type one by appropriate substitution of nitrogen atom with phosphorus, which can suppress the lattice thermal conductivity but nearly have no influence on the hole transport.

cond-mat.mtrl-sci

Large scale calculations of thermoelectric transport coefficients: a case study of γ-graphyne with point defects

Defects such as vacancies and impurities could have profound effects on the transport properties of thermoelectric materials. However, it is usually quite difficult to directly calculate the thermoelectric properties of defect-containing systems via first-principles method since very large supercell is required. In this work, based on the linear response theory and the kernel polynomial method, we present an efficient approach that can help to calculate the thermoelectric transport coefficients of a large system containing millions of atoms at arbitrary chemical potential and temperature. As a prototype example, we consider dilute vacancies and hydrogen impurities in a large scale γ-graphyne sheet and discuss their effects on the thermoelectric transport properties.

physics.comp-ph

Phonon-limited electrical transport properties of intermetallic compound YbAl3 from first-principles calculations

We combine first-principles calculations and Boltzmann transport theory to study the electrical transport properties of intermetallic compound YbAl3. To accurately predict the electronic relaxation time, we use the density functional perturbation theory and Wannier interpolation techniques which can effectively treat the electron-phonon scattering. Our calculated transport coefficients of YbAl3 are in reasonable agreement with the experimentally measured results. Strikingly, we discover that in evaluating the Seebeck coefficient of YbAl3, the scattering term has a larger contribution than the band term and should be explicitly considered in the calculations, especially for the case with localized bands near the Fermi level. Moreover, we demonstrate that by reducing the sample size to less than ~30 nm, the electronic thermal conductivity of YbAl3 can be sufficiently suppressed so that the thermoelectric figure of merit can be further enhanced.

cond-mat.mes-hall