SearcharxivSearch

arXiv subjects

Andor Menczer

Publications and source records attributed to Andor Menczer.

11 recordsLinked to original sources

Block entropy area based non-local fermionic mode optimization with gradient disentanglers

We introduce a systematic block entropy area based mode optimization algorithm for many-body quantum states of interacting fermions represented by matrix product states. From the gradient of a global cost function, the block entropy area, a long-ranged, non-interacting effective disentangler Hamiltonian is formed. We then simulate the time-dependent Schr\"odinger equation driven by the disentangler Hamiltonian by employing the time-dependent variational principle based on projector splitting, and minimize the cost function. The combination of the density matrix renormalization group with this gradient-based entanglement minimization forms an efficient low-rank iterative ground-state algorithm that also provides an optimized single-particle basis for matrix product state representation. We demonstrate the method on two-dimensional lattice models of interacting fermions and the Fe${_4}$S${_4}$ cluster, and show its robustness and superiority over earlier protocols using nearest-neighbor mode rotations and reorderings.

cond-mat.str-el

Hunting for quantum advantage in electronic structure calculations is a highly non-trivial task

In light of major developments over the past decades in both quantum computing and simulations on classical hardware, it is a serious challenge to identify a real-world problem where quantum advantage is expected to appear. In quantum chemistry, electronic structure calculations of strongly correlated, i.e. multi-reference problems, are often argued to fall into such category because of their intractability with standard methods based on mean-field theory. Therefore, providing state-of-the-art benchmark data by classical algorithms is necessary to make a decisive conclusion when such competing development directions are compared. We report cutting-edge performance results together with high accuracy ground state energy for the Fe$_4$S$_4$ molecular cluster on a CAS(54,36) model space, a problem that has been included quite recently among the list of systems in the {\it Quantum Advantage Tracker} webpage maintained by IBM and RIKEN. Pushing the limits even further, we also present CAS-SCF based orbital optimizations for unprecedented CAS sizes of up to 89 electrons in 102 orbitals [CAS(89,102)] for the Fe$_5$S$_{12}$H$_4^{5-}$ molecular system comprising twenty five open shell orbitals in its sextet ground state and an active spaces size of 331 electrons in 451 orbitals. We have achieved our results via mixed-precision spin-adapted \textit{ab initio} Density Matrix Renormalization Group (DMRG) electronic structure calculations interfaced with the ORCA program package and utilizing the NVIDIA Blackwell graphics processing unit (GPU) platform. We argue that DMRG benchmark data should be taken as a classical reference when quantum advantage is reported. In addition, full exploitation of classical hardware should also be considered since even the most advanced DMRG implementations are still in a premature stage regarding utilization of all the benefits of GPU technology.

physics.chem-ph

Mixed-precision ab initio tensor network state methods adapted for NVIDIA Blackwell technology via emulated FP64 arithmetic

We report cutting-edge performance results via mixed-precision spin adapted ab initio Density Matrix Renormalization Group (DMRG) electronic structure calculations utilizing the Ozaki scheme for emulating FP64 arithmetic through the use of fixed-point compute resources. By approximating the underlying matrix and tensor algebra with operations on a modest number of fixed-point representatives (``slices''), we demonstrate on smaller benchmark systems and for the active compounds of the FeMoco and cytochrome P450 (CYP) enzymes with complete active space (CAS) sizes of up to 113 electrons in 76 orbitals [CAS(113, 76)] and 63 electrons in 58 orbitals [CAS(63, 58)], respectively, that the chemical accuracy can be reached with mixed-precision arithmetic. We also show that, due to its variational nature, DMRG provides an ideal tool to benchmark accuracy domains, as well as the performance of new hardware developments and related numerical libraries. Detailed numerical error analysis and performance assessment are also presented for subcomponents of the DMRG algebra by systematically interpolating between double- and pseudo-half-precision. Our analyis represents the first quantum chemistry evaluation of FP64 emulation for correlated calculations capable of achieving chemical accuracy and emulation based on fixed-point arithmetic, and it paves the way for the utilization of state-of-the-art Blackwell technology in tree-like tensor network state electronic structure calculations, opening new research directions in materials sciences and beyond.

physics.chem-ph

Orbital optimization of large active spaces via AI-accelerators

We present an efficient orbital optimization procedure that combines the highly GPU accelerated, spin-adapted density matrix renormalization group (DMRG) method with the complete active space self-consistent field (CAS-SCF) approach for quantum chemistry implemented in the ORCA program package. Leveraging the computational power of the latest generation of Nvidia GPU hardware, we perform CAS-SCF based orbital optimizations for unprecedented CAS sizes of up to 82 electrons in 82 orbitals [CAS(82,82)] in molecular systems comprising of active spaces sizes of hundreds of electrons in thousands of orbitals. For both the NVIDIA DGX-A100 and DGX-H100 hardware, we provide a detailed scaling and error analysis of our DMRG-SCF approach for benchmark systems consisting of polycyclic aromatic hydrocarbons and iron-sulfur complexes of varying sizes. Our efforts demonstrate for the first time that highly accurate DMRG calculations at large bond dimensions are critical for obtaining reliably converged CAS-SCF energies. For the more challenging iron-sulfur benchmark systems, we furthermore find the optimized orbitals of a converged CAS-SCF calculation to depend more sensitively on the DMRG parameters than those for the polycyclic aromatic hydrocarbons. The ability to obtain converged CAS-SCF energies and orbitals for active spaces of such large sizes within days reduces the challenges of including the appropriate orbitals into the CAS or selecting the correct minimal CAS, and may open up entirely new avenues for tackling strongly correlated molecular systems.

physics.chem-ph

Tensor network state methods and quantum information theory for strongly correlated molecular systems

A brief pedagogical overview of recent advances in tensor network state methods are presented that have the potential to broaden their scope of application radically for strongly correlated molecular systems. These include global fermionic mode optimization, i.e., a general approach to find an optimal matrix product state (MPS) parametrization of a quantum many-body wave function with the minimum number of parameters for a given error margin, the restricted active space DMRG-RAS-X method, multi-orbital correlations and entanglement, developments on hybrid CPU-multiGPU parallelization, and an efficient treatment of non-Abelian symmetries on high-performance computing (HPC) infrastructures. Scaling analysis on NVIDIA DGX-A100 platform is also presented.

cond-mat.str-el

Cost optimized ab initio tensor network state methods: industrial perspectives

We introduce efficient solutions to optimize the cost of tree-like tensor network state method calculations when an expensive GPU-accelerated hardware is utilized. By supporting a main powerful compute node with additional auxiliary, but much cheaper nodes to store intermediate, precontracted tensor network scratch data, the IO time can be hidden behind the computation almost entirely without increasing memory peak. Our solution is based on the different bandwidths of the different communication channels, like NVLink, PCIe, InfiniBand and available storage media, which are utilized on different layers of the algorithm. This simple heterogeneous multiNode solution via asynchronous IO operation has the potential to minimize IO overhead, resulting in maximum performance rate for the main compute unit. In addition, we introduce an in-house developed massively parallel protocol to serialize and deserialize block sparse matrices and tensors, reducing data communication time tremendously. Performance profiles are presented for the spin adapted ab initio density matrix renormalization group method for corresponding U(1) bond dimension values up to 15400 on the active compounds of the FeMoco with complete active space (CAS) sizes of up to 113 electrons in 76 orbitals [CAS(113, 76)].

physics.comp-ph

Parallel implementation of the Density Matrix Renormalization Group method achieving a quarter petaFLOPS performance on a single DGX-H100 GPU node

We report cutting edge performance results for a hybrid CPU-multi GPU implementation of the spin adapted ab initio Density Matrix Renormalization Group (DMRG) method on current state-of-the-art NVIDIA DGX-H100 architectures. We evaluate the performance of the DMRG electronic structure calculations for the active compounds of the FeMoco and cytochrome P450 (CYP) enzymes with complete active space (CAS) sizes of up to 113 electrons in 76 orbitals [CAS(113, 76)] and 63 electrons in 58 orbitals [CAS(63, 58)], respectively. We achieve 246 teraFLOPS of sustained performance, an improvement of more than 2.5x compared to the performance achieved on the DGX-A100 architectures and an 80x acceleration compared to an OpenMP parallelized implementation on a 128-core CPU architecture. Our work highlights the ability of tensor network algorithms to efficiently utilize high-performance GPU hardware and shows that the combination of tensor networks with modern large-scale GPU accelerators can pave the way towards solving some of the most challenging problems in quantum chemistry and beyond.

physics.chem-ph

Global fermionic mode optimization via swap gates

An optimal choice of single-particle modes leads to an optimal tensor network-based compression of the many-body wave function for systems of indistinguishable particles, which compression can be essential for large, close-to-critical systems. The proposed mode optimization relies on the minimization of the block entropy area, a global quantity that measures entanglement of the wave function along the one-dimensional chain of modes. Extension of the two-site DMRG algorithm by nearest-neighbor mode optimizations is straightforward, but it performs poorly without systematic reordering of modes. Moreover, finding a stationary optimal point in the combination of the fixed-rank MPS manifold and the unitary group of mode rotations requires optimization for every generator, i.e., for every two-mode pair. Here, a systematic joint optimization protocol over the MPS manifold and the full unitary group is presented in which every two-mode pair is accessed by a carefully constructed sequence of two-qubit operations, including swap-gate controlled permutations. Large-scale DMRG simulations of strongly correlated two-dimensional fermionic lattice models and the multireference Fe$_4$S$_4$ transition metal cluster demonstrate rapid convergence towards the optimum, achieving low energy and low entanglement.

cond-mat.str-el

Two dimensional quantum lattice models via mode optimized hybrid CPU-GPU density matrix renormalization group method

We present a hybrid numerical approach to simulate quantum many body problems on two spatial dimensional quantum lattice models via the non-Abelian ab initio version of the density matrix renormalization group method on state-of-the-art high performance computing infrastructures. We demonstrate for the two dimensional spinless fermion model and for the Hubbard model on torus geometry that altogether several orders of magnitude in computational time can be saved by performing calculations on an optimized basis and by utilizing hybrid CPU-multiGPU parallelization. At least an order of magnitude reduction in computational complexity results from mode optimization, while a further order of reduction in wall time is achieved by massive parallelization. Our results are measured directly in FLOP and seconds. A detailed scaling analysis of the obtained performance as a function of matrix ranks and as a function of system size up to $12\times 12$ lattice topology is discussed. Our CPU-multiGPU model also tremendously accelerates the calculation of the one- and two-particle reduced density matrices, which can be used to construct various order parameters and trace quantum phase transitions with high fidelity.

cond-mat.str-el

Boosting the effective performance of massively parallel tensor network state algorithms on hybrid CPU-GPU based architectures via non-Abelian symmetries

We present novel algorithmic solutions together with implementation details utilizing non-Abelian symmetries in order to boost the current limits of tensor network state algorithms on high performance computing infrastructure. In our in-house developed hybrid CPU-multiGPU solution scheduling is decentralized, threads are autonomous and inter-thread communications are solely limited to interactions with globally visible lock-free constructs. Our custom tailored virtual memory management ensures data is produced with high spatial locality, which together with the use of specific sequences of strided batched matrix operations translates to significantly higher overall throughput. In order to lower IO overhead, an adaptive buffering technique is used to dynamically match the level of data abstraction, at which cache repositories are built and reused, to system resources. The non-Abelian symmetry related tensor algebra based on Wigner-Eckhart theorem is fully detached from the conventional tensor network layer, thus massively parallel matrix and tensor operations can be performed without additional overheads. Altogether, we have achieved an order of magnitude increase in performance with respect to results reported in arXiv:2305.05581 in terms of computational complexity and at the same time a factor of three to six in the actual performance measured in TFLOPS. Benchmark results are presented on Hilbert space dimensions up to $2.88\times10^{36}$ obtained via large-scale SU(2) spin adapted density matrix renormalization group simulations on selected strongly correlated molecular systems. These demonstrate the utilization of NVIDIA's highly specialized tensor cores, leading to performance around 110 TFLOPS on a single node supplied with eight NVIDIA A100 devices. In comparison to U(1) implementations with matching accuracy, our solution has an estimated effective performance of 250-500 TFLOPS.

physics.comp-ph

Massively Parallel Tensor Network State Algorithms on Hybrid CPU-GPU Based Architectures

The interplay of quantum and classical simulation and the delicate divide between them is in the focus of massively parallelized tensor network state (TNS) algorithms designed for high performance computing (HPC). In this contribution, we present novel algorithmic solutions together with implementation details to extend current limits of TNS algorithms on HPC infrastructure building on state-of-the-art hardware and software technologies. Benchmark results obtained via large-scale density matrix renormalization group (DMRG) simulations are presented for selected strongly correlated molecular systems addressing problems on Hilbert space dimensions up to $2.88\times10^{36}$.

quant-ph