SearcharxivSearch

arXiv subjects

Danish Khan

Publications and source records attributed to Danish Khan.

13 recordsLinked to original sources

Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory

Kohn--Sham density functional theory (DFT) underpins electronic-structure simulations, but repeated orbital diagonalizations lead to cubic scaling, restricting quantum calculations to modest scales only. Eliminating these auxiliary orbitals while retaining Kohn--Sham accuracy is the central goal of orbital-free DFT, but both analytical and machine-learning methods have so far fallen short. Prior learning approaches either try to learn the variational kinetic-energy functionals, which are ill-conditioned, or directly predict the ground state, which extrapolate poorly to larger systems. Instead, we identify the Kohn--Sham map as the right learning target for orbital-free DFT. It maps a Kohn--Sham potential directly to the corresponding density and noninteracting kinetic energy, quantities otherwise obtained through an orbital diagonalization. Focusing on the density component in this work, a domain-invariant $\mathrm{SE}(3)$-equivariant Fourier neural operator learns to predict it from the potential as input on real-space grids, enabling stable quasi-linear scaling SCFs. Trained jointly on 8,504 molecules and solids, a single model generalizes to out-of-distribution organic molecules, insulators, and metals. For the first time, the same method converges SCFs across these systems without explicitly constructing Kohn--Sham orbitals, while reproducing densities, electronic spectra, and structural observables at Kohn--Sham DFT accuracy. Linear-scaling SCFs additionally allow converging magnesium dislocation densities containing up to 82,500 valence electrons on a single GPU.

physics.chem-ph

OrbitAll: A Unified Quantum Mechanical Representation Deep Learning Framework for All Molecular Systems

We introduce OrbitAll, a geometry- and physics-informed deep learning framework that encodes any molecular system with arbitrary charges, spins, and environmental effects using electronic structure information. It utilizes spin-polarized orbital features from the underlying quantum mechanical method and combines them with SE(3)-equivariant graph neural networks. OrbitAll demonstrates superior performance and generalization in predicting charged, open-shell, and solvated molecules, and robustly extrapolates to molecules significantly larger than the training data. OrbitAll achieves chemical accuracy using 10 times fewer training data than competing AI models, with approximately $10^3$ - $10^4$ speedup compared to density functional theory. Trained on a chemically diverse dataset, OrbitAll performs robustly on challenging molecular systems, and outperforms the foundational machine-learned interatomic potential, UMA, for highly charged species, despite using 35 times less molecular data and a 50-times-smaller model. After learning solvent effects, it accurately predicts solvent-dependent reaction pathways at about 100 times lower cost than explicit-solvation simulations using UMA.

cs.LG

Roadmap on Advancements of the FHI-aims Software Package

Electronic-structure theory is the foundation of the description of materials including multiscale modeling of their properties and functions. Obviously, without sufficient accuracy at the base, reliable predictions are unlikely at any level that follows. The software package FHI-aims has proven to be a game changer for accurate free-energy calculations because of its scalability, numerical precision, and its efficient handling of density functional theory (DFT) with hybrid functionals and van der Waals interactions. It treats molecules, clusters, and extended systems (solids and liquids) on an equal footing. Besides DFT, FHI-aims also includes quantum-chemistry methods, descriptions for excited states and vibrations, and calculations of various types of transport. Recent advancements address the integration of FHI-aims into an increasing number of workflows and various artificial intelligence (AI) methods. This Roadmap describes the state-of-the-art of FHI-aims and advancements that are currently ongoing or planned.

cond-mat.mtrl-sci

Non-linear and non-empirical double hybrid density functional

We develop a non-linear and non-empirical (nlane) double hybrid density functional derived from an accurate interpolation of the adiabatic connection in density functional theory, incorporating the correct asymptotic expansions. By bridging the second-order perturbative weak correlation limit with the fully interacting limit from the semi-local SCAN functional, nlane-SCAN is free of fitted parameters while providing improved energetic predictions compared to SCAN for moderately and strongly correlated systems alike. It delivers accurate predictions for atomic total energies and multiple reaction datasets from the GMTKN55 benchmark while significantly outperforming traditional linear hybrids and double hybrids for non-covalent interactions without requiring dispersion corrections. Due to the exact constraints at the weak correlation limit, nlane-SCAN has reduced delocalization errors as evident through the SIE4x4 benchmark and bond dissociations of H\(_2^+\) and He\(_2^+\). Its proper asymptotic behavior ensures stability in strongly correlated systems, improving H\(_2\) and N\(_2\) bond dissociation profiles compared to conventional functionals.

physics.chem-ph

Quantum mechanical dataset of 836k neutral closed shell molecules with upto 5 heavy atoms from CNOFSiPSClBr

We introduce the Vector-QM24 (VQM24) dataset comprehensively covering all possible neutral closed-shell small organic and inorganic molecules with up to five heavy (\textit{p}-block) atoms: C, N, O, F, Si, P, S, Cl, Br. All valid stoichiometries, Lewis-rule-consistent graphs, and stable conformers (identified via GFN2-xTB) were enumerated combinatorially, yielding 577k conformational isomers spanning 258k constitutional isomers and 5,599 unique stoichiometries. DFT ($ω$B97X-D3/cc-pVDZ) optimizations were performed for all, and diffusion quantum Monte Carlo (DMC@PBE0(ccECP/cc-pVQZ)) energies are provided for 10,793 lowest-energy conformers with up to 4 heavy atoms. VQM24 includes structures, vibrational modes, rotational constants, thermodynamic properties (Gibbs free energies, enthalpies, ZPVEs, entropies, heat capacities), and electronic properties such as atomization, electron interaction, exchange-correlation, dispersion energies, multipole moments (dipole to hexadecapole), alchemical potentials, Mulliken charges, and wavefunctions. Machine learning models of atomization energies on this dataset reveal significantly higher complexity than QM9, with none achieving chemical accuracy. VQM24 offers a rigorous, high-fidelity benchmark for evaluating quantum machine learning models.

physics.chem-ph

Alchemical harmonic approximation based potential for iso-electronic diatomics: Foundational baseline for $Δ$-machine learning

We introduce the alchemical harmonic approximation (AHA) of the absolute electronic energy for charge-neutral iso-electronic diatomics at fixed interatomic distance $d_0$. To account for variations in distance, we combine AHA with this Ansatz for the electronic binding potential, $E(d)=(E_{u}-E_s) \left(\frac{E_c-E_s}{E_u-E_s}\right)^{\sqrt{d/d_0}}+E_s$, where $E_u,E_c,E_s$ correspond to the energies of united atom, calibration at $d_0$, and sum of infinitely separated atoms, respectively. Our model covers the entire two-dimensional electronic potential energy surface spanned by distance and difference in nuclear charge from which only one single point (with elements of nuclear charge $Z_1,Z_2$ and distance $d_0$) is drawn to calibrate $E_c$. Using reference data from pbe0/cc-pVDZ, we present numerical evidence for the electronic ground-state of all neutral diatomics with 8, 10, 12, 14 electrons. We assess the validity of our model by comparison to legacy interatomic potentials (Harmonic oscillator, Lennard-Jones, and Morse) within the most relevant range of binding (0.7 - 2.5 A), and find comparable accuracy if restricted to single diatomics, and significantly better predictive power when extrapolating to the entire iso-electronic series. We also investigated $Δ$-learning of the electronic absolute energy using our model as baseline. This baseline model results in a systematic improvement, effectively reducing training data needs for reaching chemical accuracy by up to an order of magnitude from $\sim$1000 to $\sim$100. By contrast, using AHA+Morse as a baseline hardly leads to any improvement, and sometimes even deteriorates the predictive power. Inferring the energy of unseen CO converges to a prediction error of $\sim$0.1 Ha in direct learning, and $\sim$0.04 Ha with our baseline.

physics.chem-ph

Generalized convolutional many body distribution functional representations

Modern machine learning (ML) models of chemical and materials systems with billions of parameters require vast training datasets and considerable computational efforts. Lightweight kernel or decision tree based methods, however, can be rapidly trained, leading to a considerably lower carbon footprint. We introduce generalized convolutional many-body distribution functionals (cMBDF) as highly compute and data efficient atomic representations for accurate kernels that excel in low-data regimes. Generalizing the MBDF framework, cMBDF encodes local chemical environments in a compact fashion using translationally and rotationally invariant functionals of smooth atom centered Gaussian electron density proxy distributions weighted by interaction potentials. The functional values can be efficiently evaluated by expressing them in terms of convolutions which are calculated via fast Fourier transforms and stored on pre-defined grids. In the generalized form each atomic environment is described using a set of functionals uniformly defined by three integers; many-body, derivative, weighting orders. Irrespective of size/composition, cMBDF atomic vectors remain compact and constant in size for a fixed choice of these orders controlling the structural and compositional resolution. While being up to two orders of magnitude more compact than other popular representations, cMBDF is shown to be more accurate for the learning of various quantum properties such as energies, dipole moments, homo-lumo gaps, heat-capacity, polarizability, optimal exact-exchange admixtures and basis-set scaling factors. Applicability for organic and inorganic chemistry is tested as represented by the QM7b, QM9 and VQM24 data sets. Due to its compactness, model training and testing times are reduced from 23 hours to 8 minutes, implying a corresponding reduction in carbon footprint.

physics.chem-ph

Adaptive atomic basis sets

Atomic basis sets are widely employed within quantum mechanics based simulations of matter. We introduce a machine learning model that adapts the basis set to the local chemical environment of each atom, prior to the start of self consistent field (SCF) calculations. In particular, as a proof of principle and because of their historic popularity, we have studied the Gaussian type orbitals from the Pople basis set, i.e. the STO-3G, 3-21G, 6-31G and 6-31G*. We adapt the basis by scaling the variance of the radial Gaussian functions leading to contraction or expansion of the atomic orbitals.A data set of optimal scaling factors for C, H, O, N and F were obtained by variational minimization of the Hartree-Fock (HF) energy of the smallest 2500 organic molecules from the QM9 database. Kernel ridge regression based machine learning (ML) prediction errors of the change in scaling decay rapidly with training set size, typically reaching less than 1 % for training set size 2000. Overall, we find systematically lower variance, and consequently the larger training efficiencies, when going from hydrogen to carbon to nitrogen to oxygen. Using the scaled basis functions obtained from the ML model, we conducted HF calculations for the subsequent 30'000 molecules in QM9. In comparison to the corresponding default Pople basis set results we observed improved energetics in up to 99 % of all cases. With respect to the larger basis set 6-311G(2df,2pd), atomization energy errors are lowered on average by ~31, 107, 11, and 11 kcal/mol for STO-3G, 3-21G, 6-31G and 6-31G*, respectively -- with negligible computational overhead. We illustrate the high transferability of adaptive basis sets for larger out-of-domain molecules relevant to addiction, diabetes, pain, aging.

physics.chem-ph

Adaptive hybrid density functionals

Exact exchange contributions are known to crucially affect electronic states, which in turn govern covalent bond formation and breaking in chemical species. Empirically averaging the exact exchange admixture over compositional degrees of freedom, hybrid density functional approximations have been widely successful, yet have fallen short to reach high level quantum chemistry accuracy, primarily due to delocalization errors. We propose to `adaptify` hybrid functionals by generating optimal admixture ratios of exact exchange on the fly, i.e. specifically for any chemical compound, using extremely data efficient quantum machine learning models that carry negligible overhead. The adaptive Perdew-Burke-Ernzerhof based hybrid density functional (aPBE0) is shown to yield atomization energies with sufficient accuracy to effectively cure the infamous spin gap problem in open shell systems, such as carbenes. aPBE0 further improves energetics, electron densities, and HOMO-LUMO gaps in organic molecules drawn from the QM9 and QM7b data set. Obtained with aPBE0 in a large basis, we present a revision of the entire QM9 data set (revQM9) with an estimated quality vastly superior to the original containing on average, stronger covalent binding, larger band-gaps, more localized electron densities, and larger dipole-moments. While aPBE0 is applicable in the equilibrium regime, outstanding limitations include covalent bond dissociation when going beyond the Coulson-Fisher point.

physics.chem-ph

Reducing Training Data Needs with Minimal Multilevel Machine Learning (M3L)

For many machine learning applications in science, data acquisition, not training, is the bottleneck even when avoiding experiments and relying on computation and simulation. Correspondingly, and in order to reduce cost and carbon footprint, training data efficiency is key. We introduce minimal multilevel machine learning (M3L) which optimizes training data set sizes using a loss function at multiple levels of reference data in order to minimize a combination of prediction error with overall training data acquisition costs (as measured by computational wall-times). Numerical evidence has been obtained for calculated atomization energies and electron affinities of thousands of organic molecules at various levels of theory including HF, MP2, DLPNO-CCSD(T), DFHFCABS, PNOMP2F12, and PNOCCSD(T)F12, and treating tens with basis sets TZ, cc-pVTZ, and AVTZ-F12. Our M3L benchmarks for reaching chemical accuracy in distinct chemical compound sub-spaces indicate substantial computational cost reductions by factors of $\sim$ 1.01, 1.1, 3.8, 13.8 and 25.8 when compared to heuristic sub-optimal multilevel machine learning (M2L) for the data sets QM7b, QM9$^\mathrm{LCCSD(T)}$, EGP, QM9$^\mathrm{CCSD(T)}_\mathrm{AE}$, and QM9$^\mathrm{CCSD(T)}_\mathrm{EA}$, respectively. Furthermore, we use M2L to investigate the performance for 76 density functionals when used within multilevel learning and building on the following levels drawn from the hierarchy of Jacobs Ladder:~LDA, GGA, mGGA, and hybrid functionals. Within M2L and the molecules considered, mGGAs do not provide any noticeable advantage over GGAs. Among the functionals considered and in combination with LDA, the three on average top performing GGA and Hybrid levels for atomization energies on QM9 using M3L correspond respectively to PW91, KT2, B97D, and $τ$-HCTH, B3LYP$\ast$(VWN5), TPSSH.

physics.chem-ph

Autonomous data extraction from peer reviewed literature for training machine learning models of oxidation potentials

We present an automated data-collection pipeline involving a convolutional neural network and a large language model to extract user-specified tabular data from peer-reviewed literature. The pipeline is applied to 74 reports published between 1957 and 2014 with experimentally-measured oxidation potentials for 592 organic molecules (-0.75 to 3.58 V). After data curation (solvents, reference electrodes, and missed data points), we trained multiple supervised machine learning models reaching prediction errors similar to experimental uncertainty ($\sim$0.2 V). For experimental measurements of identical molecules reported in multiple studies, we identified the most likely value based on out-of-sample machine learning predictions. Using the trained machine learning models, we then estimated oxidation potentials of $\sim$132k small organic molecules from the QM9 data set, with predicted values spanning 0.21 to 3.46 V. Analysis of the QM9 predictions in terms of plausible descriptor-property trends suggests that aliphaticity increases the oxidation potential of an organic molecule on average from $\sim$1.5 V to $\sim$2 V, while an increase in number of heavy atoms lowers it systematically. The pipeline introduced offers significant reductions in human labor otherwise required for conventional manual data collection of experimental results, and exemplifies how to accelerate scientific research through automation.

physics.chem-ph

Kernel based quantum machine learning at record rate : Many-body distribution functionals as compact representations

The feature vector mapping used to represent chemical systems is a key factor governing the superior data-efficiency of kernel based quantum machine learning (QML) models applicable throughout chemical compound space. Unfortunately, the most accurate representations require a high dimensional feature mapping, thereby imposing a considerable computational burden on model training and use. We introduce compact yet accurate, linear scaling QML representations based on atomic Gaussian many-body distribution functionals (MBDF), and their derivatives. Weighted density functions (DF) of MBDF values are used as global representations which are constant in size, i.e.~invariant with respect to the number of atoms. We report predictive performance and training data efficiency that is competitive with state of the art for two diverse datasets of organic molecules, QM9 and QMugs. Generalization capability has been investigated for atomization energies, HOMO-LUMO eigenvalues and gap, internal energies at 0 K, zero point vibrational energies, dipole moment norm, static isotropic polarizability, and heat capacity as encoded in QM9. MBDF based QM9 performance lowers the optimal Pareto front spanned between sampling and training cost to compute node minutes,~effectively sampling chemical compound space with chemical accuracy at a sampling rate of $\sim 48$ molecules per core second.

physics.chem-ph

Taking sides: Public Opinion over the Israel-Palestine Conflict in 2021

The Israel-Palestine Conflict, one of the most enduring conflicts in history, dates back to the start of 20th century, with the establishment of the British Mandate in Palestine and has deeply rooted complex issues in politics, demography, religion, and other aspects, making it harder to attain resolve. To understand the conflict in 2021, we devise an observational study to aggregate stance held by English-speaking countries. We collect Twitter data using popular hashtags around and specific to the conflict portraying opinions neutral or partial to the two parties. We use different tools and methods to classify tweets into pro-Palestinian, pro-Israel, or neutral. This paper further describes the implementation of data mining methodologies to obtain insights and reason the stance held by citizens around the conflict.

cs.SI