Searcharxiv⌕ Search

arXiv subjects

Konstantin Karandashev

Publications and source records attributed to Konstantin Karandashev.

10 recordsLinked to original sources

Calculated state-of-the art results for solvation and ionization energies of thousands of organic molecules relevant to battery design

We present high-quality reference data for two fundamentally important groups of molecular properties related to a compound's utility as a lithium battery electrolyte. The first property is energy changes associated with charge excitations of molecules, namely ionization potential and electron affinity. They were estimated for 7000 randomly chosen molecules with up to 9 non-hydrogen atoms C, N, O, and F (QM9 dataset) using the DH-HF, DF-HF-CABS, PNO-LMP2-F12, and PNO-LCCSD(T)-F12 methods as implemented in the Molpro software, and the aug-cc-pVTZ basis set. Additionally, we provide the corresponding atomization energies at these levels of theory, as well as the CPU time and disk space used during the calculations. The second property is solvation energies for 39 different solvents, which we estimate for 18361 molecules connected to battery design (Electrolyte Genome Project dataset), 309463 randomly chosen molecules with up to 17 non-hydrogen atoms C, N, O, S, and halogens (GDB17 dataset), as well as 88418 atoms-in-molecules of the ZINC database of commercially available compounds and 37772 atoms-in-molecules of GDB17. For these calculations we used the COnductor-like Screening MOdel for Real Solvents (COSMO-RS) method; we additionally provide estimates of gas-phase atomization energies, as well as information about conformers considered during the COSMO-RS calculations, namely coordinates, energies, and dipole moments.

physics.chem-ph↗

Kernel Ridge Regression for conformer ensembles made easy with Structured Orthogonal Random Features

A computationally efficient protocol for machine learning in chemical space using Boltzmann ensembles of conformers as input is proposed; the method is based on rewriting Kernel Ridge Regression expressions in terms of Structured Orthogonal Random Features, yielding physics-motivated trigonometric neural networks. To evaluate the method's utility for materials discovery, we test it on experimental datasets of two quantities related to battery electrolyte design, namely oxidation potentials in acetonitrile and hydration energies, using several popular molecular representations to demonstrate the method's flexibility. Despite only using computationally cheap forcefield calculations for conformer generation, we observe systematic decrease of machine learning error with increased training set size in all cases, with experimental accuracy reached after training on hundreds of molecules and prediction errors being comparable to state-of-the-art machine learning approaches. We also present novel versions of Huber and LogCosh loss functions that made hyperparameter optimization of the new approach more convenient.

physics.chem-ph↗

Exact sampling of molecules in chemical space

The concept of molecular similarity appears in many machine-learning algorithms based on the assumption that molecules with similar representations will also share similar properties. In this work, we propose a new way to study similarity measures in molecular graph space using a Monte Carlo approach. We enable direct sampling from the underlying distribution of chemical space without numerical approximations or complete enumeration of molecular graphs, the latter intractable for practically relevant graph sets of interest. The Monte Carlo method allows observation of several interesting fundamental properties of chemical space, such as a linear trend of average property derivatives in chemical space with respect to the property's value at the molecule of interest. The trend was observed for extensive and intensive properties, suggesting that this trend is an inherent property of chemical space.

physics.chem-ph↗

Evolutionary Monte Carlo of QM properties in chemical space: Electrolyte design

Optimizing a target function over the space of organic molecules is an important problem appearing in many fields of applied science, but also a very difficult one due to the vast number of possible molecular systems. We propose an Evolutionary Monte Carlo algorithm for solving such problems which is capable of straightforwardly tuning both exploration and exploitation characteristics of an optimization procedure while retaining favourable properties of genetic algorithms. The method, dubbed MOSAiCS (Metropolis Optimization by Sampling Adaptively in Chemical Space), is tested on problems related to optimizing components of battery electrolytes, namely minimizing solvation energy in water or maximizing dipole moment while enforcing a lower bound on the HOMO-LUMO gap; optimization was done over sets of molecular graphs inspired by QM9 and Electrolyte Genome Project (EGP) datasets. MOSAiCS reliably generated molecular candidates with good target quantity values, which were in most cases better than the ones found in QM9 or EGP. While the optimization results presented in this work sometimes required up to $10^{6}$ QM calculations and were thus only feasible thanks to computationally efficient ab initio approximations of properties of interest, we discuss possible strategies for accelerating MOSAiCS using machine learning approaches.

physics.chem-ph↗

Reducing Training Data Needs with Minimal Multilevel Machine Learning (M3L)

For many machine learning applications in science, data acquisition, not training, is the bottleneck even when avoiding experiments and relying on computation and simulation. Correspondingly, and in order to reduce cost and carbon footprint, training data efficiency is key. We introduce minimal multilevel machine learning (M3L) which optimizes training data set sizes using a loss function at multiple levels of reference data in order to minimize a combination of prediction error with overall training data acquisition costs (as measured by computational wall-times). Numerical evidence has been obtained for calculated atomization energies and electron affinities of thousands of organic molecules at various levels of theory including HF, MP2, DLPNO-CCSD(T), DFHFCABS, PNOMP2F12, and PNOCCSD(T)F12, and treating tens with basis sets TZ, cc-pVTZ, and AVTZ-F12. Our M3L benchmarks for reaching chemical accuracy in distinct chemical compound sub-spaces indicate substantial computational cost reductions by factors of $\sim$ 1.01, 1.1, 3.8, 13.8 and 25.8 when compared to heuristic sub-optimal multilevel machine learning (M2L) for the data sets QM7b, QM9$^\mathrm{LCCSD(T)}$, EGP, QM9$^\mathrm{CCSD(T)}_\mathrm{AE}$, and QM9$^\mathrm{CCSD(T)}_\mathrm{EA}$, respectively. Furthermore, we use M2L to investigate the performance for 76 density functionals when used within multilevel learning and building on the following levels drawn from the hierarchy of Jacobs Ladder:~LDA, GGA, mGGA, and hybrid functionals. Within M2L and the molecules considered, mGGAs do not provide any noticeable advantage over GGAs. Among the functionals considered and in combination with LDA, the three on average top performing GGA and Hybrid levels for atomization energies on QM9 using M3L correspond respectively to PW91, KT2, B97D, and $τ$-HCTH, B3LYP$\ast$(VWN5), TPSSH.

physics.chem-ph↗

An orbital-based representation for accurate Quantum Machine Learning

We introduce an electronic structure based representation for quantum machine learning (QML) of electronic properties throughout chemical compound space. The representation is constructed using computationally inexpensive ab initio calculations and explicitly accounts for changes in the electronic structure. We demonstrate the accuracy and flexibility of resulting QML models when applied to property labels such as total potential energy, HOMO and LUMO energies, ionization potential, and electron affinity, using as data sets for training and testing entries from the QM7b, QM7b-T, QM9, and LIBE libraries. For the latter, we also demonstrate the ability of this approach to account for molecular species of different charge and spin multiplicity, resulting in QML models that infer total potential energies based on geometry, charge, and spin as input.

physics.chem-ph↗

A combined on-the-fly/interpolation procedure for evaluating energy values needed in molecular simulations

We propose an algorithm for molecular dynamics or Monte Carlo simulations that uses an interpolation procedure to estimate potential energy values from energies and gradients evaluated previously at points of a simplicial mesh. We chose an interpolation procedure which is exact for harmonic systems and considered two possible mesh types: Delaunay triangulation and an alternative anisotropic triangulation designed to improve performance in anharmonic systems. The mesh is generated and updated on the fly during the simulation. The procedure is tested on two-dimensional quartic oscillators and on the path integral Monte Carlo evaluation of HCN/DCN equilibrium isotope effect.

physics.comp-ph↗

Accelerating equilibrium isotope effect calculations: II. Stochastic implementation of direct estimators

Path integral calculations of equilibrium isotope effects and isotopic fractionation are expensive due to the presence of path integral discretization errors, statistical errors, and thermodynamic integration errors. Whereas the discretization errors can be reduced by high-order factorization of the path integral and statistical errors by using centroid virial estimators, two recent papers proposed alternative ways to completely remove the thermodynamic integration errors: Cheng and Ceriotti [J. Chem. Phys. 141, 244112 (2015)] employed a variant of free-energy perturbation called "direct estimators," while Karandashev and Van\'ıček [J. Chem. Phys. 143, 194104 (2017)] combined the thermodynamic integration with a stochastic change of mass and piecewise-linear umbrella biasing potential. Here we combine the former approach with the stochastic change of mass in order to decrease its statistical errors when applied to larger isotope effects, and perform a thorough comparison of different methods by computing isotope effects first on a harmonic model, and then on methane and methanium, where we evaluate all isotope effects of the form $\mathrm{CH}_{\mathrm{4-x}}\mathrm{D}_{\mathrm{x}}/\mathrm{CH}_{4}$ and $\mathrm{CH}_{\mathrm{5-x}}\mathrm{D}^{+}_{\mathrm{x}}/\mathrm{CH}^{+}_{5}$, respectively. We discuss thoroughly the reasons for a surprising behavior of the original method of direct estimators, which performed well for a much larger range of isotope effects than what had been expected previously.

physics.chem-ph↗

Accelerating quantum instanton calculations of the kinetic isotope effects

Path integral implementation of the quantum instanton approximation currently belongs among the most accurate methods for computing quantum rate constants and kinetic isotope effects, but its use has been limited due to the rather high computational cost. Here we demonstrate that the efficiency of quantum instanton calculations of the kinetic isotope effects can be increased by orders of magnitude by combining two approaches: The convergence to the quantum limit is accelerated by employing high-order path integral factorizations of the Boltzmann operator, while the statistical convergence is improved by implementing virial estimators for relevant quantities. After deriving several new virial estimators for the high-order factorization and evaluating the resulting increase in efficiency, using $\mathrm{\cdot H_α+H_βH_γ\rightarrow H_αH_β+\cdot H_{γ}}$ reaction as an example, we apply the proposed method to obtain several kinetic isotope effects on $\mathrm{CH_{4}+\cdot H\rightleftharpoons\cdot CH_{3}+H_{2}}$ forward and backward reactions.

physics.chem-ph↗

Accelerating equilibrium isotope effect calculations: I. Stochastic thermodynamic integration with respect to mass

Accurate path integral Monte Carlo or molecular dynamics calculations of isotope effects have until recently been expensive because of the necessity to reduce three types of errors present in such calculations: statistical errors due to sampling, path integral discretization errors, and thermodynamic integration errors. While the statistical errors can be reduced with virial estimators and path integral discretization errors with high-order factorization of the Boltzmann operator, here we propose a method for accelerating isotope effect calculations by eliminating the integration error. We show that the integration error can be removed entirely by changing particle masses stochastically during the calculation and by using a piecewise linear umbrella biasing potential. Moreover, we demonstrate numerically that this approach does not increase the statistical error. The resulting acceleration of isotope effect calculations is demonstrated on a model harmonic system and on deuterated species of methane.

physics.chem-ph↗