SearcharxivSearch

arXiv subjects

Clemence Corminboeuf

Publications and source records attributed to Clemence Corminboeuf.

12 recordsLinked to original sources

Integer linear programming for unsupervised training set selection in molecular machine learning

Integer linear programming (ILP) is an elegant approach to solve linear optimization problems, naturally described using integer decision variables. Within the context of physics-inspired machine learning applied to chemistry, we demonstrate the relevance of an ILP formulation to select molecular training sets for predictions of size-extensive properties. We show that our algorithm outperforms existing unsupervised training set selection approaches, especially when predicting properties of molecules larger than those present in the training set. We argue that the reason for the improved performance is due to the selection that is based on the notion of local similarity (i.e., per-atom) and a unique ILP approach that finds optimal solutions efficiently. Altogether, this work provides a practical algorithm to improve the performance of physics-inspired machine learning models and offers insights into the conceptual differences with existing training set selection approaches.

physics.chem-ph

3DReact: Geometric deep learning for chemical reactions

Geometric deep learning models, which incorporate the relevant molecular symmetries within the neural network architecture, have considerably improved the accuracy and data efficiency of predictions of molecular properties. Building on this success, we introduce 3DReact, a geometric deep learning model to predict reaction properties from three-dimensional structures of reactants and products. We demonstrate that the invariant version of the model is sufficient for existing reaction datasets. We illustrate its competitive performance on the prediction of activation barriers on the GDB7-22-TS, Cyclo-23-TS and Proparg-21-TS datasets in different atom-mapping regimes. We show that, compared to existing models for reaction property prediction, 3DReact offers a flexible framework that exploits atom-mapping information, if available, as well as geometries of reactants and products (in an invariant or equivariant fashion). Accordingly, it performs systematically well across different datasets, atom-mapping regimes, as well as both interpolation and extrapolation tasks.

physics.chem-ph

Range separation of the interaction potential in intermolecular and intramolecular symmetry-adapted perturbation theory

Symmetry-adapted perturbation theory (SAPT) is a popular and versatile tool to compute and decompose noncovalent interaction energies between molecules. The intramolecular SAPT (ISAPT) variant provides a similar energy decomposition between two nonbonded fragments of the same molecule, covalently connected by a third fragment. In this work, we explore an alternative approach where the noncovalent interaction is singled out by a range separation of the Coulomb potential. We investigate two common splittings of the $1/r$ potential into long-range and short-range parts based on the Gaussian and error functions, and approximate either the entire intermolecular/interfragment interaction or only its attractive terms by the long-range contribution. These range separation schemes are tested for a number of intermolecular and intramolecular complexes. We find that the energy corrections from range-separated SAPT or ISAPT are in reasonable agreement with complete SAPT/ISAPT data. This result should be contrasted with the inability of the long-range multipole expansion to describe crucial short-range charge penetration and exchange effects; it shows that the long-range interaction potential does not just recover the asymptotic interaction energy but also provides a useful account of short-range terms. The best consistency is attained for the error-function separation applied to all interaction terms, both attractive and repulsive. This study is the first step towards a fragmentation-free decomposition of intramolecular nonbonded energy.

physics.chem-ph

SPA$^\mathrm{H}$M(a,b): encoding the density information from guess Hamiltonian in quantum machine learning representations

Recently, we introduced a class of molecular representations for kernel-based regression methods -- the spectrum of approximated Hamiltonian matrices (SPA$^\mathrm{H}$M) -- that takes advantage of lightweight one-electron Hamiltonians traditionally used as an SCF initial guess. The original SPA$^\mathrm{H}$M variant is built from occupied-orbital energies (ie, eigenvalues) and naturally contains all the information about nuclear charges, atomic positions, and symmetry requirements. Its advantages were demonstrated on datasets featuring a wide variation of charge and spin, for which traditional structure-based representations commonly fail. SPA$^\mathrm{H}$M(a,b), as introduced here, expand the eigenvalue SPA$^\mathrm{H}$M into local and transferable representations. They rely upon one-electron density matrices to build fingerprints from atomic and bond density overlap contributions inspired from preceding state-of-the-art representations. The performance and efficiency of SPA$^\mathrm{H}$M(a,b) is assessed on the predictions for datasets of prototypical organic molecules (QM7) of different charges and azoheteroarene dyes in an excited state. Overall, both SPA$^\mathrm{H}$M(a) and SPA$^\mathrm{H}$M(b) outperform state-of-the-art representations on difficult prediction tasks such as the atomic properties of charged open-shell species and of $π$-conjugated systems.

physics.chem-ph

SPA$^\mathrm{H}$M: the Spectrum of Approximated Hamiltonian Matrices representations

Physics-inspired molecular representations are the cornerstone of similarity-based learning applied to solve chemical problems. Despite their conceptual and mathematical diversity, this class of descriptors shares a common underlying philosophy: they all rely on the molecular information that determines the form of the electronic Schrödinger equation. Existing representations take the most varied forms, from non-linear functions of atom types and positions to atom densities and potential, up to complex quantum chemical objects directly injected into the ML architecture. In this work, we present the Spectrum of Approximated Hamiltonian Matrices (SPA$^\mathrm{H}$M) as an alternative pathway to construct quantum machine learning representations through leveraging the foundation of the electronic Schrödinger equation itself: the electronic Hamiltonian. As the Hamiltonian encodes all quantum chemical information at once, SPA$^\mathrm{H}$M representations not only distinguish different molecules and conformations, but also different spin, charge, and electronic states. As a proof of concept, we focus here on efficient SPA$^\mathrm{H}$M representations built from the eigenvalues of a hierarchy of well-established and readily-evaluated "guess" Hamiltonians. These SPA$^\mathrm{H}$M representations are particularly compact and efficient for kernel evaluation and their complexity is independent of the number of different atom types in the database.

physics.chem-ph

Assessing the persistence of chalcogen bonds in solution with neural network potentials

Non-covalent bonding patterns are commonly harvested as a design principle in the field of catalysis, supramolecular chemistry and functional materials to name a few. Yet, their computational description generally neglects finite temperature and environment effects, which promote competing interactions and alter their static gas-phase properties. Recently, neural network potentials (NNPs) trained on Density Functional Theory (DFT) data have become increasingly popular to simulate molecular phenomena in condensed phase with an accuracy comparable to ab initio methods. To date, most applications have centered on solid-state materials or fairly simple molecules made of a limited number of elements. Herein, we focus on the persistence and strength of chalcogen bonds involving a benzotelluradiazole in condensed phase. While the tellurium-containing heteroaromatic molecules are known to exhibit pronounced interactions with anions and lone pairs of different atoms, the relevance of competing intermolecular interactions, notably with the solvent, is complicated to monitor experimentally but also challenging to model at an accurate electronic structure level. Here, we train direct and baselined NNPs to reproduce hybrid DFT energies and forces in order to identify what are the most prevalent non-covalent interactions occurring in a solute-Cl$^-$-THF mixture. The simulations in explicit solvent highlight the clear competition with chalcogen bonds formed with the solvent and the short-range directionality of the interaction with direct consequences for the molecular properties in the solution. The comparison with other potentials (e.g., AMOEBA, direct NNP and continuum solvent model) also demonstrates that baselined NNPs offer a reliable picture of the non-covalent interaction interplay occurring in solution.

physics.chem-ph

Impact of quantum-chemical metrics on the machine learning prediction of electron density

Machine learning (ML) algorithms have undergone an explosive development impacting every aspect of computational chemistry. To obtain reliable predictions, one needs to maintain the proper balance between the black-box nature of ML frameworks and the physics of the target properties. One of the most appealing quantum-chemical properties for regression models is the electron density, and some of us recently proposed a transferable and scalable model based on the decomposition of the density onto an atom-centered basis set. The decomposition, as well as the training of the model, is at its core a minimization of some loss function, which can be arbitrarily chosen and may lead to results of different quality. Well-studied in the context of density fitting (DF), the impact of the metric on the performance of ML models has not been analyzed yet. In this work, we compare predictions obtained using the overlap and the Coulomb-repulsion metrics for both decomposition and training. As expected, the Coulomb metric used as both the DF and ML loss functions leads to the best results for the electrostatic potential and dipole moments. The origin of this difference lies in the fact that the model is not constrained to predict densities that integrate to the exact number of electrons $N$. Since an \textit{a posteriori} correction for the number of electrons decreases the errors, we proposed a modification of the model where $N$ is included directly into the kernel function, which allowed to lower the errors on the test and out-of-sample sets.

physics.chem-ph

Learning on-top: regressing the on-top pair density for real-space visualization of electron correlation

The on-top pair density [$Π(\mathrm{\mathbf{r}})$] is a local quantum-chemical property that reflects the probability of two electrons of any spin to occupy the same position in space. Being the simplest quantity related to the two-particle density matrix, the on-top pair density is a powerful indicator of electron correlation effects, and as such, it has been extensively used to combine density functional theory and multireference wavefunction theory. The widespread application of $Π(\mathrm{\mathbf{r}})$ is currently hindered by the need for post-Hartree--Fock or multireference computations for its accurate evaluation. In this work, we propose the construction of a machine learning model capable of predicting the CASSCF-quality on-top pair density of a molecule only from its structure and composition. Our model, trained on the GDB11-AD-3165 database, is able to predict with minimal error the on-top pair density of organic molecules, bypassing completely the need for $\textit{ab initio}$ computations. The accuracy of the regression is demonstrated using the on-top ratio as a visual metric of electron correlation effects and bond-breaking in real-space. In addition, we report the construction of a specialized basis set, built to fit the on-top pair density in a single atom-centered expansion. This basis, cornerstone of the regression, could be potentially used also in the same spirit of the resolution-of-the-identity approximation for the electron density.

physics.chem-ph

Simulating solvation and acidity in complex mixtures with first-principles accuracy: the case of CH$_3$SO$_3$H and H$_2$O$_2$ in phenol

We present a generally-applicable computational framework for the efficient and accurate characterization of molecular structural patterns and acid properties in explicit solvent using H$_2$O$_2$ and CH$_3$SO$_3$H in phenol as an example. In order to address the challenges posed by the complexity of the problem, we resort to a set of data-driven methods and enhanced sampling algorithms. The synergistic application of these techniques makes the first-principle estimation of the chemical properties feasible without renouncing to the use of explicit solvation, involving extensive statistical sampling. Ensembles of neural network potentials are trained on a set of configurations carefully selected out of preliminary simulations performed at a low-cost density-functional tight-binding (DFTB) level. The energy and forces of these configurations are then recomputed at the hybrid density functional theory (DFT) level and used to train the neural networks. The stability of the NN model is enhanced by using DFTB energetics as a baseline, but the efficiency of the direct NN (i.e., baseline-free) is exploited via a multiple-time step integrator. The neural network potentials are combined with enhanced sampling techniques, such as replica exchange and metadynamics, and used to characterize the relevant protonated species and dominant non-covalent interactions in the mixture, also considering nuclear quantum effects.

physics.chem-ph

Learning the energy curvature versus particle number in approximate density functionals

The average energy curvature as a function of the particle number is a molecule-specific quantity, which measures the deviation of a given functional from the exact conditions of density functional theory (DFT). Related to the lack of derivative discontinuity in approximate exchange-correlation potentials, the information about the curvature has been successfully used to restore the physical meaning of Kohn-Sham orbital eigenvalues and to develop non-empirical tuning and correction schemes for density functional approximations. In this work, we propose the construction of a machine-learning framework targeting the average energy curvature between the neutral and the radical cation state of thousands of small organic molecules (QM7 database). The applicability of the model is demonstrated in the context of system-specific gamma-tuning of the LC-$ω$PBE functional and validated against the molecular first ionization potentials at equation-of-motion (EOM) coupled-cluster references. In addition, we propose a local version of the non-linear regression model and demonstrate its transferability and predictive power by determining the optimal range-separation parameter for two large molecules relevant to the field of hole-transporting materials. Finally, we explore the underlying structure of the QM7 database with the t-SNE dimensionality-reduction algorithm and identify structural and compositional patterns that promote the deviation from the piecewise linearity condition.

physics.chem-ph

i-PI 2.0: A Universal Force Engine for Advanced Molecular Simulations

Progress in the atomic-scale modelling of matter over the past decade has been tremendous. This progress has been brought about by improvements in methods for evaluating interatomic forces that work by either solving the electronic structure problem explicitly, or by computing accurate approximations of the solution and by the development of techniques that use the Born-Oppenheimer (BO) forces to move the atoms on the BO potential energy surface. As a consequence of these developments it is now possible to identify stable or metastable states, to sample configurations consistent with the appropriate thermodynamic ensemble, and to estimate the kinetics of reactions and phase transitions. All too often, however, progress is slowed down by the bottleneck associated with implementing new optimization algorithms and/or sampling techniques into the many existing electronic-structure and empirical-potential codes. To address this problem, we are thus releasing a new version of the i-PI software. This piece of software is an easily extensible framework for implementing advanced atomistic simulation techniques using interatomic potentials and forces calculated by an external driver code. While the original version of the code was developed with a focus on path integral molecular dynamics techniques, this second release of i-PI not only includes several new advanced path integral methods, but also offers other classes of algorithms. In other words, i-PI is moving towards becoming a universal force engine that is both modular and tightly coupled to the driver codes that evaluate the potential energy surface and its derivatives.

physics.chem-ph

A Transferable Machine-Learning Model of the Electron Density

The electronic charge density plays a central role in determining the behavior of matter at the atomic scale, but its computational evaluation requires demanding electronic-structure calculations. We introduce an atom-centered, symmetry-adapted framework to machine-learn the valence charge density based on a small number of reference calculations. The model is highly transferable, meaning it can be trained on electronic-structure data of small molecules and used to predict the charge density of larger compounds with low, linear-scaling cost. Applications are shown for various hydrocarbon molecules of increasing complexity and flexibility, and demonstrate the accuracy of the model when predicting the density on octane and octatetraene after training exclusively on butane and butadiene. This transferable, data-driven model can be used to interpret experiments, initialize electronic structure calculations, and compute electrostatic interactions in molecules and condensed-phase systems.

physics.chem-ph