SearcharxivSearch

arXiv subjects

Sandip De

Publications and source records attributed to Sandip De.

15 recordsLinked to original sources

False Metallization in Short-Ranged Machine Learned Interatomic Potentials

Machine learned interatomic potentials (MLIPs) have enabled atomistic simulations with ab initio accuracy for a fraction of the computational cost. However, many widely used MLIPs are short-ranged and do not accurately capture long-ranged electrostatic interactions. At interfaces with polar solvents, such as water, this deficiency can drive unphysical long-distance dipolar alignment far away from the interface. Here we reveal that neglecting long-ranged physics leads to spurious metallization of the water layer due to artificially large fluctuations of the total solvent dipole, similar to the electron rearrangement observed to prevent polar catastrophes at polar interfaces. This metallization is eliminated in MLIPs that explicitly include long-ranged electrostatics. Our results showcase a fundamental flaw of short-ranged MLIPs, highlighting that long-ranged electrostatics are essential for studying systems with a polar-liquid component, especially if one is interested in electronic properties.

physics.chem-ph

High-quality, high-information datasets for universal atomistic machine learning

The quality, consistency, and information content of training data is often what determines the practical value of machine-learning models for atomistic simulations. Yet, many widely used electronic-structure databases are assembled having materials screening as primary goal rather than robust force-field learning, are limited in their scope to a specific class of chemical compounds, and/or employ inconsistent DFT functionals and settings. Here we introduce MAD-1.6, a highly curated dataset designed explicitly for training broadly applicable atomistic models across the periodic table at high levels of theory. MAD-1.6 extends the MAD dataset with targeted enrichment strategies that improve the coverage of chemical space to 102 elements while keeping the total number of configurations compact. All structures are computed with a single, standardized all-electron DFT workflow using the r$^2$SCAN meta-GGA functional and consistent convergence settings, ensuring uniformity across chemically heterogeneous systems. The dataset encompasses molecules, clusters, bulk crystals, surfaces, and low-dimensional structures, and its quality and consistency are further enhanced by outlier removal using uncertainty quantification. We demonstrate the high accuracy that can be achieved with the proposed dataset by training PET-MAD-1.6, a generally applicable r$^2$SCAN interatomic potential that covers 102 elements in the periodic table and achieves exceptional levels of benchmark accuracy and stability in challenging simulation protocols.

cond-mat.mtrl-sci

Massive Atomic Diversity: a compact universal dataset for atomistic machine learning

The development of machine-learning models for atomic-scale simulations has benefited tremendously from the large databases of materials and molecular properties computed in the past two decades using electronic-structure calculations. More recently, these databases have made it possible to train universal models that aim at making accurate predictions for arbitrary atomic geometries and compositions. The construction of many of these databases was however in itself aimed at materials discovery, and therefore targeted primarily to sample stable, or at least plausible, structures and to make the most accurate predictions for each compound - e.g. adjusting the calculation details to the material at hand. Here we introduce a dataset designed specifically to train machine learning models that can provide reasonable predictions for arbitrary structures, and that therefore follows a different philosophy. Starting from relatively small sets of stable structures, the dataset is built to contain massive atomic diversity (MAD) by aggressively distorting these configurations, with near-complete disregard for the stability of the resulting configurations. The electronic structure details, on the other hand, are chosen to maximize consistency rather than to obtain the most accurate prediction for a given structure, or to minimize computational effort. The MAD dataset we present here, despite containing fewer than 100k structures, has already been shown to enable training universal interatomic potentials that are competitive with models trained on traditional datasets with two to three orders of magnitude more structures. We describe in detail the philosophy and details of the construction of the MAD dataset. We also introduce a low-dimensional structural latent space that allows us to compare it with other popular datasets and that can be used as a general-purpose materials cartography tool.

cond-mat.mtrl-sci

A foundation model for atomistic materials chemistry

Atomistic simulations of matter, especially those that leverage first-principles (ab initio) electronic structure theory, provide a microscopic view of the world, underpinning much of our understanding of chemistry and materials science. Over the last decade or so, machine-learned force fields have transformed atomistic modeling by enabling simulations of ab initio quality over unprecedented time and length scales. However, early ML force fields have largely been limited by: (i) the substantial computational and human effort of developing and validating potentials for each particular system of interest; and (ii) a general lack of transferability from one chemical system to the next. Here we show that it is possible to create a general-purpose atomistic ML model, trained on a public dataset of moderate size, that is capable of running stable molecular dynamics for a wide range of molecules and materials. We demonstrate the power of the MACE-MP-0 model - and its qualitative and at times quantitative accuracy - on a diverse set of problems in the physical sciences, including properties of solids, liquids, gases, chemical reactions, interfaces and even the dynamics of a small protein. The model can be applied out of the box as a starting or "foundation" model for any atomistic system of interest and, when desired, can be fine-tuned on just a handful of application-specific data points to reach ab initio accuracy. Establishing that a stable force-field model can cover almost all materials changes atomistic modeling in a fundamental way: experienced users get reliable results much faster, and beginners face a lower barrier to entry. Foundation models thus represent a step towards democratising the revolution in atomic-scale modeling that has been brought about by ML force fields.

physics.chem-ph

Surface segregation in high-entropy alloys from alchemical machine learning

High-entropy alloys (HEAs), containing several metallic elements in near-equimolar proportions, have long been of interest for their unique mechanical properties. More recently, they have emerged as a promising platform for the development of novel heterogeneous catalysts, because of the large design space, and the synergistic effects between their components. In this work we use a machine-learning potential that can model simultaneously up to 25 transition metals to study the tendency of different elements to segregate at the surface of a HEA. We use as a starting point a potential that was previously developed using exclusively crystalline bulk phases, and show that, thanks to the physically-inspired functional form of the model, adding a much smaller number of defective configurations makes it capable of describing surface phenomena. We then present several computational studies of surface segregation, including both a simulation of a 25-element alloy, that provides a rough estimate of the relative surface propensity of the various elements, and targeted studies of CoCrFeMnNi and IrFeCoNiCu, which provide further validation of the model, and insights to guide the modeling and design of alloys for heterogeneous catalysis.

cond-mat.mtrl-sci

Accurate Energy Barriers for Catalytic Reaction Pathways: An Automatic Training Protocol for Machine Learning Force Fields

In this study, we introduce a training protocol for developing machine learning force fields (MLFFs), capable of accurately determining energy barriers in catalytic reaction pathways. The protocol is validated on the extensively explored hydrogenation of carbon dioxide to methanol over indium oxide. With the help of active learning, the final force field obtains energy barriers within 0.05 eV of Density Functional Theory. Thanks to the computational speedup, not only do we reduce the cost of routine in-silico catalytic tasks, but also find a 40\% reduction in the previously established rate-limiting step. Furthermore, we illustrate the importance of finite-temperature effects and compute free energy barriers. The transferability of the protocol is demonstrated on the experimentally relevant, yet unexplored, top-layer reduced indium oxide surface. The ability of MLFFs to enhance our understanding of extensively studied catalysts underscores the need for fast and accurate alternatives to direct ab-intio simulations.

physics.chem-ph

Modeling high-entropy transition-metal alloys with alchemical compression

Alloys composed of several elements in roughly equimolar composition, often referred to as high-entropy alloys, have long been of interest for their thermodynamics and peculiar mechanical properties, and more recently for their potential application in catalysis. They are a considerable challenge to traditional atomistic modeling, and also to data-driven potentials that for the most part have memory footprint, computational effort and data requirements which scale poorly with the number of elements included. We apply a recently proposed scheme to compress chemical information in a lower-dimensional space, which reduces dramatically the cost of the model with negligible loss of accuracy, to build a potential that can describe 25 d-block transition metals. The model shows semi-quantitative accuracy for prototypical alloys, and is remarkably stable when extrapolating to structures outside its training set. We use this framework to study element segregation in a computational experiment that simulates an equimolar alloy of all 25 elements, mimicking the seminal experiments by Cantor et al., and use our observations on the short-range order relations between the elements to define a data-driven set of Hume-Rothery rules that can serve as guidance for alloy design. We conclude with a study of three prototypical alloys, CoCrFeMnNi, CoCrFeMoNi and IrPdPtRhRu, determining their stability and the short-range order behavior of their constituents.

cond-mat.mtrl-sci

An assessment of the structural resolution of various fingerprints commonly used in machine learning

Atomic environment fingerprints are widely used in computational materials science, from machine learning potentials to the quantification of similarities between atomic configurations. Many approaches to the construction of such fingerprints, also called structural descriptors, have been proposed. In this work, we compare the performance of fingerprints based on the Overlap Matrix(OM), the Smooth Overlap of Atomic Positions (SOAP), Behler-Parrinello atom-centered symmetry functions (ACSF), modified Behler-Parrinello symmetry functions (MBSF) used in the ANI-1ccx potential and the Faber-Christensen-Huang-Lilienfeld (FCHL) fingerprint under various aspects. We study their ability to resolve differences in local environments and in particular examine whether there are certain atomic movements that leave the fingerprints exactly or nearly invariant. For this purpose, we introduce a sensitivity matrix whose eigenvalues quantify the effect of atomic displacement modes on the fingerprint. Further, we check whether these displacements correlate with the variation of localized physical quantities such as forces. Finally, we extend our examination to the correlation between molecular fingerprints obtained from the atomic fingerprints and global quantities of entire molecules.

physics.comp-ph

Chemical Shifts in Molecular Solids by Machine Learning

The calculation of chemical shifts in solids has enabled methods to determine crystal structures in powders. The dependence of chemical shifts on local atomic environments sets them among the most powerful tools for structure elucidation of powdered solids or amorphous materials. Unfortunately, this dependency comes with the cost of high accuracy first-principle calculations to qualitatively predict chemical shifts in solids. Machine learning methods have recently emerged as a way to overcome the need for explicit high accuracy first-principle calculations. However, the vast chemical and combinatorial space spanned by molecular solids, together with the strong dependency of chemical shifts of atoms on their environment, poses a huge challenge for any machine learning method. Here we propose a machine learning method based on local environments to accurately predict chemical shifts of different molecular solids and of different polymorphs within DFT accuracy (RMSE of 0.49 ppm ( 1 H), 4.3ppm ( 13 C), 13.3 ppm ( 15 N), and 17.7 ppm ( 17 O) with $R^2$ of 0.97 for 1 H, 0.99 for 13 C, 0.99 for 15 N, and 0.99 for 17 O). We also demonstrate that the trained model is able to correctly determine, based on the match between experimentally-measured and ML-predicted shifts, structures of cocaine and the drug 4-[4-(2-adamantylcarbamoyl)-5-tert-butylpyrazol-1-yl]benzoic acid in an chemical shift based NMR crystallography approach.

physics.chem-ph

Machine Learning Unifies the Modelling of Materials and Molecules

Determining the stability of molecules and condensed phases is the cornerstone of atomistic modelling, underpinning our understanding of chemical and materials properties and transformations. Here we show that a machine learning model, based on a local description of chemical environments and Bayesian statistical learning, provides a unified framework to predict atomic-scale properties. It captures the quantum mechanical effects governing the complex surface reconstructions of silicon, predicts the stability of different classes of molecules with chemical accuracy, and distinguishes active and inactive protein ligands with more than 99% reliability. The universality and the systematic nature of our framework provides new insight into the potential energy surface of materials and molecules.

cond-mat.mtrl-sci

Mapping and Classifying Molecules from a High-Throughput Structural Database

High-throughput computational materials design promises to greatly accelerate the process of discovering new materials and compounds, and of optimizing their properties. The large databases of structures and properties that result from computational searches, as well as the agglomeration of data of heterogeneous provenance leads to considerable challenges when it comes to navigating the database, representing its structure at a glance, understanding structure-property relations, eliminating duplicates and identifying inconsistencies. Here we present a case study, based on a data set of conformers of amino acids and dipeptides, of how machine-learning techniques can help addressing these issues. We will exploit a recently developed strategy to define a metric between structures, and use it as the basis of both clustering and dimensionality reduction techniques showing how these can help reveal structure-property relations, identify outliers and inconsistent structures, and rationalise how perturbations (e.g. binding of ions to the molecule) affect the stability of different conformers.

physics.chem-ph

Comparing molecules and solids across structural and alchemical space

Evaluating the (dis)similarity of crystalline, disordered and molecular compounds is a critical step in the development of algorithms to navigate automatically the configuration space of complex materials. For instance, a structural similarity metric is crucial for classifying structures, searching chemical space for better compounds and materials, and driving the next generation of machine-learning techniques for predicting the stability and properties of molecules and materials. In the last few years several strategies have been designed to compare atomic coordination environments. In particular, the Smooth Overlap of Atomic Positions (SOAP) has emerged as an elegant framework to obtain translation, rotation and permutation-invariant descriptors of groups of atoms, driven by the design of various classes of machine-learned inter-atomic potentials. Here we discuss how one can combine such local descriptors using a Regularized Entropy Match (REMatch) approach to describe the similarity of both whole molecular and bulk periodic structures, introducing powerful metrics that enable the navigation of alchemical and structural complexity within a unified framework. Furthermore, using this kernel and a ridge regression method we can predict atomization energies for a database of small organic molecules with a mean absolute error below 1kcal/mol, reaching an important milestone in the application of machine-learning techniques to the evaluation of molecular properties.

cond-mat.mtrl-sci

Glassy clusters: Relations between their dynamics and characteristic features of their energy landscape

Based on a recently introduced metric for measuring distances between configurations, we in- troduce distance-energy (DE) plots to characterize the potential energy surface (PES) of clusters. Producing such plots is computationally feasible on the density functional (DFT) level since it re- quires only a set of a few hundred stable low energy configurations including the global minimum. By comparison with standard criteria based on disconnectivity graphs and on the dynamics of Lennard- Jones clusters we show that the DE plots convey the necessary information about the character of the potential energy surface and allow to distinguish between glassy and non-glassy systems. We then apply this analysis to real systems on the DFT level and show that both glassy and non-glassy clusters can be found in simulations. It however turns out that among our investigated clusters only those can be synthesized experimentally which exhibit a non-glassy landscape.

physics.atm-clus

The effect of ionization on the global minima of small and medium sized silicon and magnesium clusters

We re-examine the question of whether the geometrical ground state of neutral and ionized clusters are identical. Using a well defined criterion for being "identical" together, the extensive sampling methods on a potential energy surface calculated by density functional theory, we show that the ground states are in general different. This behavior is to be expected whenever there are metastable configurations which are close in energy to the ground state, but it disagrees with previous studies.

physics.atm-clus

The energy landscape of fullerene materials: a comparison between boron, boron-nitride and carbon

Using the minima hopping global geometry optimization method on the density functional potential energy surface we study medium size and large boron clusters. Even though for isolated medium size clusters the ground state is a cage like structure they are unstable against external perturbations such as contact with other clusters. The energy landscape of larger boron clusters is glass like and has a large number of structures which are lower in energy than the cages. This is in contrast to carbon and boron nitride systems which can be clearly identified as structure seekers in our minima hopping runs. The differences in the potential energy landscape explain why carbon and boron nitride systems are found in nature whereas pure boron fullerenes have not been found.

physics.atm-clus