Searcharxiv⌕ Search

arXiv subjects

Miguel A. Caro

Publications and source records attributed to Miguel A. Caro.

At least 19 recordsLinked to original sources

Cluster-based Structural Similarity for Dataset Visualization and Data Selection for Machine Learning Interatomic Potentials

Machine learning interatomic potentials (MLIPs) are essential components for accelerating simulation-driven materials design. Data-efficient MLIP training relies on data-selection strategies that maximize structural diversity while limiting computationally expensive first-principles calculations. A key challenge in such strategies is evaluating structural similarity, which involves a trade-off between retaining information on individual atomic environments and reducing computational cost. Here, we propose similarity evaluation methods that achieve both representational fidelity and computational efficiency. Our method represents each structure using a small set of characteristic atomic environments identified by k-medoids clustering and computes pairwise similarity through optimal matching between these representatives or their distributions. Molecular benchmarks demonstrate that our method is approximately 50 times faster than the baseline method while also more clearly distinguishing structures with different chemical compositions. Similarity-based data-selection benchmarks demonstrate that our methods improve the data efficiency and stability of force prediction in MLIPs.

cond-mat.mtrl-sci↗

Molecular augmented dynamics: Generating experimentally consistent atomistic structures by design

A fundamental objective of materials modeling is identifying atomic structures that align with experimental observables. Conventional approaches for disordered materials involve sampling from thermodynamic ensembles and hoping for an experimental match. This process is inefficient and offers no guarantee of success. We present a method based on modified molecular dynamics, that we call molecular augmented dynamics (MAD), which identifies structures that simultaneously match multiple experimental observables and exhibit low energies as described by a machine learning interatomic potential (MLIP) trained from ab-initio data. We demonstrate its feasibility by finding representative structures of glassy carbon, nanoporous carbon, ta-C, a-C:D and a-CO$_x$ that match their respective experimental observables---X-ray diffraction, neutron diffraction, pair distribution function and X-ray photoelectron spectroscopy data---using the same initial structure and underlying MLIP. The method is general, accepting any experimental observable whose simulated counterpart can be cast as a function of differentiable atomic descriptors. This method enables a computational "microscope" into experimental structures.

cond-mat.mtrl-sci↗

Linear-scaling calculation of experimental observables for molecular augmented dynamics simulations

Aligning theoretical atomistic structural models of materials with available experimental data presents a significant challenge for disordered systems. The configurational space to navigate is vast, and faithful realizations require large system sizes with quantum-mechanical accuracy in order to capture the distribution of structural motifs present in experiment. Traditional equilibrium sampling approaches offer no guarantee of generating structures that coincide with experimental data for such systems. An efficient means to search for such structures is molecular augmented dynamics (MAD) [arXiv:2508.17132], a modified molecular dynamics method that can generate ab-initio accurate, low-energy structures through a multi-objective optimization of the interatomic potential energy and the experimental potential. The computational scaling of this method depends on both the scaling of the interatomic potential and that of the experimental potential. We present the general equations for MAD with linear-scaling formulations for calculating and matching X-ray/neutron diffraction and local observables, e.g., the core-electron binding energies used in X-ray photoelectron spectroscopy. MAD simulations can both find metastable structures compatible with non-equilibrium experimental synthesis and lower energy structures than alternative computational sampling protocols, like the melt-quench approach. In addition, generalizing the virial tensor with the experimental forces enables generalized barostatting, allowing one to find structures whose density matches that compatible with the experimental observables. Scaling tests with the TurboGAP code demonstrate their linear-scaling nature for both CPU and GPU implementations, the latter of which has a 100$\times$ speedup compared to the CPU.

cond-mat.mtrl-sci↗

Density-dependent sodium-storage mechanisms in hard carbon materials

Understanding the sodium-storage mechanism in hard carbon (HC) anodes is crucial for advancing sodium-ion battery (SIB) technology. However, the intrinsic complexity of HC microstructures and their interactions with sodium remain not fully elucidated. We present a multiscale methodology that integrates grand-canonical Monte Carlo (GCMC) simulations with a machine-learning interatomic potential based on the Gaussian approximation potential (GAP) framework to investigate sodium insertion mechanisms in hard carbons with different levels of porosity, achieved by simulating structural models with densities ranging from 0.7 to 1.9 g cm$^{-3}$. Structural and thermodynamic analyses reveal the interplay between pore size and accessibility and the relative contributions of adsorption, intercalation, and pore filling to the overall storage capacity. Low-density carbons favor pore-filling, achieving extremely high capacities at near-zero voltages, whereas high-density carbons primarily store sodium through adsorption and intercalation, leading to lower but more stable capacities. Intermediate-density carbons ($1.3-1.6$ g cm$^{-3}$) provide the most balanced performance, combining moderate capacity (480 and 310 mAh g$^{-1}$), safe operating voltages, and minimal volume expansion ($<10$\%). These findings establish a direct correlation between carbon density and electrochemical behavior, providing atomic-scale insight into how hard carbon morphology governs sodium-storage. The proposed framework offers a rational design principle for optimizing HC-based SIB anodes toward high energy density and long-term cycling stability.

cond-mat.mtrl-sci↗

Improved capabilities of the TurboGAP code for radiation induced cascade simulations: an illustration with silicon

TurboGAP is a software package designed for efficient molecular dynamics simulations using Gaussian Approximation Potential (GAP) machine-learning interatomic potentials (MLIP). In this work, we enhance the capabilities of TurboGAP for radiation damage simulations by implementing a two-temperature molecular dynamics model, based on electron density-dependent coupling of electronic and atomic subsystems. Additionally, we implement adaptive calculation of the timestep and grouping of atoms for cell-border cooling. Our implementation incorporates electronic stopping power either through a traditional friction-based model or a more realistic first-principles-derived model. By combining the computational efficiency of TurboGAP with the accuracy of GAP MLIP, we perform cascade simulations in silicon with primary knock-on atom (PKA) energies up to 10 keV. Our simulations scale to systems containing up to 1 million atoms. We study the generation and clustering of radiation-induced defects. We also calculate ion-beam mixing and compare our results with the experimental data, discussing how the GAP-MLIP along with the inclusion of a realistic electronic stopping model improves the prediction of experimental mixing values.

physics.app-ph↗

Density dependence of thermal conductivity in nanoporous and amorphous carbon with machine-learned molecular dynamics

Disordered forms of carbon are an important class of materials for applications such as thermal management. However, a comprehensive theoretical understanding of the structural dependence of thermal transport and the underlying microscopic mechanisms is lacking. Here we study the structure-dependent thermal conductivity of disordered carbon by employing molecular dynamics (MD) simulations driven by a machine-learned interatomic potential based on the efficient neuroevolution potential approach. Using large-scale MD simulations, we generate realistic nanoporous carbon (NP-C) samples with density varying from $0.3$ to $1.5$ g cm$^{-3}$ dominated by sp$^2$ motifs, and amorphous carbon (a-C) samples with density varying from $1.5$ to $3.5$ g cm$^{-3}$ exhibiting mixed sp$^2$ and sp$^3$ motifs. Structural properties including short- and medium-range order are characterized by atomic coordination, pair correlation function, angular distribution function and structure factor. Using the homogeneous nonequilibrium MD method and the associated quantum-statistical correction scheme, we predict a linear and a superlinear density dependence of thermal conductivity for NP-C and a-C, respectively, in good agreement with relevant experiments. The distinct density dependences are attributed to the different impacts of the sp$^2$ and sp$^3$ motifs on the spectral heat capacity, vibrational mean free paths and group velocity. We additionally highlight the significant role of structural order in regulating the thermal conductivity of disordered carbon.

cond-mat.mtrl-sci↗

On-the-fly reparametrization of pairwise dispersion interactions for accurate and efficient molecular dynamics: Phase diagram of white phosphorus

Accurate estimation of the contribution from dispersion interactions to the total energy is important for many molecular systems and low-dimensional solids. In this work we demonstrate how the recently developed linear-scaling many-body dispersion correction method can be efficiently applied in molecular dynamics simulations while keeping high accuracy. This is achieved by reparametrization of the effective pairwise dispersion interactions on the fly during the simulation. We demonstrate this method by computing order-disorder and solid-liquid transitions of the phase diagram of white phosphorus (P$_4$).

physics.chem-ph↗

Unifying the description of hydrocarbons and hydrogenated carbon materials with a chemically reactive machine learning interatomic potential

We present a general-purpose machine learning (ML) interatomic potential for carbon and hydrogen which is capable of simulating various materials and molecules composed of these elements. This ML interatomic potential is trained using the Gaussian approximation potential (GAP) framework and an extensive dataset of C-H configurations obtained from density functional theory. The dataset is constructed through iterative training and structure-search techniques that generate a broad range of configurations to comprehensively sample the potential energy surface. Furthermore, the dataset is supplemented with relevant bulk, molecular, and high-pressure structures. Finally, long-range van der Waals interactions are added as a locally parametrized model. The accuracy and generality of the potential are validated through the analysis of different simulations under a wide range of conditions, including weak interactions, high temperature, and high pressure. We show that our CH GAP model describes different problems such as the formation of simple and complex alkanes, aromatic hydrocarbons, hydrogenated amorphous carbon (a-C:H), and CH systems at extreme conditions, while retaining good accuracy for pure carbon materials. We use this model to generate hydrocarbons of different sizes and complexity without prior knowledge of organic chemistry rules, and to highlight intrinsic limitations to the simultaneous description on intra and intermolecular interactions within a single computational framework. Our general-purpose ML interatomic potential has the capability to significantly advance research in the field of H-containing carbon materials and compounds, particularly in the areas where longer dynamics, reactivity and large-scale effects may be important.

physics.chem-ph↗

Atom-wise formulation of the many-body dispersion problem for linear-scaling van der Waals corrections

A common approach to modeling dispersion interactions and overcoming the inaccurate description of long-range correlation effects in electronic structure calculations is the use of pairwise-additive potentials, as in the Tkatchenko-Scheffler [Phys. Rev. Lett. 102, 073005 (2009)] method. In previous work [Phys. Rev. B 104, 054106 (2021)], we have shown how these are amenable to highly efficient atomistic simulation by machine learning their local parametrization. However, the atomic polarizability and the electron correlation energy have a complex and non-local many-body character and some of the dispersion effects in complex systems are not sufficiently described by these types of pairwise-additive potentials. Currently, one of the most widely used rigorous descriptions of the many-body effects is based on the many-body dispersion (MBD) model [Phys. Rev. Lett. 108, 236402 (2012)]. In this work, we show that the MBD model can also be locally parametrized to derive a local approximation for the highly non-local many-body effects. With this local parametrization, we develop an atom-wise formulation of MBD that we refer to as linear MBD (lMBD), as this decomposition enables linear scaling with system size. This model provides a transparent and controllable approximation to the full MBD model with tunable convergence parameters for a fraction of the computational cost observed in electronic structure calculations with popular density-functional theory codes. We show that our model scales linearly with the number of atoms in the system and is easily parallelizable. Furthermore, we show how using the same machinery already established in previous work for predicting Hirshfeld volumes with machine learning enables access to large-scale simulations with MBD-level corrections.

cond-mat.mtrl-sci↗

Anomalous Enhancement of the Electrocatalytic Hydrogen Evolution Reaction in AuPt Nanoclusters

Energy- and resource-efficient electrocatalytic water splitting is of paramount importance to enable sustainable hydrogen production. The best bulk catalyst for the hydrogen evolution reaction (HER), i.e., platinum, is one of the scarcest elements on Earth. The use of raw material for HER can be dramatically reduced by utilizing nanoclusters. In addition, nanoalloying can further improve the performance of these nanoclusters. In this paper, we present results for HER on nanometer-sized ligand-free AuPt nanoclusters grafted on carbon nanotubes. These results demonstrate excellent monodispersity and a significant reduction of the overpotential for the electrocatalytic HER. We utilize atomistic machine learning techniques to elucidate the atomic-scale origin of the synergistic effect between Pt and Au. We show that the presence of surface Au atoms, known to be poor HER catalysts, in a Pt(core)/AuPt(shell) nanocluster structure, drives an anomalous enhancement of the inherently high catalytic activity of Pt atoms.

physics.chem-ph↗

Syngas conversion to higher alcohols via wood-framed Cu/Co-carbon catalyst

Syngas conversion into higher alcohols represents a promising avenue for transforming coal or biomass into liquid fuels. However, the commercialization of this process has been hindered by the high cost, low activity, and inadequate C$_{2+}$OH selectivity of catalysts. Herein, we have developed Cu/Co carbon wood catalysts, offering a cost-effective and stable alternative with exceptional selectivity for catalytic conversion. The formation of Cu/Co nanoparticles was found, influenced by water-1,2-propylene glycol ratios in the solution, resulting in bidisperse nanoparticles. The catalyst exhibited a remarkable CO conversion rate of 74.8% and a selectivity of 58.7% for C$_{2+}$OH, primarily comprising linear primary alcohols. This catalyst demonstrated enduring stability and selectivity under industrial conditions, maintaining its efficacy for up to 350 h of operation. We also employed density functional theory (DFT) to analyze selectivity, particularly focusing on the binding strength of CO, a crucial precursor for subsequent reactions leading to the formation of CH$_3$OH. DFT identified the pathway of CH$_x$ and CO coupling, ultimately yielding C$_2$H$_5$OH. This computational understanding, coupled with high performance of the Cu/Co-carbon wood catalyst, paves ways for the development of catalytically selective materials tailored for higher alcohols production, thereby ushering in new possibility in this field.

physics.chem-ph↗

Cluster-based multidimensional scaling embedding tool for data visualization

We present a new technique for visualizing high-dimensional data called cluster MDS (cl-MDS), which addresses a common difficulty of dimensionality reduction methods: preserving both local and global structures of the original sample in a single 2-dimensional visualization. Its algorithm combines the well-known multidimensional scaling (MDS) tool with the $k$-medoids data clustering technique, and enables hierarchical embedding, sparsification and estimation of 2-dimensional coordinates for additional points. While cl-MDS is a generally applicable tool, we also include specific recipes for atomic structure applications. We apply this method to non-linear data of increasing complexity where different layers of locality are relevant, showing a clear improvement in their retrieval and visualization quality.

cs.GR↗

Accelerated First-Principles Exploration of Structure and Reactivity in Graphene Oxide

Graphene oxide (GO) materials are widely studied, and yet their atomic-scale structures remain to be fully understood. Here we show that the chemical and configurational space of GO can be rapidly explored by advanced machine-learning methods, combining on-the-fly acceleration for first-principles molecular dynamics with message-passing neural-network potentials. The first step allows for the rapid sampling of chemical structures with very little prior knowledge required; the second step affords state-of-the-art accuracy and predictive power. We apply the method to the thermal reduction of GO, which we describe in a realistic (ten-nanometre scale) structural model. Our simulations are consistent with recent experimental findings and help to rationalise them in atomistic and mechanistic detail. More generally, our work provides a platform for routine, accurate, and predictive simulations of diverse carbonaceous materials.

physics.chem-ph↗

Experiment-driven atomistic materials modeling: A case study combining X-ray photoelectron spectroscopy and machine learning potentials to infer the structure of oxygen-rich amorphous carbon

An important yet challenging aspect of atomistic materials modeling is reconciling experimental and computational results. Conventional approaches involve generating numerous configurations through molecular dynamics or Monte Carlo structure optimization and selecting the one with the closest match to experiment. However, this inefficient process is not guaranteed to succeed. We introduce a general method to combine atomistic machine learning (ML) with experimental observables that produces atomistic structures compatible with experiment by design. We use this approach in combination with grand-canonical Monte Carlo within a modified Hamiltonian formalism, to generate configurations that agree with experimental data and are chemically sound (low in energy). We apply our approach to understand the atomistic structure of oxygenated amorphous carbon (a-CO$_{x}$), an intriguing carbon-based material, to answer the question of how much oxygen can be added to carbon before it fully decomposes into CO and CO$_2$. Utilizing an ML-based X-ray photoelectron spectroscopy (XPS) model trained from $GW$ and density functional theory (DFT) data, in conjunction with an ML interatomic potential, we identify a-CO$_{x}$ structures compliant with experimental XPS predictions that are also energetically favorable with respect to DFT. Employing a network analysis, we accurately deconvolve the XPS spectrum into motif contributions, both revealing the inaccuracies inherent to experimental XPS interpretation and granting us atomistic insight into the structure of a-CO$_{x}$. This method generalizes to multiple experimental observables and allows for the elucidation of the atomistic structure of materials directly from experimental data, thereby enabling experiment-driven materials modeling with a degree of realism previously out of reach.

cond-mat.mtrl-sci↗

Gaussian Approximation Potentials: theory, software implementation and application examples

Gaussian Approximation Potentials are a class of Machine Learned Interatomic Potentials routinely used to model materials and molecular systems on the atomic scale. The software implementation provides the means for both fitting models using ab initio data and using the resulting potentials in atomic simulations. Details of the GAP theory, algorithms and software are presented, together with detailed usage examples to help new and existing users. We review some recent developments to the GAP framework, including MPI parallelisation of the fitting code enabling its use on thousands of CPU cores and compression of descriptors to eliminate the poor scaling with the number of different chemical elements.

cond-mat.mtrl-sci↗

Searching for iron nanoparticles with a general-purpose Gaussian approximation potential

We present a general-purpose machine learning Gaussian approximation potential (GAP) for iron that is applicable to all bulk crystal structures found experimentally under diverse thermodynamic conditions, as well as surfaces and nanoparticles (NPs). By studying its phase diagram, we show that our GAP remains stable at extreme conditions, including those found in the Earth's core. The new GAP is particularly accurate for the description of NPs. We use it to identify new low-energy NPs, whose stability is verified by performing density functional theory calculations on the GAP structures. Many of these NPs are lower in energy than those previously available in the literature up to $N_\text{atoms}=100$. We further extend the convex hull of available stable structures to $N_\text{atoms}=200$. For these NPs, we study characteristic surface atomic motifs using data clustering and low-dimensional embedding techniques. With a few exceptions, e.g., at magic numbers $N_\text{atoms}=59$, $65$, $76$ and $78$, we find that iron tends to form irregularly shaped NPs without a dominant surface character or characteristic atomic motif, and no reminiscence of crystalline features. We hypothesize that the observed disorder stems from an intricate balance and competition between the stable bulk motif formation, with bcc structure, and the stable surface motif formation, with fcc structure. We expect these results to improve our understanding of the fundamental properties and structure of low-dimensional forms of iron, and to facilitate future work in the field of iron-based catalysis.

cond-mat.mtrl-sci↗

A general-purpose machine learning Pt interatomic potential for an accurate description of bulk, surfaces and nanoparticles

A Gaussian approximation machine learning interatomic potential for platinum is presented. It has been trained on DFT data computed for bulk, surfaces and nanostructured platinum, in particular nanoparticles. Across the range of tested properties, which include bulk elasticity, surface energetics and nanoparticle stability, this potential shows excellent transferability and agreement with DFT, providing state-of-the-art accuracy at low computational cost. We showcase the possibilities for modeling of Pt systems enabled by this potential with two examples: the pressure-temperature phase diagram of Pt calculated using nested sampling and a study of the spontaneous crystallization of a large Pt nanoparticle based on classical dynamics simulations over several nanoseconds.

cond-mat.mtrl-sci↗

Machine learning based modeling of disordered elemental semiconductors: understanding the atomic structure of a-Si and a-C

Disordered elemental semiconductors, most notably a-C and a-Si, are ubiquitous in a myriad of different applications. These exploit their unique mechanical and electronic properties. In the past couple of decades, density functional theory (DFT) and other quantum mechanics-based computational simulation techniques have been successful at delivering a detailed understanding of the atomic and electronic structure of crystalline semiconductors. Unfortunately, the complex structure of disordered semiconductors sets the time and length scales required for DFT simulation of these materials out of reach. In recent years, machine learning (ML) approaches to atomistic modeling have been developed that provide an accurate approximation of the DFT potential energy surface for a small fraction of the computational time. These ML approaches have now reached maturity and are starting to deliver the first conclusive insights into some of the missing details surrounding the intricate atomic structure of disordered semiconductors. In this Topical Review we give a brief introduction to ML atomistic modeling and its application to amorphous semiconductors. We then take a look at how ML simulations have been used to improve our current understanding of the atomic structure of a-C and a-Si.

cond-mat.mtrl-sci↗