SearcharxivSearch

arXiv subjects

Patrick Rinke

Publications and source records attributed to Patrick Rinke.

At least 19 recordsLinked to original sources

Charting the thermodynamic stability of hybrid perovskite alloys with machine learning

Alloy-based perovskite solar cells offer tunable properties and improved stability, but their complexity has impeded accurate modeling, hindering development. We present a machine-learning (ML) accelerated atomistic modeling approach for the phase stability of (Cs/FA)Pb(Br/I)3 and (Cs/FA)Sn(Br/I)3 perovskites, with FA being formamidinium. To make such quaternary alloys tractable, we adopt a two-level ML strategy, combining 1) graph neural network interatomic potentials trained on density functional theory data for efficient structure relaxations with 2) secondary ML models for direct energy prediction from unrelaxed structures. These models enable computations of free energy landscapes across compositions and phases, capturing alloy disorder and FA molecular orientations. Our results reveal narrower stable composition regions for the Sn-based system compared to its Pb-based counterpart, limiting options for compositional engineering. Maximum stability occurs at high I content, and no stabilization is observed near the center of the composition space. Our results guide the design of stable perovskites.

cond-mat.mtrl-sci

Benchmarking machine-learned interatomic potentials for molecular infrared spectroscopy

Machine learning has transformed the field of atomistic simulations by enabling the development of interatomic potentials that are computationally efficient and highly accurate. These advances have opened the door to modeling molecular vibrations and predicting infrared spectra with near ab-initio accuracy at a fraction of the computational cost. Among these approaches, message-passing neural networks (MPNNs) have emerged as a particularly powerful class of models for representing complex atomic interactions. In this study, we benchmark five MPNN architectures, SchNet, FieldSchNet, SO3Net, PaiNN, and MACE, for predicting infrared spectra of small organic molecules. SchNet and FieldSchNet are invariant models, while SO3Net, PaiNN, and MACE are equivariant, explicitly accounting for rotational symmetries in molecular representations. We evaluate their performance in terms of computational efficiency, accuracy, and robustness. All models accurately predict properties, such as energies, forces, and dipole moments, required for infrared spectra calculations. They also capture harmonic frequencies and infrared spectra derived from molecular dynamics with high fidelity for molecules in the training set. However, SchNet and FieldSchNet show limited transferability to unseen systems, while SO3Net, PaiNN, and MACE generalize more effectively. In terms of computational efficiency, SchNet is the most efficient and FieldSchNet enables field-dependent response modeling but with higher cost. PaiNN achieves the best balance between accuracy and efficiency, MACE provides the highest spectral accuracy and transferability, and SO3Net performs between PaiNN and MACE.

physics.chem-ph

Selectivity- and Activity-Aware Catalyst Descriptors for CO$_2$ Hydrogenation on Alloy Nanocatalysts using Machine-Learned Force Fields

Adsorption energy distributions (AEDs) have emerged as a powerful and increasingly adopted descriptor for catalytic performance in high-entropy alloys and, more recently, in conventional metallic alloy nanocrystal catalysts. By accounting for diverse adsorption sites and crystallographic facets, AEDs more fully represent nanoparticle-based catalytic surfaces and show strong promise for accelerating rational design and discovery of heterogeneous catalysts, especially for CO$_2$ hydrogenation. However, previous high-throughput screenings have not provided simultaneous facet-level insight into catalytic activity, product selectivity, and thermodynamic stability. Here, we implement and extend the AED framework to individual crystallographic facets, combining facet-resolved adsorption fingerprints with Wulff-derived facet abundances and an interpretable latent-space analysis of C1-product selectivity. We employ universal machine-learned force fields trained on Open Catalyst Project data to compute adsorption energies across 226 experimentally observed metals, binary alloys, and previously unexplored ternary alloys, encompassing ~1.4 million adsorption configurations on >2,600 crystallographically distinct surfaces. Using distribution-based similarity analysis and principal component analysis of AED moments, we identify composition-facet combinations with adsorption landscapes associated with activity and qualitative selectivity trends toward methanol, methane, formic acid, and CO. Our framework links surface structure to adsorption-based catalytic trends, providing experimentally testable composition-facet candidates for further investigation and validation.

cond-mat.mtrl-sci

Bayesian Optimization for Mixed-Variable Problems in the Natural Sciences

Optimizing expensive black-box objectives over mixed search spaces is a common challenge across the natural sciences. Bayesian optimization (BO) offers sample-efficient strategies through probabilistic surrogate models and acquisition functions. However, its effectiveness diminishes in mixed or high-cardinality discrete spaces, where gradients are unavailable and optimizing the acquisition function becomes computationally demanding. In this work, we generalize the probabilistic reparameterization (PR) approach of Daulton et al. to handle non-equidistant discrete variables, enabling gradient-based optimization in fully mixed-variable settings with Gaussian process (GP) surrogates. With real-world scientific optimization tasks in mind, we conduct systematic benchmarks on synthetic and experimental objectives to obtain an optimized kernel formulations and demonstrate the robustness of our generalized PR method. We additionally show that, when combined with a modified BO workflow, our approach can efficiently optimize highly discontinuous and discretized objective landscapes. This work establishes a practical BO framework for addressing fully mixed optimization problems in the natural sciences, and is particularly well suited to autonomous laboratory settings where noise, discretization, and limited data are inherent.

cs.LG

Role of photonic interference in exciton-mediated magneto-optic responses

Coupled optical and magnetic excitations can give rise to remarkably strong magneto-optic responses. This is particularly evident in van der Waals magnets, such as the antiferromagnet CrSBr, where excitons and magnons emerge from the same electronic orbitals. While previous work has primarily focused on uncovering the magneto-electric origin of the resulting exciton-magnon interactions, the influence of photonic effects has received comparatively little attention. Here, we use numerical simulations to disentangle exciton-magnon coupling from the exciton-mediated magnon-photon interactions observed in optical experiments. Our simulations show the strong dependence of these interactions on photonic interference and dispersion effects near excitonic resonances. Such effects shape the optical response to coherent magnons and make it intrinsically non-linear in the magnon-induced exciton energy shift. Thermal magnons, which have a particularly pronounced impact on excitons, are found to even produce qualitatively different trends in optical signatures. Depending on weak or strong coupling of excitons and photons, the same exciton-magnon interaction can lead to a red-shift of optical modes, a nearly vanishing response, or their blue-shift. Finally, we demonstrate first steps towards optimizing the multi-parameter problem of efficient magnon-photon transduction using a machine-learning approach.

cond-mat.mtrl-sci

Predicting the Thermal Behavior of Semiconductor Defects with Equivariant Neural Networks

The presence of defects strongly influences semiconductor behavior. However, predicting the electronic properties of defective materials at finite temperatures remains computationally expensive even with density functional theory due to the large number of atoms in the simulation cell and the multitude of thermally accessible configurations. Here, we present a neural network-based framework to investigate the electronic properties of defective semiconductors at finite temperatures efficiently. We develop an active learning approach that integrates two advanced equivariant graph neural networks: MACE for atomic energies and forces and DeepH-E3 for the electronic Hamiltonian. Focusing on representative point defects in GaAs, we demonstrate computational accuracy comparable to density functional theory at a fraction of the computational cost, predicting the temperature-dependent band gap of defective GaAs directly from larger scale molecular dynamics trajectories with an accuracy of few tens of meV. Our results highlight the potential of equivariant neural networks for accurate atomic-scale predictions in complex, dynamically evolving materials.

cond-mat.mtrl-sci

An interpretable molecular descriptor for machine learning predictions in atmospheric science

The study of aerosol formation and chemistry using machine learning is limited by the lack of molecular descriptors suited to atmospheric compounds. Interpretable models are particularly affected because they often rely on dictionary-based descriptors tied to specific molecular substructures, which currently fail to capture the full range of organic atmospheric compounds, including large, highly oxidized molecules common in the atmosphere. We introduce ATMOMACCS, an interpretable descriptor combining the 166 binary keys of the MACCS fingerprint with motifs inspired by the SIMPOL method for estimating saturation vapor pressures. We show that ATMOMACCS based models improve predictions of saturation vapor pressures (7-8 % error reduction), equilibrium partition coefficients (5 % and 9 % error reduction), glass transition temperatures (22 % error reduction), and enthalpy of vaporization (61 % error reduction) on four datasets with atmospheric compounds. Feature analysis shows that saturation vapor pressure and partition coefficients are governed by carbon number and oxygen-related features, whereas other phase-transition properties (e.g., enthalpy of vaporization, glass transition temperature) depend on carbon-hydrogen bond types and the presence of heteroatoms other than oxygen. This highlights the generalizability of ATMOMACCS across different datasets and properties as an interpretable molecular descriptor.

physics.chem-ph

MACE4IRmol: An uncertainty-aware foundation model for molecular infrared spectroscopy

Machine-learned interatomic potentials (MLIPs) have shown significant promise in predicting infrared spectra with high fidelity. However, the absence of general-purpose MLIPs that simultaneously span broad chemical diversity and provide reliable uncertainty estimates has limited their wider applicability. In this work, we introduce MACE4IRmol, an uncertainty-aware foundation model ensemble built on the MACE architecture. MACE4IRmol is trained on ~16 million molecular geometries and the corresponding density-functional theory (DFT) energies, forces, and dipole moments from the QCML dataset. The training data encompasses approximately 80 elements and a diverse set of molecules, including organic and inorganic compounds, and metal complexes. Importantly, MACE4IRmol is formulated as an ensemble of models to enable uncertainty quantification, which helps improve robustness in chemically diverse systems. Within this ensemble, separate models are trained with and without explicit dispersion corrections, allowing systematic assessment of van der Waals effects. In addition, MACE4IRmol delivers accurate predictions of energies, forces, dipole moments, and infrared spectra at a fraction of the computational cost of DFT, while enabling the explicit inclusion of nuclear quantum effects in infrared spectrum simulations. By combining generality, accuracy, efficiency, and uncertainty estimation, MACE4IRmol opens the door to rapid and reliable infrared spectra prediction for complex and diverse molecular systems.

physics.chem-ph

Leveraging active learning-enhanced machine-learned interatomic potential for efficient infrared spectra prediction

Infrared (IR) spectroscopy is a pivotal analytical tool as it provides real-time molecular insight into material structures and enables the observation of reaction intermediates in situ. However, interpreting IR spectra often requires high-fidelity simulations, such as density functional theory based ab-initio molecular dynamics, which are computationally expensive and therefore limited in the tractable system size and complexity. In this work, we present a novel active learning-based framework, implemented in the open-source software package PALIRS, for efficiently predicting the IR spectra of small catalytically relevant organic molecules. PALIRS leverages active learning to train a machine-learned interatomic potential, which is then used for machine learning-assisted molecular dynamics simulations to calculate IR spectra. PALIRS reproduces IR spectra computed with ab-initio molecular dynamics accurately at a fraction of the computational cost. PALIRS further agrees well with available experimental data not only for IR peak positions but also for their amplitudes. This advancement with PALIRS enables high-throughput prediction of IR spectra, facilitating the exploration of larger and more intricate catalytic systems and aiding the identification of novel reaction pathways.

physics.chem-ph

Efficient dataset generation for machine learning perovskite alloys

Lead-based perovskite solar cells have reached high efficiencies, but toxicity and lack of stability hinder their wide-scale adoption. These issues have been partially addressed through compositional engineering of perovskite materials, but the vast complexity of the perovskite materials space poses a significant obstacle to exploration. We previously demonstrated how machine learning (ML) can accelerate property predictions for the CsPb(Cl/Br)$_3$ perovskite alloy. However, the substantial computational demand of density functional theory (DFT) calculations required for model training prevents applications to more complex materials. Here, we introduce a data-efficient scheme to facilitate model training, validated initially on CsPb(Cl/Br)$_3$ data and extended to the ternary alloy CsSn(Cl/Br/I)$_3$. Our approach employs clustering to construct a compact yet diverse initial dataset of atomic structures. We then apply a two-stage active learning approach to first improve the reliability of the ML-based structure relaxations and then refine accuracy near equilibrium structures. Tests for CsPb(Cl/Br)$_3$ demonstrate that our scheme reduces the number of required DFT calculations during the different parts of our proposed model training method by up to 20% and 50%. The fitted model for CsSn(Cl/Br/I)$_3$ is robust and highly accurate, evidenced by the convergence of all ML-based structure relaxations in our tests and an average relaxation error of only 0.5 meV/atom.

cond-mat.mtrl-sci

Design Rules for Optimizing Quaternary Mixed-Metal Chalcohalides

Quaternary mixed-metal M(II)2M(III)Ch2X3 chalcohalides are an emerging material class for photovoltaic absorbers that combines the beneficial optoelectronic properties of lead-based halide perovskites with the stability of metal chalcogenides. Inspired by the recent discovery of lead-free mixed-metal chalcohalides materials, we utilized a combination of density functional theory and machine learning to determine compositional trends and chemical design rules in the lead-free and lead-based materials spaces. We explored a total of 54 M(II)2M(III)Ch2X3 materials with M(II) = Sn, Pb, M(III) = In, Sb, Bi, Ch = S, Se, Te, and X = Cl, Br, I per phase (Cmcm, Cmc21 , and P21/c). The P21/c phase is the equilibrium phase at low temperatures, followed by Cmc21 and Cmcm. The fundamental band gaps in Cmcm and Cmc21 are smaller than those in P21/c, but direct band gaps are more common in Cmcm and Cmc21. The effective electron masses in P21/c are significantly larger compared to Cmcm and Cmc21, while the effective hole masses are nearly the same across all three phases. Using random forest regression, we found that the two electron acceptor sites (Ch and X) are crucial in shaping the properties of mixed-metal chalcohalide compounds. Furthermore, the electron donor sites (M(II) and M(III)) can be used to finetune the material properties to desired applications. These design rules enable precise tailoring of mixed-metal chalcohalide compounds for a variety of applications.

cond-mat.mtrl-sci

Exploring Noncollinear Magnetic Energy Landscapes with Bayesian Optimization

The investigation of magnetic energy landscapes and the search for ground states of magnetic materials using ab initio methods like density functional theory (DFT) is a challenging task. Complex interactions, such as superexchange and spin-orbit coupling, make these calculations computationally expensive and often lead to non-trivial energy landscapes. Consequently, a comprehensive and systematic investigation of large magnetic configuration spaces is often impractical. We approach this problem by utilizing Bayesian Optimization, an active machine learning scheme that has proven to be efficient in modeling unknown functions and finding global minima. Using this approach we can obtain the magnetic contribution to the energy as a function of one or more spin canting angles with relatively small numbers of DFT calculations. To assess the capabilities and the efficiency of the approach we investigate the noncollinear magnetic energy landscapes of selected materials containing 3d, 5d and 5f magnetic ions: Ba$_3$MnNb$_2$O$_9$, LaMn$_2$Si$_2$, $\beta$-MnO$_2$, Sr$_2$IrO$_4$, UO$_2$ and Ba$_2$NaOsO$_6$. By comparing our results to previous ab initio studies that followed more conventional approaches, we observe significant improvements in efficiency.

cond-mat.mtrl-sci

Machine Learning Accelerated Descriptor Design for Catalyst Discovery in CO$_2$ to Methanol Conversion

Transforming CO$_2$ into methanol represents a crucial step towards closing the carbon cycle, with thermoreduction technology nearing industrial application. However, obtaining high methanol yields and ensuring the stability of heterocatalysts remain significant challenges. Herein, we present a sophisticated computational framework to accelerate the discovery of thermal heterogeneous catalysts, using machine-learned force fields. We propose a new catalytic descriptor, termed adsorption energy distribution, that aggregates the binding energies for different catalyst facets, binding sites, and adsorbates. The descriptor is versatile and can be adjusted to a specific reaction through careful choice of the key-step reactants and reaction intermediates. By applying unsupervised machine learning and statistical analysis to a dataset comprising nearly 160 metallic alloys, we offer a powerful tool for catalyst discovery. We propose new promising candidates such as ZnRh and ZnPt$_3$, which to our knowledge, have not yet been tested, and discuss their possible advantage in terms of stability.

physics.chem-ph

Precision benchmarks for solids: G0W0 calculations with different basis sets

The GW approximation within many-body perturbation theory is the state of the art for computing quasiparticle energies in solids. Typically, Kohn-Sham (KS) eigenvalues and eigenfunctions, obtained from a Density Functional Theory (DFT) calculation are used as a starting point to build the Green's function G and the screened Coulomb interaction W, yielding the one-shot G0W0 selfenergy if no further update of these quantities are made. Multiple implementations exist for both the DFT and the subsequent G0W0 calculation, leading to possible differences in quasiparticle energies. In the present work, the G0W0 quasiparticle energies for states close to the band gap are calculated for six crystalline solids, using four different codes: Abinit, exciting, FHI-aims, and GPAW. This comparison helps to assess the impact of basis-set types (planewaves versus localized orbitals) and the treatment of core and valence electrons (all-electron full potentials versus pseudopotentials). The impact of unoccupied states as well as the algorithms for solving the quasiparticle equation are also briefly discussed. For the KS-DFT band gaps, we observe good agreement between all codes, with differences not exceeding 0.1 eV, while the G0W0 results deviate on the order of 0.1-0.3 eV. Between all-electron codes (FHI-aims and exciting), the agreement is better than 15 meV for KS-DFT and, with one exception, about 0.1 eV for G0W0 band gaps.

cond-mat.mtrl-sci

Active Learning of Molecular Data for Task-Specific Objectives

Active learning (AL) has shown promise for being a particularly data-efficient machine learning approach. Yet, its performance depends on the application and it is not clear when AL practitioners can expect computational savings. Here, we carry out a systematic AL performance assessment for three diverse molecular datasets and two common scientific tasks: compiling compact, informative datasets and targeted molecular searches. We implemented AL with Gaussian processes (GP) and used the many-body tensor as molecular representation. For the first task, we tested different data acquisition strategies, batch sizes and GP noise settings. AL was insensitive to the acquisition batch size and we observed the best AL performance for the acquisition strategy that combines uncertainty reduction with clustering to promote diversity. However, for optimal GP noise settings, AL did not outperform randomized selection of data points. Conversely, for targeted searches, AL outperformed random sampling and achieved data savings up to 64%. Our analysis provides insight into this task-specific performance difference in terms of target distributions and data collection strategies. We established that the performance of AL depends on the relative distribution of the target molecules in comparison to the total dataset distribution, with the largest computational savings achieved when their overlap is minimal.

cs.LG

Similarity-Based Analysis of Atmospheric Organic Compounds for Machine Learning Applications

The formation of aerosol particles in the atmosphere impacts air quality and climate change, but many of the organic molecules involved remain unknown. Machine learning could aid in identifying these compounds through accelerated analysis of molecular properties and detection characteristics. However, such progress is hindered by the current lack of curated datasets for atmospheric molecules and their associated properties. To tackle this challenge, we propose a similarity analysis that connects atmospheric compounds to existing large molecular datasets used for machine learning development. We find a small overlap between atmospheric and non-atmospheric molecules using standard molecular representations in machine learning applications. The identified out-of-domain character of atmospheric compounds is related to their distinct functional groups and atomic composition. Our investigation underscores the need for collaborative efforts to gather and share more molecular-level atmospheric chemistry data. The presented similarity based analysis can be used for future dataset curation for machine learning development in the atmospheric sciences.

physics.ao-ph

Learning Relevant Contextual Variables Within Bayesian Optimization

Contextual Bayesian Optimization (CBO) efficiently optimizes black-box functions with respect to design variables, while simultaneously integrating contextual information regarding the environment, such as experimental conditions. However, the relevance of contextual variables is not necessarily known beforehand. Moreover, contextual variables can sometimes be optimized themselves at an additional cost, a setting overlooked by current CBO algorithms. Cost-sensitive CBO would simply include optimizable contextual variables as part of the design variables based on their cost. Instead, we adaptively select a subset of contextual variables to include in the optimization, based on the trade-off between their relevance and the additional cost incurred by optimizing them compared to leaving them to be determined by the environment. We learn the relevance of contextual variables by sensitivity analysis of the posterior surrogate model while minimizing the cost of optimization by leveraging recent developments on early stopping for BO. We empirically evaluate our proposed Sensitivity-Analysis-Driven Contextual BO (SADCBO) method against alternatives on both synthetic and real-world experiments, together with extensive ablation studies, and demonstrate a consistent improvement across examples.

cs.LG

Validation of the GreenX library time-frequency component for efficient GW and RPA calculations

Electronic structure calculations based on many-body perturbation theory (e.g. GW or the random-phase approximation (RPA)) require function evaluations in the complex time and frequency domain, for example inhomogeneous Fourier transforms or analytic continuation from the imaginary axis to the real axis. For inhomogeneous Fourier transforms, the time-frequency component of the GreenX library provides time-frequency grids that can be utilized in low-scaling RPA and GW implementations. In addition, the adoption of the compact frequency grids provided by our library also reduces the computational overhead in RPA implementations with conventional scaling. In this work, we present low-scaling GW and conventional RPA benchmark calculations using the GreenX grids with different codes (FHI-aims, CP2K and ABINIT) for molecules, two-dimensional materials and solids. Very small integration errors are observed when using 30 time-frequency points for our test cases, namely $<10^{-8}$ eV/electron for the RPA correlation energies, and 10 meV for the GW quasiparticle energies.

physics.comp-ph