Searcharxiv⌕ Search

arXiv subjects

Andrew S. Rosen

Publications and source records attributed to Andrew S. Rosen.

12 recordsLinked to original sources

Democratizing Atomistic Simulation Workflows for the AI Era with the Quantum Accelerator

We present the Quantum Accelerator (QuAcc), an open-source workflow library for atomistic simulations with an emphasis on quantum-mechanical calculations. QuAcc provides predefined workflow recipes spanning first-principles electronic-structure methods, semiempirical and tight-binding approaches, classical potentials, and foundation machine-learned interatomic potentials (MLIPs). A central design feature of QuAcc is its separation of domain-specific scientific logic from the workflow engine used to orchestrate and execute calculations. Workflows are written as ordinary Python functions and can be executed with multiple supported workflow engines without modifying the underlying source code, lowering the barrier to developing and contributing new workflows. QuAcc also streamlines the evaluation of foundation MLIPs by providing a unified platform for generating ab initio reference calculations consistent with the model of interest, mitigating methodological drift when assessing model performance. Together, these features make QuAcc a flexible and accessible framework for atomistic simulation workflows that have become central to the current era of machine learning and artificial intelligence.

cond-mat.mtrl-sci↗

Towards a Metal-Organic Framework with Pore-Confined Electrons

Electrides are an unconventional class of materials in which electrons are localized in crystallographic void spaces rather than solely around atomic nuclei, giving rise to appealing properties such as low work functions, strong electron-donating character, and even superconductivity. Here, we use ab initio methods to investigate metal-organic framework (MOF) electrides, a new class of materials that combines the interstitial electrons of electrides with the permanent porosity and chemical tunability of MOFs. These materials host pore-confined electrons: occupied electronic states localized in the pore space and with bands slightly below or crossing through the Fermi level. Using density functional theory calculations, we establish several design rules for stabilizing pore-confined electrons in MOFs via an anion-electron exchange process and identify candidate MOF electrides. As a proof-of-concept, we also demonstrate that the pore-confined electrons can directly facilitate chemical reactions, substantially lowering the activation barrier for H2 dissociation without requiring adsorption at a surface site. We envision that pore-confined electrons in nanoporous materials may enable a fundamentally new type of catalysis in which chemical reactions take place in the pore space, driven by electron-centered active sites.

cond-mat.mtrl-sci↗

Fine-Tuning a Universal Machine-Learned Interatomic Potential for Oxygen Plasma Interactions with WS$_2$

Molecular dynamics simulation of plasma-surface interactions requires an interatomic potential that is simultaneously accurate, computationally efficient, and able to describe many elements and bonding types in reactive systems. In principle, a foundation model for machine-learned interatomic potential (MLIP) can meet these demands. We explore the use of the Universal Models for Atoms (UMA) model, developed by Meta FAIR, for the interactions of oxygen plasma species on a multilayer of WS$_2$, a promising 2D material. Starting from the pretrained uma-s-1p1 model under the Open Catalyst 2020 (OC20) task, we apply an iterative fine-tuning loop. Even in the absence of fine-tuning, the pretrained model reproduces the production-scale observables of interest, namely, chemisorbed S and O coverage under 15 eV O$^+$ and O$_2^+$ bombardment. These results were obtained without spin polarization and Hubbard $U$ correction. Nonetheless, fine-tuning with spin polarization and a Hubbard $U$ correction reduces the energy and force mean absolute error (MAE) to $4.5\times10^{-3}$ eV/atom and $0.076$ eV/angstrom, respectively.

cond-mat.mtrl-sci↗

The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models

Machine learning (ML) models hold the promise of transforming atomic simulations by delivering quantum chemical accuracy at a fraction of the computational cost. Realization of this potential would enable high-throughout, high-accuracy molecular screening campaigns to explore vast regions of chemical space and facilitate ab initio simulations at sizes and time scales that were previously inaccessible. However, a fundamental challenge to creating ML models that perform well across molecular chemistry is the lack of comprehensive data for training. Despite substantial efforts in data generation, no large-scale molecular dataset exists that combines broad chemical diversity with a high level of accuracy. To address this gap, Meta FAIR introduces Open Molecules 2025 (OMol25), a large-scale dataset composed of more than 100 million density functional theory (DFT) calculations at the $ω$B97M-V/def2-TZVPD level of theory, representing billions of CPU core-hours of compute. OMol25 uniquely blends elemental, chemical, and structural diversity including: 83 elements, a wide-range of intra- and intermolecular interactions, explicit solvation, variable charge/spin, conformers, and reactive structures. There are ~83M unique molecular systems in OMol25 covering small molecules, biomolecules, metal complexes, and electrolytes, including structures obtained from existing datasets. OMol25 also greatly expands on the size of systems typically included in DFT datasets, with systems of up to 350 atoms. In addition to the public release of the data, we provide baseline models and a comprehensive set of model evaluations to encourage community engagement in developing the next-generation ML models for molecular chemistry.

physics.chem-ph↗

A foundation model for atomistic materials chemistry

Atomistic simulations of matter, especially those that leverage first-principles (ab initio) electronic structure theory, provide a microscopic view of the world, underpinning much of our understanding of chemistry and materials science. Over the last decade or so, machine-learned force fields have transformed atomistic modeling by enabling simulations of ab initio quality over unprecedented time and length scales. However, early ML force fields have largely been limited by: (i) the substantial computational and human effort of developing and validating potentials for each particular system of interest; and (ii) a general lack of transferability from one chemical system to the next. Here we show that it is possible to create a general-purpose atomistic ML model, trained on a public dataset of moderate size, that is capable of running stable molecular dynamics for a wide range of molecules and materials. We demonstrate the power of the MACE-MP-0 model - and its qualitative and at times quantitative accuracy - on a diverse set of problems in the physical sciences, including properties of solids, liquids, gases, chemical reactions, interfaces and even the dynamics of a small protein. The model can be applied out of the box as a starting or "foundation" model for any atomistic system of interest and, when desired, can be fine-tuned on just a handful of application-specific data points to reach ab initio accuracy. Establishing that a stable force-field model can cover almost all materials changes atomistic modeling in a fundamental way: experienced users get reliable results much faster, and beginners face a lower barrier to entry. Foundation models thus represent a step towards democratising the revolution in atomic-scale modeling that has been brought about by ML force fields.

physics.chem-ph↗

An accurate and efficient framework for modelling the surface chemistry of ionic materials

Quantum-mechanical simulations can offer atomic-level insights into chemical processes on surfaces. This understanding is crucial for the rational design of new solid catalysts as well as materials to store energy and mitigate greenhouse gases. However, achieving the accuracy needed for reliable predictions has proven challenging. Density functional theory (DFT), the workhorse quantum-mechanical method, can often lead to inconsistent predictions, necessitating accurate methods from correlated wave-function theory (cWFT). However, the high computational demands and significant user intervention associated with cWFT have traditionally made it impractical to carry out for surfaces. In this work, we address this challenge, presenting an automated framework which leverages multilevel embedding approaches, to apply accurate cWFT methods to the surfaces of ionic materials with computational costs approaching DFT. With this framework, we have reproduced experimental adsorption enthalpies for a diverse set of 19 adsorbate-surface systems. Moreover, we resolve debates on the adsorption configuration of several systems, while offering benchmarks to assess DFT. This framework is open-source, making it possible to more routinely apply cWFT to complex problems involving the surfaces of ionic materials.

physics.chem-ph↗

Machine Learned Potential for High-Throughput Phonon Calculations of Metal-Organic Frameworks

Metal-organic frameworks (MOFs) are highly porous and versatile materials studied extensively for applications such as carbon capture and water harvesting. However, computing phonon-mediated properties in MOFs, like thermal expansion and mechanical stability, remains challenging due to the large number of atoms per unit cell, making traditional Density Functional Theory (DFT) methods impractical for high-throughput screening. Recent advances in machine learning potentials have led to foundation atomistic models, such as MACE-MP-0, that accurately predict equilibrium structures but struggle with phonon properties of MOFs. In this work, we developed a workflow for computing phonons in MOFs within the quasi-harmonic approximation with a fine-tuned MACE model, MACE-MP-MOF0. The model was trained on a curated dataset of 127 representative and diverse MOFs. The fine-tuned MACE-MP-MOF0 improves the accuracy of phonon density of states and corrects the imaginary phonon modes of MACE-MP-0, enabling high-throughput phonon calculations with state-of-the-art precision. The model successfully predicts thermal expansion and bulk moduli in agreement with DFT and experimental data for several well-known MOFs. These results highlight the potential of MACE-MP-MOF0 in guiding MOF design for applications in energy storage and thermoelectrics.

cond-mat.mtrl-sci↗

Deep Learning of ab initio Hessians for Transition State Optimization

Identifying transition states -- saddle points on the potential energy surface connecting reactant and product minima -- is central to predicting kinetic barriers and understanding chemical reaction mechanisms. In this work, we train an equivariant neural network potential, NewtonNet, on an ab initio dataset of thousands of organic reactions from which we derive the analytical Hessians from the fully differentiable machine learning (ML) model. By reducing the computational cost by several orders of magnitude relative to the Density Functional Theory (DFT) ab initio source, we can afford to use the learned Hessians at every step for the saddle point optimizations. We have implemented our ML Hessian algorithm in Sella, an open source software package designed to optimize atomic systems to find saddle point structures, in order to compare transition state optimization against quasi-Newton Hessian updates using DFT or the ML model. We show that the full ML Hessian robustly finds the transition states of 240 unseen organic reactions, even when the quality of the initial guess structures are degraded, while reducing the number of optimization steps to convergence by 2--3$\times$ compared to the quasi-Newton DFT and ML methods. All data generation, NewtonNet model, and ML transition state finding methods are available in an automated workflow.

physics.chem-ph↗

Investigating the Behavior of Diffusion Models for Accelerating Electronic Structure Calculations

We present an investigation into diffusion models for molecular generation, with the aim of better understanding how their predictions compare to the results of physics-based calculations. The investigation into these models is driven by their potential to significantly accelerate electronic structure calculations using machine learning, without requiring expensive first-principles datasets for training interatomic potentials. We find that the inference process of a popular diffusion model for de novo molecular generation is divided into an exploration phase, where the model chooses the atomic species, and a relaxation phase, where it adjusts the atomic coordinates to find a low-energy geometry. As training proceeds, we show that the model initially learns about the first-order structure of the potential energy surface, and then later learns about higher-order structure. We also find that the relaxation phase of the diffusion model can be re-purposed to sample the Boltzmann distribution over conformations and to carry out structure relaxations. For structure relaxations, the model finds geometries with ~10x lower energy than those produced by a classical force field for small organic molecules. Initializing a density functional theory (DFT) relaxation at the diffusion-produced structures yields a >2x speedup to the DFT relaxation when compared to initializing at structures relaxed with a classical force field.

physics.chem-ph↗

MOFDiff: Coarse-grained Diffusion for Metal-Organic Framework Design

Metal-organic frameworks (MOFs) are of immense interest in applications such as gas storage and carbon capture due to their exceptional porosity and tunable chemistry. Their modular nature has enabled the use of template-based methods to generate hypothetical MOFs by combining molecular building blocks in accordance with known network topologies. However, the ability of these methods to identify top-performing MOFs is often hindered by the limited diversity of the resulting chemical space. In this work, we propose MOFDiff: a coarse-grained (CG) diffusion model that generates CG MOF structures through a denoising diffusion process over the coordinates and identities of the building blocks. The all-atom MOF structure is then determined through a novel assembly algorithm. Equivariant graph neural networks are used for the diffusion model to respect the permutational and roto-translational symmetries. We comprehensively evaluate our model's capability to generate valid and novel MOF structures and its effectiveness in designing outstanding MOF materials for carbon capture applications with molecular simulations.

physics.chem-ph↗

Structured information extraction from complex scientific text with fine-tuned large language models

Intelligently extracting and linking complex scientific information from unstructured text is a challenging endeavor particularly for those inexperienced with natural language processing. Here, we present a simple sequence-to-sequence approach to joint named entity recognition and relation extraction for complex hierarchical information in scientific text. The approach leverages a pre-trained large language model (LLM), GPT-3, that is fine-tuned on approximately 500 pairs of prompts (inputs) and completions (outputs). Information is extracted either from single sentences or across sentences in abstracts/passages, and the output can be returned as simple English sentences or a more structured format, such as a list of JSON objects. We demonstrate that LLMs trained in this way are capable of accurately extracting useful records of complex scientific knowledge for three representative tasks in materials chemistry: linking dopants with their host materials, cataloging metal-organic frameworks, and general chemistry/phase/morphology/application information extraction. This approach represents a simple, accessible, and highly-flexible route to obtaining large databases of structured knowledge extracted from unstructured text. An online demo is available at http://www.matscholar.com/info-extraction.

cs.CL↗

Realizing the Data-Driven, Computational Discovery of Metal-Organic Framework Catalysts

Metal-organic frameworks (MOFs) have been widely investigated for challenging catalytic transformations due to their well-defined structures and high degree of synthetic tunability. These features, at least in principle, make MOFs ideally suited for a computational approach towards catalyst design and discovery. Nonetheless, the widespread use of data science and machine learning to accelerate the discovery of MOF catalysts has yet to be substantially realized. In this review, we provide an overview of recent work that sets the stage for future high-throughput computational screening and machine learning studies involving MOF catalysts. This is followed by a discussion of several challenges currently facing the broad adoption of data-centric approaches in MOF computational catalysis, and we share possible solutions that can help propel the field forward.

cond-mat.mtrl-sci↗