SearcharxivSearch

arXiv subjects

Brian DeCost

Publications and source records attributed to Brian DeCost.

At least 19 recordsLinked to original sources

Bridging Atomistic Simulation and Experimental Processing Timescales with Goal-Directed Deep Reinforcement Learning

Atomic-scale modeling has advanced rapidly through integration of machine learning, yet a key bottleneck remains. Even with an accurate potential energy surface and a clear target material, we still lack a practical atomistic dynamics framework that can simulate how materials form under realistic synthesis and processing conditions. Many processing transformations are governed by rare events in non-idealized evolving environments, while direct molecular dynamics is limited by femtosecond timesteps and short accessible trajectories. Existing acceleration methods often require prior mechanistic knowledge, including reaction coordinates, collective variables, event tables, or pathway guesses, which is rarely available in real experiments. Here we present an E(3)-equivariant deep reinforcement learning framework that enables goal-directed pathway discovery without hand-crafted reaction coordinates. The framework introduces a complementary operating mode for atomistic simulation in which realistic, non-idealized environments can be addressed directly while retaining kinetic plausibility through barrier-aware rewards. As a challenging benchmark, we target silicon dry oxidation, where rare-event pathways in amorphous SiO2 are effectively inaccessible to conventional atomistic methods. We treat an O2 molecule as an agent that performs continuous rigid-body translations and rotations in a Si/a-SiO2 environment. The agent is trained with an episode-level objective that rewards verified O2 dissociation while preferring low effective activation barriers. We demonstrate that the learned policy discovers kinetically favorable O2 diffusion and dissociation pathways in a disordered Si/a-SiO2 environment, progressively improving success rate while reducing effective activation barriers over training. We also discuss how the approach can be generalized to other processing and synthesis problems.

cond-mat.mtrl-sci

Materials Acceleration Platform for Electrochemistry: a Platform for Autonomous Electrochemistry

Corrosion testing is slow, labor-intensive, and sensitive to operator technique, limiting the generation of large, high-quality datasets for data-driven materials discovery. The Materials Acceleration Platform for Electrochemistry (MAP-E) is an autonomous, high-throughput system, capable of performing parallel electrochemical experiments. It integrates robotic liquid handling, sample transfer with a multi-channel potentiostatic control to extract corrosion metrics without human intervention. Validation against an ASTM G61-analog benchmark demonstrates good reproducibility, with a standard deviation of 75 mV in pitting potential across 32 automated measurements. The platform was then employed to autonomously construct pH-chloride stability diagrams for 304 stainless steel using an uncertainty-driven sampling strategy on a Gaussian process surrogate model. This approach reduces operator involvement and accelerates the exploration of environmental spaces. The MAP-E establishes a framework for autonomous electrochemical experimentation, enabling generation of corrosion datasets that inform materials discovery, alloy design, and durability assessment in service environments.

cond-mat.mtrl-sci

When Active Learning Fails, Uncalibrated Out of Distribution Uncertainty Quantification Might Be the Problem

Efficiently and meaningfully estimating prediction uncertainty is important for exploration in active learning campaigns in materials discovery, where samples with high uncertainty are interpreted as containing information missing from the model. In this work, the effect of different uncertainty estimation and calibration methods are evaluated for active learning when using ensembles of ALIGNN, eXtreme Gradient Boost, Random Forest, and Neural Network model architectures. We compare uncertainty estimates from ALIGNN deep ensembles to loss landscape uncertainty estimates obtained for solubility, bandgap, and formation energy prediction tasks. We then evaluate how the quality of the uncertainty estimate impacts an active learning campaign that seeks model generalization to out-of-distribution data. Uncertainty calibration methods were found to variably generalize from in-domain data to out-of-domain data. Furthermore, calibrated uncertainties were generally unsuccessful in reducing the amount of data required by a model to improve during an active learning campaign on out-of-distribution data when compared to random sampling and uncalibrated uncertainties. The impact of poor-quality uncertainty persists for random forest and eXtreme Gradient Boosting models trained on the same data for the same tasks, indicating that this is at least partially intrinsic to the data and not due to model capacity alone. Analysis of the target, in-distribution uncertainty, out-of-distribution uncertainty, and training residual distributions suggest that future work focus on understanding empirical uncertainties in the feature input space for cases where ensemble prediction variances do not accurately capture the missing information required for the model to generalize.

cond-mat.mtrl-sci

Intrinsic Direct Air Capture

We present new metrics to evaluate solid sorbent materials for Direct Air Capture (DAC). These new metrics provide a theoretical upper bound on CO2 captured per energy as well as a theoretical upper limit on the purity of the captured CO2. These new metrics are based entirely on intrinsic material properties and are therefore agnostic to the design of the DAC system. These metrics apply to any adsorption-refresh cycle design. In this work we demonstrate the use of these metrics with the example of temperature-pressure swing refresh cycles. The main requirement for applying these metrics is to describe the equilibrium uptake (along with a few other materials properties) of each species in terms of the thermodynamic variables (e.g. temperature, pressure). We derive these metrics from thermodynamic energy balances. To apply these metrics on a set of examples, we first generated approximations of the necessary materials properties for 11 660 metal-organic framework materials (MOFs). We find that the performance of the sorbents is highly dependent on the path through thermodynamic parameter space. These metrics allow for: 1) finding the optimum materials given a particular refresh cycle, and 2) finding the optimum refresh cycles given a particular sorbent. Applying these metrics to the database of MOFs lead to the following insights: 1) start cold - the equilibrium uptake of CO2 diverges from that of N2 at lower temperatures, and 2) selectivity of CO2 vs other gases at any one point in the cycle does not matter - what matters is the relative change in uptake along the cycle.

cond-mat.mtrl-sci

Probing out-of-distribution generalization in machine learning for materials

Scientific machine learning (ML) endeavors to develop generalizable models with broad applicability. However, the assessment of generalizability is often based on heuristics. Here, we demonstrate in the materials science setting that heuristics based evaluations lead to substantially biased conclusions of ML generalizability and benefits of neural scaling. We evaluate generalization performance in over 700 out-of-distribution tasks that features new chemistry or structural symmetry not present in the training data. Surprisingly, good performance is found in most tasks and across various ML models including simple boosted trees. Analysis of the materials representation space reveals that most tasks contain test data that lie in regions well covered by training data, while poorly-performing tasks contain mainly test data outside the training domain. For the latter case, increasing training set size or training time has marginal or even adverse effects on the generalization performance, contrary to what the neural scaling paradigm assumes. Our findings show that most heuristically-defined out-of-distribution tests are not genuinely difficult and evaluate only the ability to interpolate. Evaluating on such tasks rather than the truly challenging ones can lead to an overestimation of generalizability and benefits of scaling.

cond-mat.mtrl-sci

Efficient first principles based modeling via machine learning: from simple representations to high entropy materials

High-entropy materials (HEMs) have recently emerged as a significant category of materials, offering highly tunable properties. However, the scarcity of HEM data in existing density functional theory (DFT) databases, primarily due to computational expense, hinders the development of effective modeling strategies for computational materials discovery. In this study, we introduce an open DFT dataset of alloys and employ machine learning (ML) methods to investigate the material representations needed for HEM modeling. Utilizing high-throughput DFT calculations, we generate a comprehensive dataset of 84k structures, encompassing both ordered and disordered alloys across a spectrum of up to seven components and the entire compositional range. We apply descriptor-based models and graph neural networks to assess how material information is captured across diverse chemical-structural representations. We first evaluate the in-distribution performance of ML models to confirm their predictive accuracy. Subsequently, we demonstrate the capability of ML models to generalize between ordered and disordered structures, between low-order and high-order alloys, and between equimolar and non-equimolar compositions. Our findings suggest that ML models can generalize from cost-effective calculations of simpler systems to more complex scenarios. Additionally, we discuss the influence of dataset size and reveal that the information loss associated with the use of unrelaxed structures could significantly degrade the generalization performance. Overall, this research sheds light on several critical aspects of HEM modeling and offers insights for data-driven atomistic modeling of HEMs.

cond-mat.mtrl-sci

Leveraging Domain Adaptation for Accurate Machine Learning Predictions of New Halide Perovskites

We combine graph neural networks (GNN) with an inexpensive and reliable structure generation approach based on the bond-valence method (BVM) to train accurate machine learning models for screening 222,960 halide perovskites using statistical estimates of the DFT/PBE formation energy (Ef), and the PBE and HSE band gaps (Eg). The GNNs were fined tuned using domain adaptation (DA) from a source model, which yields a factor of 1.8 times improvement in Ef and 1.2 - 1.35 times improvement in HSE Eg compared to direct training (i.e., without DA). Using these two ML models, 48 compounds were identified out of 222,960 candidates as both stable and that have an HSE Eg that is relevant for photovoltaic applications. For this subset, only 8 have been reported to date, indicating that 40 compounds remain unexplored to the best of our knowledge and therefore offer opportunities for potential experimental examination.

cond-mat.mtrl-sci

Learning material synthesis-process-structure-property relationship by data fusion: Bayesian Coregionalization N-Dimensional Piecewise Function Learning

Autonomous materials research labs require the ability to combine and learn from diverse data streams. This is especially true for learning material synthesis-process-structure-property relationships, key to accelerating materials optimization and discovery as well as accelerating mechanistic understanding. We present the Synthesis-process-structure-property relAtionship coreGionalized lEarner (SAGE) algorithm. A fully Bayesian algorithm that uses multimodal coregionalization to merge knowledge across data sources to learn synthesis-process-structure-property relationships. SAGE outputs a probabilistic posterior for the relationships including the most likely relationships given the data.

cs.LG

Recent progress in the JARVIS infrastructure for next-generation data-driven materials design

The Joint Automated Repository for Various Integrated Simulations (JARVIS) infrastructure at the National Institute of Standards and Technology (NIST) is a large-scale collection of curated datasets and tools with more than 80000 materials and millions of properties. JARVIS uses a combination of electronic structure, artificial intelligence (AI), advanced computation and experimental methods to accelerate materials design. Here we report some of the new features that were recently included in the infrastructure such as: 1) doubling the number of materials in the database since its first release, 2) including more accurate electronic structure methods such as Quantum Monte Carlo, 3) including graph neural network-based materials design, 4) development of unified force-field, 5) development of a universal tight-binding model, 6) addition of computer-vision tools for advanced microscopy applications, 7) development of a natural language processing tool for text-generation and analysis, 8) debuting a large-scale benchmarking endeavor, 9) including quantum computing algorithms for solids, 10) integrating several experimental datasets and 11) staging several community engagement and outreach events. New classes of materials, properties, and workflows added to the database include superconductors, two-dimensional (2D) magnets, magnetic topological materials, metal-organic frameworks, defects, and interface systems. The rich and reliable datasets, tools, documentation, and tutorials make JARVIS a unique platform for modern materials design. JARVIS ensures openness of data and tools to enhance reproducibility and transparency and to promote a healthy and collaborative scientific environment.

cond-mat.mtrl-sci

Approaches for Uncertainty Quantification of AI-predicted Material Properties: A Comparison

The development of large databases of material properties, together with the availability of powerful computers, has allowed machine learning (ML) modeling to become a widely used tool for predicting material performances. While confidence intervals are commonly reported for such ML models, prediction intervals, i.e., the uncertainty on each prediction, are not as frequently available. Here, we investigate three easy-to-implement approaches to determine such individual uncertainty, comparing them across ten ML quantities spanning energetics, mechanical, electronic, optical, and spectral properties. Specifically, we focused on the Quantile approach, the direct machine learning of the prediction intervals and Ensemble methods.

cond-mat.mtrl-sci

Accelerating Defect Predictions in Semiconductors Using Graph Neural Networks

Here, we develop a framework for the prediction and screening of native defects and functional impurities in a chemical space of Group IV, III-V, and II-VI zinc blende (ZB) semiconductors, powered by crystal Graph-based Neural Networks (GNNs) trained on high-throughput density functional theory (DFT) data. Using an innovative approach of sampling partially optimized defect configurations from DFT calculations, we generate one of the largest computational defect datasets to date, containing many types of vacancies, self-interstitials, anti-site substitutions, impurity interstitials and substitutions, as well as some defect complexes. We applied three types of established GNN techniques, namely Crystal Graph Convolutional Neural Network (CGCNN), Materials Graph Network (MEGNET), and Atomistic Line Graph Neural Network (ALIGNN), to rigorously train models for predicting defect formation energy (DFE) in multiple charge states and chemical potential conditions. We find that ALIGNN yields the best DFE predictions with root mean square errors around 0.3 eV, which represents a prediction accuracy of 98 % given the range of values within the dataset, improving significantly on the state-of-the-art. Models are tested for different defect types as well as for defect charge transition levels. We further show that GNN-based defective structure optimization can take us close to DFT-optimized geometries at a fraction of the cost of full DFT. DFT-GNN models enable prediction and screening across thousands of hypothetical defects based on both unoptimized and partially-optimized defective structures, helping identify electronically active defects in technologically-important semiconductors.

cond-mat.mtrl-sci

Emulating Expert Insight: A Robust Strategy for Optimal Experimental Design

The challenge of optimal design of experiments (DOE) pervades materials science, physics, chemistry, and biology. Bayesian optimization has been used to address this challenge in vast sample spaces, although it requires framing experimental campaigns through the lens of maximizing some observable. This framing is insufficient for epistemic research goals that seek to comprehensively analyze a sample space, without an explicit scalar objective (e.g., the characterization of a wafer or sample library). In this work, we propose a flexible formulation of scientific value that recasts a dataset of input conditions and higher-dimensional observable data into a continuous, scalar metric. Intuitively, the scientific value function measures where observables change significantly, emulating the perspective of experts driving an experiment, and can be used in collaborative analysis tools or as an objective for optimization techniques. We demonstrate this technique by exploring simulated phase boundaries from different observables, autonomously driving a variable temperature measurement of a ferroelectric material, and providing feedback from a nanoparticle synthesis campaign. The method is seamlessly compatible with existing optimization tools, can be extended to multi-modal and multi-fidelity experiments, and can integrate existing models of an experimental system. Because of its flexibility, it can be deployed in a range of experimental settings for autonomous or accelerated experiments.

cond-mat.mtrl-sci

On the redundancy in large material datasets: efficient and robust learning with less data

Extensive efforts to gather materials data have largely overlooked potential data redundancy. In this study, we present evidence of a significant degree of redundancy across multiple large datasets for various material properties, by revealing that up to 95 % of data can be safely removed from machine learning training with little impact on in-distribution prediction performance. The redundant data is related to over-represented material types and does not mitigate the severe performance degradation on out-of-distribution samples. In addition, we show that uncertainty-based active learning algorithms can construct much smaller but equally informative datasets. We discuss the effectiveness of informative data in improving prediction performance and robustness and provide insights into efficient data acquisition and machine learning training. This work challenges the "bigger is better" mentality and calls for attention to the information richness of materials data rather than a narrow emphasis on data volume.

cond-mat.mtrl-sci

AutoEIS: automated Bayesian model selection and analysis for electrochemical impedance spectroscopy

Electrochemical Impedance Spectroscopy (EIS) is a powerful tool for electrochemical analysis; however, its data can be challenging to interpret. Here, we introduce a new open-source tool named AutoEIS that assists EIS analysis by automatically proposing statistically plausible equivalent circuit models (ECMs). AutoEIS does this without requiring an exhaustive mechanistic understanding of the electrochemical systems. We demonstrate the generalizability of AutoEIS by using it to analyze EIS datasets from three distinct electrochemical systems, including thin-film oxygen evolution reaction (OER) electrocatalysis, corrosion of self-healing multi-principal components alloys, and a carbon dioxide reduction electrolyzer device. In each case, AutoEIS identified competitive or in some cases superior ECMs to those recommended by experts and provided statistical indicators of the preferred solution. The results demonstrated AutoEIS's capability to facilitate EIS analysis without expert labels while diminishing user bias in a high-throughput manner. AutoEIS provides a generalized automated approach to facilitate EIS analysis spanning a broad suite of electrochemical applications with minimal prior knowledge of the system required. This tool holds great potential in improving the efficiency, accuracy, and ease of EIS analysis and thus creates an avenue to the widespread use of EIS in accelerating the development of new electrochemical materials and devices.

cond-mat.mtrl-sci

Why is EXAFS analysis for multicomponent metals so hard? Challenges and opportunities for measuring ordering in complex concentrated alloys using x-ray absorption spectroscopy

Short range order is a critical driver of properties (e.g. corrosion resistance and tensile strength) in multicomponent alloys such as complex concentrated alloys (CCAs). Extended x-ray absorption fine structure (EXAFS) is a powerful technique well suited for quantifying this short range order.Here, we described in detail the characteristics of CCAs that make the already challenging task of analyzing EXAFS data even more difficult. We then illustrate novel paths towards robust and scalable quantitative SRO analysis which will accelerate the scientific understanding and development of CCAs.

cond-mat.mtrl-sci

AtomVision: A machine vision library for atomistic images

Computer vision techniques have immense potential for materials design applications. In this work, we introduce an integrated and general-purpose AtomVision library that can be used to generate, curate scanning tunneling microscopy (STM) and scanning transmission electron microscopy (STEM) datasets and apply machine learning techniques. To demonstrate the applicability of this library, we 1) generate and curate an atomistic image dataset of about 10000 materials, 2) develop and compare convolutional and graph neural network models to classify the Bravais lattices, 3) develop fully convolutional neural network using U-Net architecture to pixelwise classify atom vs background, 4) use generative adversarial network for super-resolution, 5) curate a natural language processing based image dataset using open-access arXiv dataset, and 6) integrate the computational framework with experimental microscopy tools. AtomVision library is available at https://github.com/usnistgov/atomvision.

cond-mat.mtrl-sci

A critical examination of robustness and generalizability of machine learning prediction of materials properties

Recent advances in machine learning (ML) methods have led to substantial improvement in materials property prediction against community benchmarks, but an excellent benchmark score may not imply good generalization of performance. Here we show that ML models trained on the Materials Project 2018 (MP18) dataset can have severely degraded prediction performance on new compounds in the Materials Project 2021 (MP21) dataset. We document performance degradation in graph neural networks and traditional descriptor-based ML models for both quantitative and qualitative predictions. We find the source of the predictive degradation is due to the distribution shift between the MP18 and MP21 versions. This is revealed by the uniform manifold approximation and projection (UMAP) of the feature space. We then show that the performance degradation issue can be foreseen using a few simple tools. Firstly, the UMAP can be used to investigate the connectivity and relative proximity of the training and test data within feature space. Secondly, the disagreement between multiple ML models on the test data can illuminate out-of-distribution samples. We demonstrate that the simple yet efficient UMAP-guided and query-by-committee acquisition strategies can greatly improve prediction accuracy through adding only 1~\% of the test data. We believe this work provides valuable insights for building materials databases and ML models that enable better prediction robustness and generalizability.

cond-mat.mtrl-sci

Unified Graph Neural Network Force-field for the Periodic Table

Classical force fields (FF) based on machine learning (ML) methods show great potential for large scale simulations of materials. MLFFs have hitherto largely been designed and fitted for specific systems and are not usually transferable to chemistries beyond the specific training set. We develop a unified atomisitic line graph neural network-based FF (ALIGNN-FF) that can model both structurally and chemically diverse materials with any combination of 89 elements from the periodic table. To train the ALIGNN-FF model, we use the JARVIS-DFT dataset which contains around 75000 materials and 4 million energy-force entries, out of which 307113 are used in the training. We demonstrate the applicability of this method for fast optimization of atomic structures in the crystallography open database and by predicting accurate crystal structures using genetic algorithm for alloys.

cond-mat.mtrl-sci