SearcharxivSearch

arXiv subjects

Raymundo Arroyave

Publications and source records attributed to Raymundo Arroyave.

At least 19 recordsLinked to original sources

Portfolio-Based Constrained Multi-Objective Bayesian Optimization for Materials Design

Materials discovery and design campaigns can be formulated as constrained multi-objective Bayesian optimization (CMOBO) problems, within which each experimental decision negotiates between two coupled but competing goals: discovering feasible candidates and refining the underlying Pareto front. Here we recast acquisition-function choice as an adaptive policy-selection problem over a portfolio of conventional and feasibility-focused acquisition functions. This was done using two controllers: UCB-Bandit, a modified UCB multi-armed bandit, and Agentic-Switch, a multi-agent decision system driven by a large language model (LLM). Both were evaluated against fixed-policy baselines in silico across five synthetic benchmark functions and two materials design case studies. The adaptive policies performed competitively in terms of both cumulative feasibility count and feasible hypervolume improvement, while each individual acquisition function performed well for only one metric, suggesting that adaptive policies are better suited for constrained materials science problems.

cond-mat.mtrl-sci

From MLIPs to Microstructure: A High-Throughput Computational Framework to Design Spinodal Alloys in High-Dimensional Composition Spaces via Analytic Derivatives of CALPHAD Model Predictions

Identifying regions of design space subject to spinodal decomposition is a critical component of alloy design in high-dimensional composition spaces. In cases where designers are seeking to exploit spinodal microstructures to tailor alloy properties, prediction of microstructure evolution and morphology is also needed. In this work, we present a Machine Learning Interatomic Potential (MLIP)-trained, CALPHAD-based, open-source workflow for high-throughput microstructure stability analysis and visualization. In this workflow, coherent strain contributions are captured via high-throughput MLIP elastic constant calculations. To predict microstructure morphology for compositions of interest, MLIP-generated thermodynamic models are fed into an elasto-chemical phase field simulation. Both stability analyses and phase-field simulations utilize analytically-derived Gibbs energy Hessians to improve computational efficiency and accuracy over finite difference approximations. We demonstrate this workflow by investigating microstructure stability in the Hf-Nb-Ti-V quaternary system.

cond-mat.mtrl-sci

Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design

Closed-loop materials discovery iterates between proposing candidate structures and evaluating their properties, and property evaluation dominates the cost. In the generative variant, a learned prior proposes candidate crystals and a property oracle scores them; we ask whether a cheap probabilistic surrogate can triage the generator's output, and what such a surrogate must do well. Across three architecturally distinct pretrained diffusion priors (MatterGen, CrystalFlow, ADiT) and two targets (room-temperature heat capacity and bulk modulus), we insert a Gaussian process acquisition gate between structure generation and the oracle in an RL-steered generative workflow. The gate matches or exceeds ungated fine-tuning of the generative model while capping oracle calls at a fixed per-cycle budget. Budget-matched ablations isolate the mechanism. At an identical four-call budget, ranking-based selection outperforms arbitrary selection, confirming that the gain comes from the surrogate's choice; the gate comes within $\sim$9\% of exhaustive oracle spending at roughly one-fifth of the calls. A density-functional-theory check of the bulk-modulus discoveries confirms the learned oracle to within 2.5\% on average and the surrogate's ranking of the generated structures at Spearman $ρ= 0.94$. A cross-factorial benchmark of surrogate performance spanning mechanical, electronic, and vibrational properties identifies pretrained ORB embeddings with a Gaussian process as the most reliable combination, which we adopt as the building blocks of the proposed workflow. The complete pipeline is released as open-source software.

cond-mat.mtrl-sci

Ground-State Structure Search of Defective High-Entropy Alloys Using Machine-Learning Potentials and Monte Carlo Sampling

Resolving the atomic-scale structure of defective high-entropy alloys (HEAs) containing interstitial species remains a major computational challenge due to the vast configurational space and the limitations of existing methods. Here we introduce PAIPAI (Package for Alloy Interstitial Predictions using Artificial Intelligence), a Monte Carlo framework coupled with machine-learning interatomic potentials (MLIPs) that searches for ground-state atomic configurations in HEAs with defects and interstitials. PAIPAI employs a dual-worker architecture-fast workers for rapid configurational screening and slow workers for high-accuracy refinement-coordinated through a shared waiting pool, enabling efficient parallel sampling. We demonstrate PAIPAI through three case studies: (i) surface segregation in a Ti-V-Cr-Re slab; (ii) interstitial oxygen and boron aggregation in bulk BCC Nb-Ti-Ta-Hf; and (iii) coupled metallic and interstitial segregation at grain boundaries in Nb-Ti-Ta-Hf. In all cases, Monte Carlo-optimized structures are significantly lower in energy than any configuration obtained by random sampling, and MLIP energy rankings are validated against density functional theory calculations. PAIPAI provides a general and efficient framework for predicting atomic ordering, segregation, and interstitial behavior in complex, defective HEA systems.

cond-mat.mtrl-sci

DataScribe: An AI-Native, Policy-Aligned Web Platform for Multi-Objective Materials Design and Discovery

The acceleration of materials discovery requires digital platforms that go beyond data repositories to embed learning, optimization, and decision-making directly into research workflows. We introduce DataScribe, an AI-native, cloud-based materials discovery platform that unifies heterogeneous experimental and computational data through ontology-backed ingestion and machine-actionable knowledge graphs. The platform integrates FAIR-compliant metadata capture, schema and unit harmonization, uncertainty-aware surrogate modeling, and native multi-objective multi-fidelity Bayesian optimization, enabling closed-loop propose-measure-learn workflows across experimental and computational pipelines. DataScribe functions as an application-layer intelligence stack, coupling data governance, optimization, and explainability rather than treating them as downstream add-ons. We validate the platform through case studies in electrochemical materials and high-entropy alloys, demonstrating end-to-end data fusion, real-time optimization, and reproducible exploration of multi-objective trade spaces. By embedding optimization engines, machine learning, and unified access to public and private scientific data directly within the data infrastructure, and by supporting open, free use for academic and non-profit researchers, DataScribe functions as a general-purpose application-layer backbone for laboratories of any scale, including self-driving laboratories and geographically distributed materials acceleration platforms, with built-in support for performance, sustainability, and supply-chain-aware objectives.

cs.LG

Deep Gaussian Process-based Cost-Aware Batch Bayesian Optimization for Complex Materials Design Campaigns

The accelerating pace and expanding scope of materials discovery demand optimization frameworks that efficiently navigate vast, nonlinear design spaces while judiciously allocating limited evaluation resources. We present a cost-aware, batch Bayesian optimization scheme powered by deep Gaussian process (DGP) surrogates and a heterotopic querying strategy. Our DGP surrogate, formed by stacking GP layers, models complex hierarchical relationships among high-dimensional compositional features and captures correlations across multiple target properties, propagating uncertainty through successive layers. We integrate evaluation cost into an upper-confidence-bound acquisition extension, which, together with heterotopic querying, proposes small batches of candidates in parallel, balancing exploration of under-characterized regions with exploitation of high-mean, low-variance predictions across correlated properties. Applied to refractory high-entropy alloys for high-temperature applications, our framework converges to optimal formulations in fewer iterations with cost-aware queries than conventional GP-based BO, highlighting the value of deep, uncertainty-aware, cost-sensitive strategies in materials campaigns.

cond-mat.mtrl-sci

Data Driven Insights into Composition Property Relationships in FCC High Entropy Alloys

Structural High Entropy Alloys (HEAs) are crucial in advancing technology across various sectors, including aerospace, automotive, and defense industries. However, the scarcity of integrated chemistry, process, structure, and property data presents significant challenges for predictive property modeling. Given the vast design space of these alloys, uncovering the underlying patterns is essential yet difficult, requiring advanced methods capable of learning from limited and heterogeneous datasets. This work presents several sensitivity analyses, highlighting key elemental contributions to mechanical behavior, including insights into the compositional factors associated with brittle and fractured responses observed during nanoindentation testing in the BIRDSHOT center NiCoFeCrVMnCuAl system dataset. Several encoder decoder based chemistry property models, carefully tuned through Bayesian multi objective hyperparameter optimization, are evaluated for mapping alloy composition to six mechanical properties. The models achieve competitive or superior performance to conventional regressors across all properties, particularly for yield strength and the UTS/YS ratio, demonstrating their effectiveness in capturing complex composition property relationships.

cond-mat.mtrl-sci

Construction and Tuning of CALPHAD Models Using Machine-Learned Interatomic Potentials and Experimental Data: A Case Study of the Pt-W System

This work introduces PhaseForgePlus -- a computationally efficient, fully open-source workflow for physically-informed CALPHAD model generation and parameter fitting. Using the Pt-W system as an example, we show that the integration of Machine Learning Potentials into the Alloy Theoretic Automated Toolkit can produce physically grounded Gibbs energy descriptions requiring only slight adjustments to produce accurate phase diagrams. Employing the Jansson derivative method in the context of experimental observations, such adjustments can be efficiently and robustly determined through gradient-informed optimization procedures.

cond-mat.mtrl-sci

Machine Learning Potentials for Alloys: A Detailed Workflow to Predict Phase Diagrams and Benchmark Accuracy

High-entropy alloys (HEAs) have attracted increasing attention due to their unique structural and functional properties. In the study of HEAs, thermodynamic properties and phase stability play a crucial role, making phase diagram calculations significantly important. However, phase diagram calculations with conventional CALPHAD assessments based on experimental or ab-initio data can be expensive. With the emergence of machine-learning interatomic potentials (MLIPs), we have developed a program named PhaseForge, which integrates MLIPs into the Alloy Theoretic Automated Toolkit (ATAT) framework using our MLIP calculation library, MaterialsFramework, to enable efficient exploration of alloy phase diagrams. Moreover, our workflow can also serve as a benchmarking tool for evaluating the quality of different MLIPs.

cond-mat.mtrl-sci

Accurate and Uncertainty-Aware Multi-Task Prediction of HEA Properties Using Prior-Guided Deep Gaussian Processes

Surrogate modeling techniques have become indispensable in accelerating the discovery and optimization of high-entropy alloys(HEAs), especially when integrating computational predictions with sparse experimental observations. This study systematically evaluates the fitting performance of four prominent surrogate models conventional Gaussian Processes(cGP), Deep Gaussian Processes(DGP), encoder-decoder neural networks for multi-output regression and XGBoost applied to a hybrid dataset of experimental and computational properties in the AlCoCrCuFeMnNiV HEA system. We specifically assess their capabilities in predicting correlated material properties, including yield strength, hardness, modulus, ultimate tensile strength, elongation, and average hardness under dynamic and quasi-static conditions, alongside auxiliary computational properties. The comparison highlights the strengths of hierarchical and deep modeling approaches in handling heteroscedastic, heterotopic, and incomplete data commonly encountered in materials informatics. Our findings illustrate that DGP infused with machine learning-based prior outperform other surrogates by effectively capturing inter-property correlations and input-dependent uncertainty. This enhanced predictive accuracy positions advanced surrogate models as powerful tools for robust and data-efficient materials design.

cs.LG

A Materials Foundation Model via Hybrid Invariant-Equivariant Architectures

Machine learning interatomic potentials (MLIPs) can predict energy, force, and stress of materials and enable a wide range of downstream discovery tasks. A key design choice in MLIPs involves the trade-off between invariant and equivariant architectures. Invariant models offer computational efficiency but may not perform as well, especially when predicting high-order outputs. In contrast, equivariant models can capture high-order symmetries, but are computationally expensive. In this work, we propose HIENet, a hybrid invariant-equivariant materials interatomic potential model that integrates both invariant and equivariant message passing layers, while provably satisfying key physical constraints. HIENet achieves state-of-the-art performance with considerable computational speedups over prior models. Experimental results on both common benchmarks and downstream materials discovery tasks demonstrate the efficiency and effectiveness of HIENet.

cs.LG

Analytical Gradient-Based Optimization of CALPHAD Model Parameters

The calibration of CALPHAD (CALculation of PHAse Diagrams) models involves the solution of a very challenging high-dimensional multiobjective optimization problem. Traditional approaches to parameter fitting predominantly rely on gradient-free methods, which while robust, are computationally inefficient and often scale poorly with model complexity. In this work, we introduce and demonstrate a generalizable framework for analytic gradient-based optimization of the parameters of the CALPHAD model enabled by the recently formalized Jansson derivative technique. This method allows for efficient evaluation of gradients of thermodynamic properties at equilibrium with respect to model parameters, even in the presence of arbitrarily complex internal degrees of freedom. Leveraging these semi-analytic gradients, we employ the conjugate gradient (CG) method to optimize thermodynamic model parameters for four binary alloy systems: Cu-Mg, Fe-Ni, Cr-Ni, and Cr-Fe. Across all systems, CG achieves comparable or superior optimality relative to Bayesian ensemble Markov Chain Monte Carlo (MCMC) with improvements in computational efficiency ranging from one to three orders of magnitude. Our results establish a new paradigm for CALPHAD assessments in which high fidelity data-rich model calibration becomes tractable using deterministic gradient-informed algorithms.

cond-mat.mtrl-sci

Accelerated Multi-Objective Alloy Discovery through Efficient Bayesian Methods: Application to the FCC Alloy Space

This study introduces BIRDSHOT, an integrated Bayesian materials discovery framework designed to efficiently explore complex compositional spaces while optimizing multiple material properties. We applied this framework to the CoCrFeNiVAl FCC high entropy alloy (HEA) system, targeting three key performance objectives: ultimate tensile strength/yield strength ratio, hardness, and strain rate sensitivity. The experimental campaign employed an integrated cyber-physical approach that combined vacuum arc melting (VAM) for alloy synthesis with advanced mechanical testing, including tensile and high-strain-rate nanoindentation testing. By incorporating batch Bayesian optimization schemes that allowed the parallel exploration of the alloy space, we completed five iterative design-make-test-learn loops, identifying a non-trivial three-objective Pareto set in a high-dimensional alloy space. Notably, this was achieved by exploring only 0.15% of the feasible design space, representing a significant acceleration in discovery rate relative to traditional methods. This work demonstrates the capability of BIRDSHOT to navigate complex, multi-objective optimization challenges and highlights its potential for broader application in accelerating materials discovery.

cond-mat.mtrl-sci

Physics-Informed Gaussian Process Classification for Constraint-Aware Alloy Design

Alloy design can be framed as a constraint-satisfaction problem. Building on previous methodologies, we propose equipping Gaussian Process Classifiers (GPCs) with physics-informed prior mean functions to model the boundaries of feasible design spaces. Through three case studies, we highlight the utility of informative priors for handling constraints on continuous and categorical properties. (1) Phase Stability: By incorporating CALPHAD predictions as priors for solid-solution phase stability, we enhance model validation using a publicly available XRD dataset. (2) Phase Stability Prediction Refinement: We demonstrate an in silico active learning approach to efficiently correct phase diagrams. (3) Continuous Property Thresholds: By embedding priors into continuous property models, we accelerate the discovery of alloys meeting specific property thresholds via active learning. In each case, integrating physics-based insights into the classification framework substantially improved model performance, demonstrating an efficient strategy for constraint-aware alloy design.

cond-mat.mtrl-sci

Compositionally Grading Alloy Stacking Fault Energy using Autonomous Path Planning and Additive Manufacturing with Elemental Powders

Compositionally graded alloys (CGAs) are often proposed for use in structural components where the combination of two or more alloys within a single part can yield substantial enhancement in performance and functionality. For these applications, numerous design methodologies have been developed, one of the most sophisticated being the application of path planning algorithms originally designed for robotics to solve CGA design problems. In addition to the traditional application to structural components, this work proposes and demonstrates the application of this CGA design framework to rapid alloy design, synthesis, and characterization. A composition gradient in the CoCrFeNi alloy space was planned between the maximum and minimum stacking fault energy (SFE) as predicted by a previously developed model in a face-centered cubic (FCC) high entropy alloy (HEA) space. The path was designed to be monotonic in SFE and avoid regions that did not meet FCC phase fraction and solidification range constraints predicted by CALculation of PHAse Diagrams (CALPHAD). Compositions from the path were selected to produce a linear gradient in SFE, and the CGA was built using laser directed energy deposition (L-DED). The resulting gradient was characterized for microstructure and mechanical properties, including hardness, elastic modulus, and strain rate sensitivity. Despite being predicted to contain a single FCC phase throughout the gradient, part of the CGA underwent a martensitic transformation, thereby demonstrating a limitation of using equilibrium CALPHAD calculations for phase stability predictions. More broadly, this demonstrates the ability of the methods employed to bring attention to blind spots in alloy models.

cond-mat.mtrl-sci

Towards Autonomous Experimentation: Bayesian Optimization over Problem Formulation Space for Accelerated Alloy Development

Accelerated discovery in materials science demands autonomous systems capable of dynamically formulating and solving design problems. In this work, we introduce a novel framework that leverages Bayesian optimization over a problem formulation space to identify optimal design formulations in line with decision-maker preferences. By mapping various design scenarios to a multi attribute utility function, our approach enables the system to balance conflicting objectives such as ductility, yield strength, density, and solidification range without requiring an exact problem definition at the outset. We demonstrate the efficacy of our method through an in silico case study on a Mo-Nb-Ti-V-W alloy system targeted for gas turbine engine blade applications. The framework converges on a sweet spot that satisfies critical performance thresholds, illustrating that integrating problem formulation discovery into the autonomous design loop can significantly streamline the experimental process. Future work will incorporate human feedback to further enhance the adaptability of the system in real-world experimental settings.

eess.SY

Microstructure-Aware Bayesian Materials Design

In this study, we propose a novel microstructure-sensitive Bayesian optimization (BO) framework designed to enhance the efficiency of materials discovery by explicitly incorporating microstructural information. Traditional materials design approaches often focus exclusively on direct chemistry-process-property relationships, overlooking the critical role of microstructures. To address this limitation, our framework integrates microstructural descriptors as latent variables, enabling the construction of a comprehensive process-structure-property mapping that improves both predictive accuracy and optimization outcomes. By employing the active subspace method for dimensionality reduction, we identify the most influential microstructural features, thereby reducing computational complexity while maintaining high accuracy in the design process. This approach also enhances the probabilistic modeling capabilities of Gaussian processes, accelerating convergence to optimal material configurations with fewer iterations and experimental observations. We demonstrate the efficacy of our framework through synthetic and real-world case studies, including the design of Mg$_2$Sn$_x$Si$_{1-x}$ thermoelectric materials for energy conversion. Our results underscore the critical role of microstructures in linking processing conditions to material properties, highlighting the potential of a microstructure-aware design paradigm to revolutionize materials discovery. Furthermore, this work suggests that since incorporating microstructure awareness improves the efficiency of Bayesian materials discovery, microstructure characterization stages should be integral to automated -- and eventually autonomous -- platforms for materials development.

cond-mat.mtrl-sci

Decoding Non-Linearity and Complexity: Deep Tabular Learning Approaches for Materials Science

Materials data, especially those related to high-temperature properties, pose significant challenges for machine learning models due to extreme skewness, wide feature ranges, modality, and complex relationships. While traditional models like tree-based ensembles (e.g., XGBoost, LightGBM) are commonly used for tabular data, they often struggle to fully capture the subtle interactions inherent in materials science data. In this study, we leverage deep learning techniques based on encoder-decoder architectures and attention-based models to handle these complexities. Our results demonstrate that XGBoost achieves the best loss value and the fastest trial duration, but deep encoder-decoder learning like Disjunctive Normal Form architecture (DNF-nets) offer competitive performance in capturing non-linear relationships, especially for highly skewed data distributions. However, convergence rates and trial durations for deep model such as CNN is slower, indicating areas for further optimization. The models introduced in this study offer robust and hybrid solutions for enhancing predictive accuracy in complex materials datasets.

cond-mat.mtrl-sci