SearcharxivSearch

arXiv subjects

Lucas Foppa

Publications and source records attributed to Lucas Foppa.

12 recordsLinked to original sources

Unveiling the Core of Materials Properties via SISSO and Sensitivity Analysis

Interpretable AI can reveal physical principles governing intricate materials properties by uncovering explicit relationships between physical parameters and target properties. The sure-independence screening and sparsifying operator (SISSO) symbolic-regression approach identifies analytical expressions that correlate a target property with a small set of parameters, termed materials genes, selected from a large pool of candidates. However, multiple gene combinations can yield equally accurate SISSO models, with individual genes contributing with different weights. Here, we establish a derivative-based sensitivity analysis that resolves the non-uniqueness of symbolic-regression descriptions, enhances interpretability, thereby enabling deeper physical insight. This analysis reveals how distinct gene combinations encode equivalent information and identifies valence orbital radii, nuclear charges, and their products as the key quantities governing the equilibrium lattice constant of perovskites.

cond-mat.mtrl-sci

A Critical Examination of Active Learning Workflows in Materials Science

Active learning (AL) plays a critical role in materials science, enabling applications such as the construction of machine-learning interatomic potentials for atomistic simulations and the operation of self-driving laboratories. Despite its widespread use, the reliability and effectiveness of AL workflows depend on implicit design assumptions that are rarely examined systematically. Here, we critically assess AL workflows deployed in materials science and investigate how key design choices, such as surrogate models, sampling strategies, uncertainty quantification and evaluation metrics, relate to their performance. By identifying common pitfalls and discussing practical mitigation strategies, we provide guidance to practitioners for the efficient design, assessment, and interpretation of AL workflows in materials science.

cond-mat.mtrl-sci

Roadmap on Advancements of the FHI-aims Software Package

Electronic-structure theory is the foundation of the description of materials including multiscale modeling of their properties and functions. Obviously, without sufficient accuracy at the base, reliable predictions are unlikely at any level that follows. The software package FHI-aims has proven to be a game changer for accurate free-energy calculations because of its scalability, numerical precision, and its efficient handling of density functional theory (DFT) with hybrid functionals and van der Waals interactions. It treats molecules, clusters, and extended systems (solids and liquids) on an equal footing. Besides DFT, FHI-aims also includes quantum-chemistry methods, descriptions for excited states and vibrations, and calculations of various types of transport. Recent advancements address the integration of FHI-aims into an increasing number of workflows and various artificial intelligence (AI) methods. This Roadmap describes the state-of-the-art of FHI-aims and advancements that are currently ongoing or planned.

cond-mat.mtrl-sci

Materials Database from All-electron Hybrid Functional DFT Calculations

Materials databases built from calculations based on density functional approximations play an important role in the discovery of materials with improved properties. Most databases thus constructed rely on the generalized gradient approximation (GGA) for electron exchange and correlation. This limits the reliability of these databases, as well as the artificial intelligence (AI) models trained on them, for certain classes of materials and properties which are not well described by GGA. In this paper, we describe a database of 7,024 inorganic materials presenting diverse structures and compositions generated using hybrid functional calculations enabled by their efficient implementation in the all-electron code FHI-aims. The database is used to evaluate the thermodynamic and electrochemical stability of oxides relevant to catalysis and energy related applications. We illustrate how the database can be used to train AI models for material properties using the sure-independence screening and sparsifying operator (SISSO) approach.

cond-mat.mtrl-sci

Materials-Discovery Workflows Guided by Symbolic Regression: Identifying Acid-Stable Oxides for Electrocatalysis

The efficiency of active learning (AL) approaches to identify materials with desired properties relies on the knowledge of a few parameters describing the property. However, these parameters are unknown if the property is governed by a high intricacy of many atomistic processes. Here, we develop an AL workflow based on the sure-independence screening and sparsifying operator (SISSO) symbolic-regression approach. SISSO identifies the few, key parameters correlated with a given materials property via analytical expressions, out of many offered primary features. Crucially, we train ensembles of SISSO models in order to quantify mean predictions and their uncertainty, enabling the use of SISSO in AL. By combining bootstrap sampling to obtain training datasets with Monte-Carlo feature dropout, the high prediction errors observed by a single SISSO model are improved. Besides, the feature dropout procedure alleviates the overconfidence issues observed in the widely used bagging approach. We demonstrate the SISSO-guided AL workflow by identifying acid-stable oxides for water splitting using high-quality DFT-HSE06 calculations. From a pool of 1470 materials, 12 acid-stable materials are identified in only 30 AL iterations. The materials property maps provided by SISSO along with the uncertainty estimates reduce the risk of missing promising portions of the materials space that were overlooked in the initial, possibly biased dataset.

cond-mat.mtrl-sci

Coherent Collections of Rules Describing Exceptional Materials Identified with a Multi-Objective Optimization of Subgroups

Useful materials are often statistically exceptional and they might be overlooked by AI models that attempt to describe all materials simultaneously. These global models perform well for the majority of (useless) materials, but they do not necessarily capture the useful ones. Subgroup discovery (SGD) identifies rules describing subsets of materials (SGs) associated to exceptional values, e.g., high values, of a materials property of interest. Thus, SGD can better capture exceptional materials compared to most widely used AI techniques. Previous works focused on the SG that maximizes an objective function that establishes one tradeoff between the size of the SG and the exceptionality of the distribution of property values in the SG. However, this optimization does not give a unique solution, but many SGs typically have similar objective-function values. Here, we identify a Pareto region of SGs presenting a multitude of size-exceptionality tradeoffs. The approach is demonstrated by the learning of rules describing perovskites with high bulk modulus. These rules are used to screen a large space of perovskites and to efficiently identify materials with bulk modulus up to 13 % higher than the highest value of the training set.

cond-mat.mtrl-sci

Roadmap on Data-Centric Materials Science

Science is and always has been based on data, but the terms "data-centric" and the "4th paradigm of" materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of Artificial Intelligence (AI) and its subset Machine Learning (ML), has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

cond-mat.mtrl-sci

From Prediction to Action: Critical Role of Performance Estimation for Machine-Learning-Driven Materials Discovery

Materials discovery driven by statistical property models is an iterative decision process, during which an initial data collection is extended with new data proposed by a model-informed acquisition function--with the goal to maximize a certain "reward" over time, such as the maximum property value discovered so far. While the materials science community achieved much progress in developing property models that predict well on average with respect to the training distribution, this form of in-distribution performance measurement is not directly coupled with the discovery reward. This is because an iterative discovery process has a shifting reward distribution that is over-proportionally determined by the model performance for exceptional materials. We demonstrate this problem using the example of bulk modulus maximization among double perovskite oxides. We find that the in-distribution predictive performance suggests random forests as superior to Gaussian process regression, while the results are inverse in terms of the discovery rewards. We argue that the lack of proper performance estimation methods from pre-computed data collections is a fundamental problem for improving data-driven materials discovery, and we propose a novel such estimator that, in contrast to na\"ive reward estimation, successfully predicts Gaussian processes with the "expected improvement" acquisition function as the best out of four options in our demonstrational study for double perovskites. Importantly, it does so without requiring the over thousand ab initio computations that were needed to confirm this prediction.

cond-mat.mtrl-sci

Towards a Multi-Objective Optimization of Subgroups for the Discovery of Materials with Exceptional Performance

Artificial intelligence (AI) can accelerate the design of materials by identifying correlations and complex patterns in data. However, AI methods commonly attempt to describe the entire, immense materials space with a single model, while it is typical that different mechanisms govern the materials behaviors across the materials space. The subgroup-discovery (SGD) approach identifies local rules describing exceptional subsets of data with respect to a given target. Thus, SGD can focus on mechanisms leading to exceptional performance. However, the identification of appropriate SG rules requires a careful consideration of the generality-exceptionality tradeoff. Here, we discuss challenges to advance the SGD approach in materials science and analyse the tradeoff between exceptionality and generality based on a Pareto front of SGD solutions.

cond-mat.mtrl-sci

Hierarchical symbolic regression for identifying key physical parameters correlated with bulk properties of perovskites

Symbolic regression identifies key physical parameters describing materials properties by uncovering correlations as nonlinear analytical expressions. However, the pool of expressions grows rapidly with complexity, compromising its efficiency. We tackle this challenge by a hierarchical approach: identified expressions are used as input parameters for obtaining more complex expressions. Crucially, this framework can transfer knowledge among properties, highlighting physical relationships. We demonstrate this strategy by using the Sure-Independence-Screening-and-Sparsifying-Operator (SISSO) approach to identify expressions correlated with the lattice constant and cohesive energy, which are then used to model the bulk modulus of ABO3 perovskites.

cond-mat.mtrl-sci

Identifying outstanding transition-metal-alloy heterogeneous catalysts for the oxygen reduction and evolution reactions via subgroup discovery

In order to estimate the reactivity of a large number of potentially complex heterogeneous catalysts while searching for novel and more efficient materials, physical as well as data-centric models have been developed for a faster evaluation of adsorption energies compared to first-principles calculations. However, global models designed to describe as many materials as possible might overlook the very few compounds that have the appropriate adsorption properties to be suitable for a given catalytic process. Here, the subgroup-discovery (SGD) local artificial-intelligence approach is used to identify the key descriptive parameters and constrains on their values, the so-called SG rules, which particularly describe transition-metal surfaces with outstanding adsorption properties for the oxygen reduction and evolution reactions. We start from a data set of 95 oxygen adsorption energy values evaluated by density-functional-theory calculations for several monometallic surfaces along with 16 atomic, bulk and surface properties as candidate descriptive parameters. From this data set, SGD identifies constraints on the most relevant parameters describing materials and adsorption sites that (i) result in O adsorption energies within the Sabatier-optimal range required for the oxygen reduction reaction and (ii) present the largest deviations from the linear scaling relations between O and OH adsorption energies, which limit the performance in the oxygen evolution reaction. The SG rules not only reflect the local underlying physicochemical phenomena that result in the desired adsorption properties but also guide the challenging design of alloy catalysts.

cond-mat.mtrl-sci

Materials genes of heterogeneous catalysis from clean experiments and artificial intelligence

Heterogeneous catalysis is an example of a complex materials function, governed by an intricate interplay of several processes, e.g., the different surface chemical reactions, and the dynamic re-structuring of the catalyst material at reaction conditions. Modelling the full catalytic progression via first-principles statistical mechanics is impractical, if not impossible. Instead, we show here how a tailored artificial-intelligence approach can be applied, even to a small number of materials, to model catalysis and determine the key descriptive parameters ("materials genes") reflecting the processes that trigger, facilitate, or hinder catalyst performance. We start from a consistent experimental set of "clean data", containing nine vanadium-based oxidation catalysts. These materials were synthesized, fully characterized, and tested according to standardized protocols. By applying the symbolic-regression SISSO approach, we identify correlations between the few most relevant materials properties and their reactivity. This approach highlights the underlying physicochemical processes, and accelerates catalyst design.

cond-mat.mtrl-sci