SearcharxivSearch

arXiv subjects

Luca M. Ghiringhelli

Publications and source records attributed to Luca M. Ghiringhelli.

At least 19 recordsLinked to original sources

Machine-learning octet $AB$-type binary compounds across chemical space with domain knowledge of the interatomic bond

The prediction of the structural stability of octet $AB$-type binary compounds is a classical materials informatics problem. The challenge is to capture the relative stability of 4-fold coordinated atoms in zincblende ($\beta$-ZnS) structure and 6-fold coordinated atoms in rocksalt (NaCl) structure, modulated by charge transfer and atomic-size differences. Previous structure maps and machine-learning approaches used atomic features such as valence-electron count, ionization potential and atomic radii, using either physical intuition or symbolic regression. Here, we demonstrate that explicitly incorporating the domain knowledge of the interatomic bonds can significantly and systematically improve the prediction of $\beta$-ZnS/NaCl stability. We encode this bonding information through a coarse-grained representation of the local electronic structure obtained by a recursive solution of a tight-binding bond model. The underlying pairwise Hamiltonians are taken from downfolded eigenstates of density-functional theory calculations for diatomic molecules and thereby include domain knowledge of the bond between specific $A-B$ pairs. The benefit of this description is demonstrated with an ensemble of independently trained Kernel Ridge or symbolic regression models combined with sequential feature selection. The obtained models are compared to a previous symbolic-regression model using the same set of \emph{ab initio} calculations for octet binaries as training data. We find a significant improvement in the prediction of the formation energy difference of $AB$ compounds as compared to previous works and demonstrate that an increasing amount of bond-informed recursion features improves the predictive accuracy.

cond-mat.mtrl-sci

Roadmap on Advancements of the FHI-aims Software Package

Electronic-structure theory is the foundation of the description of materials including multiscale modeling of their properties and functions. Obviously, without sufficient accuracy at the base, reliable predictions are unlikely at any level that follows. The software package FHI-aims has proven to be a game changer for accurate free-energy calculations because of its scalability, numerical precision, and its efficient handling of density functional theory (DFT) with hybrid functionals and van der Waals interactions. It treats molecules, clusters, and extended systems (solids and liquids) on an equal footing. Besides DFT, FHI-aims also includes quantum-chemistry methods, descriptions for excited states and vibrations, and calculations of various types of transport. Recent advancements address the integration of FHI-aims into an increasing number of workflows and various artificial intelligence (AI) methods. This Roadmap describes the state-of-the-art of FHI-aims and advancements that are currently ongoing or planned.

cond-mat.mtrl-sci

A high-performance and portable implementation of the SISSO method for CPUs and GPUs

SISSO (sure-independence screening and sparsifying operator) is an artificial intelligence (AI) method based on symbolic regression and compressed sensing widely used in materials science research. SISSO++ is its C++ implementation that employs MPI and OpenMP for parallelization, rendering it well-suited for high-performance computing (HPC) environments. As heterogeneous hardware becomes mainstream in the HPC and AI fields, we chose to port the SISSO++ code to GPUs using the Kokkos performance-portable library. Kokkos allows us to maintain a single codebase for both Nvidia and AMD GPUs, significantly reducing the maintenance effort. In this work, we summarize the necessary code changes we did to achieve hardware and performance portability. This is accompanied by performance benchmarks on Nvidia and AMD GPUs. We demonstrate the speedups obtained from using GPUs across the three most time-consuming parts of our code.

cs.PF

How big is Big Data?

Big data has ushered in a new wave of predictive power using machine learning models. In this work, we assess what {\it big} means in the context of typical materials-science machine-learning problems. This concerns not only data volume, but also data quality and veracity as much as infrastructure issues. With selected examples, we ask (i) how models generalize to similar datasets, (ii) how high-quality datasets can be gathered from heterogenous sources, (iii) how the feature set and complexity of a model can affect expressivity, and (iv) what infrastructure requirements are needed to create larger datasets and train models on them. In sum, we find that big data present unique challenges along very different aspects that should serve to motivate further work.

stat.ML

Roadmap on Data-Centric Materials Science

Science is and always has been based on data, but the terms "data-centric" and the "4th paradigm of" materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of Artificial Intelligence (AI) and its subset Machine Learning (ML), has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

cond-mat.mtrl-sci

Combining genetic algorithm and compressed sensing for features and operators selection in symbolic regression

Symbolic-inference methods have recently found a broad application in materials science. In particular, the Sure-Independence Screening and Sparsifying Operator (SISSO) performs symbolic regression and classification by adopting compressed sensing for the selection of an optimized subset of features and mathematical operators out of a given set of candidates. However, SISSO becomes computationally unpractical when the set of candidate features and operators exceeds the size of few tens. In the present work, we combine SISSO with a genetic algorithm (GA) for the global search of the optimal subset of features and operators. We demonstrate that GA-SISSO efficiently finds more accurate predictive models than the original SISSO, due to the possibility to access a larger input feature and operator space. GA-SISSO was applied for the search of the model for the prediction of carbon-dioxide adsorption energies on semiconductor oxides. The obtained with GA-SISSO model has much higher accuracy compared to models previously discussed in the literature (based solely on the O 2p-band center). The analysis of features importance shows that, besides the O 2p-band center, the contribution of the electrostatic potential above adsorption sites and the surface formation energies are also important.

cond-mat.mtrl-sci

Automatic Identification of Crystal Structures and Interfaces via Artificial-Intelligence-based Electron Microscopy

Characterizing crystal structures and interfaces down to the atomic level is an important step for designing advanced materials. Modern electron microscopy routinely achieves atomic resolution and is capable to resolve complex arrangements of atoms with picometer precision. Here, we present AI-STEM, an automatic, artificial-intelligence based method, for accurately identifying key characteristics from atomic-resolution scanning transmission electron microscopy (STEM) images of polycrystalline materials. The method is based on a Bayesian convolutional neural network (BNN) that is trained only on simulated images. AI-STEM automatically and accurately identifies crystal structure, lattice orientation, and location of interface regions in synthetic and experimental images. The model is trained on cubic and hexagonal crystal structures, yielding classifications and uncertainty estimates, while no explicit information on structural patterns at the interfaces is included during training. This work combines principles from probabilistic modeling, deep learning, and information theory, enabling automatic analysis of experimental, atomic-resolution images.

cond-mat.mtrl-sci

On the Uncertainty Estimates of Equivariant-Neural-Network-Ensembles Interatomic Potentials

Machine-learning (ML) interatomic potentials (IPs) trained on first-principles datasets are becoming increasingly popular since they promise to treat larger system sizes and longer time scales, compared to the {\em ab initio} techniques producing the training data. Estimating the accuracy of MLIPs and reliably detecting when predictions become inaccurate is key for enabling their unfailing usage. In this paper, we explore this aspect for a specific class of MLIPs, the equivariant-neural-network (ENN) IPs using the ensemble technique for quantifying their prediction uncertainties. We critically examine the robustness of uncertainties when the ENN ensemble IP (ENNE-IP) is applied to the realistic and physically relevant scenario of predicting local-minima structures in the configurational space. The ENNE-IP is trained on data for liquid silicon, created by density-functional theory (DFT) with the generalized gradient approximation (GGA) for the exchange-correlation functional. Then, the ensemble-derived uncertainties are compared with the actual errors (comparing the results of the ENNE-IP with those of the underlying DFT-GGA theory) for various test sets, including liquid silicon at different temperatures and out-of-training-domain data such as solid phases with and without point defects as well as surfaces. Our study reveals that the predicted uncertainties are generally overconfident and hold little quantitative predictive power for the actual errors.

cond-mat.mtrl-sci

Shared Metadata for Data-Centric Materials Science

The expansive production of data in materials science, their widespread sharing and repurposing requires educated support and stewardship. In order to ensure that this need helps rather than hinders scientific work, the implementation of the FAIR-data principles (Findable, Accessible, Interoperable, and Reusable) must not be too narrow. Besides, the wider materials-science community ought to agree on the strategies to tackle the challenges that are specific to its data, both from computations and experiments. In this paper, we present the result of the discussions held at the workshop on "Shared Metadata and Data Formats for Big-Data Driven Materials Science". We start from an operative definition of metadata, and what features a FAIR-compliant metadata schema should have. We will mainly focus on computational materials-science data and propose a constructive approach for the FAIRification of the (meta)data related to ground-state and excited-states calculations, potential-energy sampling, and generalized workflows. Finally, challenges with the FAIRification of experimental (meta)data and materials-science ontologies are presented together with an outlook of how to meet them.

cond-mat.mtrl-sci

Accelerating Materials-Space Exploration for Thermal Insulators by Mapping Materials Properties via Artificial Intelligence

Reliable artificial-intelligence models have the potential to accelerate the discovery of materials with optimal properties for various applications, including superconductivity, catalysis, and thermoelectricity. Advancements in this field are often hindered by the scarcity and quality of available data and the significant effort required to acquire new data. For such applications, reliable surrogate models that help guide materials space exploration using easily accessible materials properties are urgently needed. Here, we present a general, data-driven framework that provides quantitative predictions as well as qualitative rules for steering data creation for all datasets via a combination of symbolic regression and sensitivity analysis. We demonstrate the power of the framework by generating an accurate analytic model for the lattice thermal conductivity using only 75 experimentally measured values. By extracting the most influential material properties from this model, we are then able to hierarchically screen 732 materials and find 80 ultra-insulating materials.

cond-mat.mtrl-sci

Uncertainty Quantification in Deep Neural Networks through Statistical Inference on Latent Space

Uncertainty-quantification methods are applied to estimate the confidence of deep-neural-networks classifiers over their predictions. However, most widely used methods are known to be overconfident. We address this problem by developing an algorithm that exploits the latent-space representation of data points fed into the network, to assess the accuracy of their prediction. Using the latent-space representation generated by the fraction of training set that the network classifies correctly, we build a statistical model that is able to capture the likelihood of a given prediction. We show on a synthetic dataset that commonly used methods are mostly overconfident. Overconfidence occurs also for predictions made on data points that are outside the distribution that generated the training data. In contrast, our method can detect such out-of-distribution data points as inaccurately predicted, thus aiding in the automatic detection of outliers.

cs.LG

Recent advances in the SISSO method and their implementation in the SISSO++ code

Accurate and explainable artificial-intelligence (AI) models are promising tools for the acceleration of the discovery of new materials, ore new applications for existing materials. Recently, symbolic regression has become an increasingly popular tool for explainable AI because it yields models that are relatively simple analytical descriptions of target properties. Due to its deterministic nature, the sure-independence screening and sparsifying operator (SISSO) method is a particularly promising approach for this application. Here we describe the new advancements of the SISSO algorithm, as implemented into SISSO++, a C++ code with Python bindings. We introduce a new representation of the mathematical expressions found by SISSO. This is a first step towards introducing ``grammar'' rules into the feature creation step. Importantly, by introducing a controlled non-linear optimization to the feature creation step we expand the range of possible descriptors found by the methodology. Finally, we introduce refinements to the solver algorithms for both regression and classification, that drastically increase the reliability and efficiency of SISSO. For all of these improvements to the basic SISSO algorithm, we not only illustrate their potential impact, but also fully detail how they operate both mathematically and computationally.

physics.data-an

The NOMAD Artificial-Intelligence Toolkit: Turning materials-science data into knowledge and understanding

We present the Novel-Materials-Discovery (NOMAD) Artificial-Intelligence (AI) Toolkit, a web-browser-based infrastructure for the interactive AI-based analysis of materials-science findable, accessible, interoperable, and reusable (FAIR) data. The AI Toolkit readily operates on the FAIR data stored in the central server of the NOMAD Archive, the largest database of materials-science data worldwide, as well as locally stored, users' owned data. The NOMAD Oasis, a local, stand alone server can be also used to run the AI Toolkit. By using Jupyter notebooks that run in a web-browser, the NOMAD data can be queried and accessed; data mining, machine learning, and other AI techniques can be then applied to analyse them. This infrastructure brings the concept of reproducibility in materials science to the next level, by allowing researchers to share not only the data contributing to their scientific publications, but also all the developed methods and analytics tools. Besides reproducing published results, users of the NOMAD AI toolkit can modify the Jupyter notebooks towards their own research work.

cond-mat.mtrl-sci

TCMI: a non-parametric mutual-dependence estimator for multivariate continuous distributions

The identification of relevant features, i.e., the driving variables that determine a process or the properties of a system, is an essential part of the analysis of data sets with a large number of variables. A mathematical rigorous approach to quantifying the relevance of these features is mutual information. Mutual information determines the relevance of features in terms of their joint mutual dependence to the property of interest. However, mutual information requires as input probability distributions, which cannot be reliably estimated from continuous distributions such as physical quantities like lengths or energies. Here, we introduce total cumulative mutual information (TCMI), a measure of the relevance of mutual dependences that extends mutual information to random variables of continuous distribution based on cumulative probability distributions. TCMI is a non-parametric, robust, and deterministic measure that facilitates comparisons and rankings between feature sets with different cardinality. The ranking induced by TCMI allows for feature selection, i.e., the identification of variable sets that are nonlinear statistically related to a property of interest, taking into account the number of data samples as well as the cardinality of the set of variables. We evaluate the performance of our measure with simulated data, compare its performance with similar multivariate-dependence measures, and demonstrate the effectiveness of our feature-selection method on a set of standard data sets and a typical scenario in materials science.

stat.ML

Artifcial-intelligence-driven discovery of catalyst \textit{genes} with application to CO2 activation on semiconductor oxides

Catalytic-materials design requires predictive modeling of the interaction between catalyst and reactants. This is challenging due to the complexity and diversity of structure-property relationships across the chemical space. Here, we report a strategy for a rational design of catalytic materials using the artifcial intelligence approach (AI) subgroup discovery. We identify catalyst \textit{genes} (features) that correlate with mechanisms that trigger, facilitate, or hinder the activation of carbon dioxide (CO$_2$) towards a chemical conversion. The AI model is trained on frst-principles data for a broad family of oxides. We demonstrate that surfaces of experimentally identifed good catalysts consistently exhibit combinations of \textit{genes} resulting in a strong elongation of a C-O bond. The same combinations of \textit{genes} also minimize the OCO-angle, the previously proposed indicator of activation, albeit under the constraint that the Sabatier principle is satisfed. Based on these fndings, we propose a set of new promising catalyst materials for CO$_2$ conversion.

cond-mat.mtrl-sci

Hierarchical symbolic regression for identifying key physical parameters correlated with bulk properties of perovskites

Symbolic regression identifies key physical parameters describing materials properties by uncovering correlations as nonlinear analytical expressions. However, the pool of expressions grows rapidly with complexity, compromising its efficiency. We tackle this challenge by a hierarchical approach: identified expressions are used as input parameters for obtaining more complex expressions. Crucially, this framework can transfer knowledge among properties, highlighting physical relationships. We demonstrate this strategy by using the Sure-Independence-Screening-and-Sparsifying-Operator (SISSO) approach to identify expressions correlated with the lattice constant and cohesive energy, which are then used to model the bulk modulus of ABO3 perovskites.

cond-mat.mtrl-sci

Ab initio approach for thermodynamic surface phases with full consideration of anharmonic effects -- the example of hydrogen at Si(100)

A reliable description of surfaces structures in a reactive environment is crucial to understand materials functions. We present a first-principles theory of replica-exchange grand-canonical-ensemble molecular dynamics (REGC-MD) and apply it to evaluate phase equilibria of surfaces in reactive gas-phase environment. We identify the different surface phases and locate phase boundaries including triple and critical points. The approach is demonstrated by addressing open questions for the Si(100) surface in contact with a hydrogen atmosphere. In the range from 300 to 1 000 K, we find 25 distinct thermodynamically stable surface phases, for which we also provide microscopic descriptions. Most of the identified phases, including few order-disorder phase transitions, have not yet been observed experimentally. The REGC-MD-derived phase diagram shows significant, qualitative differences to the description by the state-of-the-art "ab initio atomistic thermodynamics" approach.

cond-mat.mtrl-sci

Numerical Quality Control for DFT-based Materials Databases

Electronic-structure theory is a strong pillar of materials science. Many different computer codes that employ different approaches are used by the community to solve various scientific problems. Still, the precision of different packages has only recently been scrutinized thoroughly, focusing on a specific task, namely selecting a popular density functional, and using unusually high, extremely precise numerical settings for investigating 71 monoatomic crystals. Little is known, however, about method- and code-specific uncertainties that arise under numerical settings that are commonly used in practice. We shed light on this issue by investigating the deviations in total and relative energies as a function of computational parameters. Using typical settings for basis sets and k-grids, we compare results for 71 elemental and 63 binary solids obtained by three different electronic-structure codes that employ fundamentally different strategies. On the basis of the observed trends, we propose a simple, analytical model for the estimation of the errors associated with the basis-set incompleteness. We cross-validate this model using ternary systems obtained from the NOMAD Repository and discuss how our approach enables the comparison of the heterogeneous data present in computational materials databases.

physics.comp-ph