Searcharxiv⌕ Search

arXiv subjects

Matthias Scheffler

Publications and source records attributed to Matthias Scheffler.

At least 37 records · Page 2Linked to original sources

Self-interaction corrected SCAN functional for molecules and solids in the numeric atom-center orbital framework

Semilocal density-functional approximations (DFAs), including the state-of-the-art SCAN functional, are plagued by the self-interaction error (SIE). While this error is explicitly defined only for one-electron systems, it has inspired the self-interaction correction method proposed by Perdew and Zunger (PZ-SIC), which has shown promise in mitigating the many-electron SIE. However, the PZ-SIC method is known for its significant numerical instability. In this study, we introduce a novel constraint that facilitates self-consistent localization of the SIC orbitals in the spirit of Edmiston-Ruedenberg orbitals [Rev. Mod. Phys. 35, 457 (1963)]. Our practical implementation within the all-electron numeric atom-centered orbitals code FHI-aims guarantees efficient and stable convergence of the self-consistent PZ-SIC equations for both molecules and solids. We further demonstrate that our PZ-SIC approach effectively mitigates the SIE in the meta-GGA SCAN functional, significantly improving the accuracy for ionization potentials, charge-transfer energies, and band gaps for a diverse selection of molecules and solids. However, our PZ-SIC method does have its limitations. It can not improve the already accurate SCAN results for properties such as cohesive energies, lattice constants, and bulk modulus in our test sets. This highlights the need for new-generation DFAs with more comprehensive applicability.

cond-mat.mtrl-sci↗

From Prediction to Action: Critical Role of Performance Estimation for Machine-Learning-Driven Materials Discovery

Materials discovery driven by statistical property models is an iterative decision process, during which an initial data collection is extended with new data proposed by a model-informed acquisition function--with the goal to maximize a certain "reward" over time, such as the maximum property value discovered so far. While the materials science community achieved much progress in developing property models that predict well on average with respect to the training distribution, this form of in-distribution performance measurement is not directly coupled with the discovery reward. This is because an iterative discovery process has a shifting reward distribution that is over-proportionally determined by the model performance for exceptional materials. We demonstrate this problem using the example of bulk modulus maximization among double perovskite oxides. We find that the in-distribution predictive performance suggests random forests as superior to Gaussian process regression, while the results are inverse in terms of the discovery rewards. We argue that the lack of proper performance estimation methods from pre-computed data collections is a fundamental problem for improving data-driven materials discovery, and we propose a novel such estimator that, in contrast to naïve reward estimation, successfully predicts Gaussian processes with the "expected improvement" acquisition function as the best out of four options in our demonstrational study for double perovskites. Importantly, it does so without requiring the over thousand ab initio computations that were needed to confirm this prediction.

cond-mat.mtrl-sci↗

Towards a Multi-Objective Optimization of Subgroups for the Discovery of Materials with Exceptional Performance

Artificial intelligence (AI) can accelerate the design of materials by identifying correlations and complex patterns in data. However, AI methods commonly attempt to describe the entire, immense materials space with a single model, while it is typical that different mechanisms govern the materials behaviors across the materials space. The subgroup-discovery (SGD) approach identifies local rules describing exceptional subsets of data with respect to a given target. Thus, SGD can focus on mechanisms leading to exceptional performance. However, the identification of appropriate SG rules requires a careful consideration of the generality-exceptionality tradeoff. Here, we discuss challenges to advance the SGD approach in materials science and analyse the tradeoff between exceptionality and generality based on a Pareto front of SGD solutions.

cond-mat.mtrl-sci↗

On the Uncertainty Estimates of Equivariant-Neural-Network-Ensembles Interatomic Potentials

Machine-learning (ML) interatomic potentials (IPs) trained on first-principles datasets are becoming increasingly popular since they promise to treat larger system sizes and longer time scales, compared to the {\em ab initio} techniques producing the training data. Estimating the accuracy of MLIPs and reliably detecting when predictions become inaccurate is key for enabling their unfailing usage. In this paper, we explore this aspect for a specific class of MLIPs, the equivariant-neural-network (ENN) IPs using the ensemble technique for quantifying their prediction uncertainties. We critically examine the robustness of uncertainties when the ENN ensemble IP (ENNE-IP) is applied to the realistic and physically relevant scenario of predicting local-minima structures in the configurational space. The ENNE-IP is trained on data for liquid silicon, created by density-functional theory (DFT) with the generalized gradient approximation (GGA) for the exchange-correlation functional. Then, the ensemble-derived uncertainties are compared with the actual errors (comparing the results of the ENNE-IP with those of the underlying DFT-GGA theory) for various test sets, including liquid silicon at different temperatures and out-of-training-domain data such as solid phases with and without point defects as well as surfaces. Our study reveals that the predicted uncertainties are generally overconfident and hold little quantitative predictive power for the actual errors.

cond-mat.mtrl-sci↗

Shared Metadata for Data-Centric Materials Science

The expansive production of data in materials science, their widespread sharing and repurposing requires educated support and stewardship. In order to ensure that this need helps rather than hinders scientific work, the implementation of the FAIR-data principles (Findable, Accessible, Interoperable, and Reusable) must not be too narrow. Besides, the wider materials-science community ought to agree on the strategies to tackle the challenges that are specific to its data, both from computations and experiments. In this paper, we present the result of the discussions held at the workshop on "Shared Metadata and Data Formats for Big-Data Driven Materials Science". We start from an operative definition of metadata, and what features a FAIR-compliant metadata schema should have. We will mainly focus on computational materials-science data and propose a constructive approach for the FAIRification of the (meta)data related to ground-state and excited-states calculations, potential-energy sampling, and generalized workflows. Finally, challenges with the FAIRification of experimental (meta)data and materials-science ontologies are presented together with an outlook of how to meet them.

cond-mat.mtrl-sci↗

Accelerating Materials-Space Exploration for Thermal Insulators by Mapping Materials Properties via Artificial Intelligence

Reliable artificial-intelligence models have the potential to accelerate the discovery of materials with optimal properties for various applications, including superconductivity, catalysis, and thermoelectricity. Advancements in this field are often hindered by the scarcity and quality of available data and the significant effort required to acquire new data. For such applications, reliable surrogate models that help guide materials space exploration using easily accessible materials properties are urgently needed. Here, we present a general, data-driven framework that provides quantitative predictions as well as qualitative rules for steering data creation for all datasets via a combination of symbolic regression and sensitivity analysis. We demonstrate the power of the framework by generating an accurate analytic model for the lattice thermal conductivity using only 75 experimentally measured values. By extracting the most influential material properties from this model, we are then able to hierarchically screen 732 materials and find 80 ultra-insulating materials.

cond-mat.mtrl-sci↗

Extrapolation to complete basis-set limit in density-functional theory by quantile random-forest models

The numerical precision of density-functional-theory (DFT) calculations depends on a variety of computational parameters, one of the most critical being the basis-set size. The ultimate precision is reached with an infinitely large basis set, i.e., in the limit of a complete basis set (CBS). Our aim in this work is to find a machine-learning model that extrapolates finite basis-size calculations to the CBS limit. We start with a data set of 63 binary solids investigated with two all-electron DFT codes, exciting and FHI-aims, which employ very different types of basis sets. A quantile-random-forest model is used to estimate the total-energy correction with respect to a fully converged calculation as a function of the basis-set size. The random-forest model achieves a symmetric mean absolute percentage error of lower than 25% for both codes and outperforms previous approaches in the literature. Our approach also provides prediction intervals, which quantify the uncertainty of the models' predictions.

physics.comp-ph↗

Recent advances in the SISSO method and their implementation in the SISSO++ code

Accurate and explainable artificial-intelligence (AI) models are promising tools for the acceleration of the discovery of new materials, ore new applications for existing materials. Recently, symbolic regression has become an increasingly popular tool for explainable AI because it yields models that are relatively simple analytical descriptions of target properties. Due to its deterministic nature, the sure-independence screening and sparsifying operator (SISSO) method is a particularly promising approach for this application. Here we describe the new advancements of the SISSO algorithm, as implemented into SISSO++, a C++ code with Python bindings. We introduce a new representation of the mathematical expressions found by SISSO. This is a first step towards introducing ``grammar'' rules into the feature creation step. Importantly, by introducing a controlled non-linear optimization to the feature creation step we expand the range of possible descriptors found by the methodology. Finally, we introduce refinements to the solver algorithms for both regression and classification, that drastically increase the reliability and efficiency of SISSO. For all of these improvements to the basic SISSO algorithm, we not only illustrate their potential impact, but also fully detail how they operate both mathematically and computationally.

physics.data-an↗

Ab initio Green-Kubo simulations of heat transport in solids: Method and implementation

Ab initio Green-Kubo (aiGK) simulations of heat transport in solids allow for assessing lattice thermal conductivity in anharmonic or complex materials from first principles. In this work, we present a detailed account of their practical application and evaluation with an emphasis on noise reduction and finite-size corrections in semiconductors and insulators. To account for such corrections, we propose strategies in which all necessary numerical parameters are chosen based on the dynamical properties displayed during molecular dynamics simulations in order to minimize manual intervention. This paves the way for applying the aiGK method in semi-automated and high-throughput frameworks. The proposed strategies are presented and demonstrated for computing the lattice thermal conductivity at room temperature in the mildly anharmonic periclase MgO, and for the strongly anharmonic marshite CuI.

cond-mat.mtrl-sci↗

Anharmonicity in Thermal Insulators: An Analysis from First Principles

The anharmonicity of atomic motion limits the thermal conductivity in crystalline solids. However, a microscopic understanding of the mechanisms active in strong thermal insulators is lacking. In this letter, we classify 465 experimentally known materials with respect to their anharmonicity and perform fully anharmonic ab initio Green-Kubo calculations for 58 of them, finding 28 thermal insulators with $κ< 10$ W/mK including 6 with ultralow $κ\lesssim 1$ W/mK. Our analysis reveals that the underlying strong anharmonic dynamics is driven by the exploration of meta-stable intrinsic defect geometries. This is at variance with the frequently applied perturbative approach, in which the dynamics is assumed to evolve around a single stable geometry.

cond-mat.mtrl-sci↗

Heat flux for semi-local machine-learning potentials

The Green-Kubo (GK) method is a rigorous framework for heat transport simulations in materials. However, it requires an accurate description of the potential-energy surface and carefully converged statistics. Machine-learning potentials can achieve the accuracy of first-principles simulations while allowing to reach well beyond their simulation time and length scales at a fraction of the cost. In this paper, we explain how to apply the GK approach to the recent class of message-passing machine-learning potentials, which iteratively consider semi-local interactions beyond the initial interaction cutoff. We derive an adapted heat flux formulation that can be implemented using automatic differentiation without compromising computational efficiency. The approach is demonstrated and validated by calculating the thermal conductivity of zirconium dioxide across temperatures.

cond-mat.mtrl-sci↗

The NOMAD Artificial-Intelligence Toolkit: Turning materials-science data into knowledge and understanding

We present the Novel-Materials-Discovery (NOMAD) Artificial-Intelligence (AI) Toolkit, a web-browser-based infrastructure for the interactive AI-based analysis of materials-science findable, accessible, interoperable, and reusable (FAIR) data. The AI Toolkit readily operates on the FAIR data stored in the central server of the NOMAD Archive, the largest database of materials-science data worldwide, as well as locally stored, users' owned data. The NOMAD Oasis, a local, stand alone server can be also used to run the AI Toolkit. By using Jupyter notebooks that run in a web-browser, the NOMAD data can be queried and accessed; data mining, machine learning, and other AI techniques can be then applied to analyse them. This infrastructure brings the concept of reproducibility in materials science to the next level, by allowing researchers to share not only the data contributing to their scientific publications, but also all the developed methods and analytics tools. Besides reproducing published results, users of the NOMAD AI toolkit can modify the Jupyter notebooks towards their own research work.

cond-mat.mtrl-sci↗

Roadmap on Electronic Structure Codes in the Exascale Era

Electronic structure calculations have been instrumental in providing many important insights into a range of physical and chemical properties of various molecular and solid-state systems. Their importance to various fields, including materials science, chemical sciences, computational chemistry and device physics, is underscored by the large fraction of available public supercomputing resources devoted to these calculations. As we enter the exascale era, exciting new opportunities to increase simulation numbers, sizes, and accuracies present themselves. In order to realize these promises, the community of electronic structure software developers will however first have to tackle a number of challenges pertaining to the efficient use of new architectures that will rely heavily on massive parallelism and hardware accelerators. This roadmap provides a broad overview of the state-of-the-art in electronic structure calculations and of the various new directions being pursued by the community. It covers 14 electronic structure codes, presenting their current status, their development priorities over the next five years, and their plans towards tackling the challenges and leveraging the opportunities presented by the advent of exascale computing.

cond-mat.mtrl-sci↗

TCMI: a non-parametric mutual-dependence estimator for multivariate continuous distributions

The identification of relevant features, i.e., the driving variables that determine a process or the properties of a system, is an essential part of the analysis of data sets with a large number of variables. A mathematical rigorous approach to quantifying the relevance of these features is mutual information. Mutual information determines the relevance of features in terms of their joint mutual dependence to the property of interest. However, mutual information requires as input probability distributions, which cannot be reliably estimated from continuous distributions such as physical quantities like lengths or energies. Here, we introduce total cumulative mutual information (TCMI), a measure of the relevance of mutual dependences that extends mutual information to random variables of continuous distribution based on cumulative probability distributions. TCMI is a non-parametric, robust, and deterministic measure that facilitates comparisons and rankings between feature sets with different cardinality. The ranking induced by TCMI allows for feature selection, i.e., the identification of variable sets that are nonlinear statistically related to a property of interest, taking into account the number of data samples as well as the cardinality of the set of variables. We evaluate the performance of our measure with simulated data, compare its performance with similar multivariate-dependence measures, and demonstrate the effectiveness of our feature-selection method on a set of standard data sets and a typical scenario in materials science.

stat.ML↗

Artifcial-intelligence-driven discovery of catalyst \textit{genes} with application to CO2 activation on semiconductor oxides

Catalytic-materials design requires predictive modeling of the interaction between catalyst and reactants. This is challenging due to the complexity and diversity of structure-property relationships across the chemical space. Here, we report a strategy for a rational design of catalytic materials using the artifcial intelligence approach (AI) subgroup discovery. We identify catalyst \textit{genes} (features) that correlate with mechanisms that trigger, facilitate, or hinder the activation of carbon dioxide (CO$_2$) towards a chemical conversion. The AI model is trained on frst-principles data for a broad family of oxides. We demonstrate that surfaces of experimentally identifed good catalysts consistently exhibit combinations of \textit{genes} resulting in a strong elongation of a C-O bond. The same combinations of \textit{genes} also minimize the OCO-angle, the previously proposed indicator of activation, albeit under the constraint that the Sabatier principle is satisfed. Based on these fndings, we propose a set of new promising catalyst materials for CO$_2$ conversion.

cond-mat.mtrl-sci↗

FAIR data enabling new horizons for materials research

The prosperity and lifestyle of our society are very much governed by achievements in condensed matter physics, chemistry and materials science, because new products for sectors such as energy, the environment, health, mobility and information technology (IT) rely largely on improved or even new materials. Examples include solid-state lighting, touchscreens, batteries, implants, drug delivery and many more. The enormous amount of research data produced every day in these fields represents a gold mine of the twenty-first century. This gold mine is, however, of little value if these data are not comprehensively characterized and made available. How can we refine this feedstock; that is, turn data into knowledge and value? For this, a FAIR (findable, accessible, interoperable and reusable) data infrastructure is a must. Only then can data be readily shared and explored using data analytics and artificial intelligence (AI) methods. Making data 'findable and AI ready' (a forward-looking interpretation of the acronym) will change the way in which science is carried out today. In this Perspective, we discuss how we can prepare to make this happen for the field of materials science.

cond-mat.mtrl-sci↗

Interface to high-performance periodic coupled-cluster theory calculations with atom-centered, localized basis functions

Coupled cluster (CC) theory is often considered the gold standard of quantum-chemistry. For solids, however, the available software is scarce. We present CC-aims, which can interface ab initio codes with localized atomic orbitals and the CC for solids (CC4S) code by the group of A. Grüneis. CC4S features a continuously growing selection of wave function-based methods including perturbation and CC theory. The CC-aims interface was developed for the FHI-aims code (https://fhi-aims.org) but is implemented such that other codes may use it as a starting point for corresponding interfaces. As CC4S offers treatment of both molecular and periodic systems, the CC-aims interface is a valuable tool, where DFT is either too inaccurate or too unreliable, in theoretical chemistry and materials science alike.

cond-mat.mtrl-sci↗