SearcharxivSearch

arXiv subjects

Giovanni Pizzi

Publications and source records attributed to Giovanni Pizzi.

At least 19 recordsLinked to original sources

Correcting DFT formation energies towards experimental accuracy using foundational MLIPs and latent-feature delta-learning

Crystal structure databases curated by high-throughput density functional theory calculations typically serve as the starting point for computational materials discovery efforts. Thermodynamic stability data, such as formation energies and the energy above the convex hull, are important quantities to guide the search for novel materials, enabling filtering for (meta)stable structures. Here, we present the thermodynamic stability of the fully open-source, reproducible, and experimentally focused Materials Cloud three-dimensional crystals database (MC3D). We compare against two other DFT databases, the Open Quantum Materials Database (OQMD) and the Materials Project (MP), as well as against experimental formation enthalpies. We then demonstrate how recent foundational machine learning interatomic potentials (MLIPs) trained at the r$^2$SCAN level (specifically, we test PET-OMATPES here) can be leveraged to improve the agreement of formation energies with experiment, reducing the mean absolute error by more than 40% relative to GGA without requiring any additional DFT calculation. Our results validate and extend the established practice of combining PBEsol geometries with meta-GGA energies to the era of foundational MLIPs. Finally, we train classical machine learning models to further correct the formation energies in a delta-learning framework, where we use the information-rich latent features of the foundational MLIP. These models further reduce the mean absolute error below 50 meV/atom, bringing it down to values comparable with the experimental uncertainty itself. Notably, compared to purely compositional features, the latent features (combined with carefully tuned regularization) simultaneously reduce the prediction error and limit the impact of the learned corrections on the relative phase stability.

cond-mat.mtrl-sci

optimade-maker: Automated generation of interoperable materials APIs from static datasets

Atomistic structural data are central to materials science, condensed matter physics, and chemistry, and are increasingly digitised across diverse repositories and databases. Interoperable access to these heterogeneous data sources enables reusable clients and tools, and is essential for cross-database analyses and data-driven materials discovery. Toward this aim, the OPTIMADE (Open Databases Integration for Materials Design) specification defines a standard REST API for atomistic structures and related properties. However, deploying and maintaining compliant services remains technically demanding and poses a significant barrier for many data providers. Here, we present optimade-maker, a lightweight toolkit for the automated generation of OPTIMADE-compliant APIs directly from raw atomistic structure and property data. The toolkit supports a wide range of raw datasets, enables conversion to a standardised OPTIMADE data representation, and allows for rapid deployment of APIs in both local and production environments. We further demonstrate it through an automated service on the Materials Cloud Archive, which automatically creates and publishes OPTIMADE APIs for contributed datasets, enabling immediate discoverability and interoperability. In addition, we implement data transformation pipelines for the Cambridge Structural Database (CSD) and the Inorganic Crystal Structure Database (ICSD), enabling unified access to these curated resources through the OPTIMADE framework. By lowering the technical barriers to interoperable data publication, optimade-maker represents an important step toward a scalable, FAIR materials data ecosystem integrating both community-contributed and curated databases.

cs.DB

Score-based diffusion models for accurate crystal-structure inpainting and reconstruction of hydrogen positions

Generative AI models, such as score-based diffusion models, have recently advanced the field of computational materials science by enabling the generation of new materials with desired properties. In addition, these models could also be leveraged to reconstruct crystal structures for which partial information is available. One relevant example is the reliable determination of atomic positions occupied by hydrogen atoms in hydrogen-containing crystalline materials. While crucial to the analysis and prediction of many materials properties, the identification of hydrogen positions can however be difficult and expensive, as it is challenging in X-ray scattering experiments and often requires dedicated neutron scattering measurements. As a consequence, inorganic crystallographic databases frequently report lattice structures where hydrogen atoms have been either omitted or inserted with heuristics or by chemical intuition. Here, we combine diffusion models from the field of materials science with techniques originally developed in computer vision for image inpainting. We present how this knowledge transfer across domains enables a much faster and more accurate completion of host structures, compared to unconditioned diffusion models or previous approaches solely based on DFT. Overall, our approach exceeds a success rate of 97% in terms of finding a structural match or predicting a more stable configuration than the initial reference, when starting both from structures that were already relaxed with DFT, or directly from the experimentally determined host structures.

cond-mat.mtrl-sci

Accelerating discovery across scientific disciplines through reproducible workflows with AiiDAlab

With ever-increasing computational capabilities, robust and automated research workflows have become essential for orchestrating large numbers of interdependent simulations. However, significant technical expertise is still required to configure execution environments, define calculation inputs, interpret outputs, and manage the complexity of parallel code execution on remote machines. To address these challenges, we developed AiiDAlab, a Jupyter-based web platform powered by the AiiDA computational infrastructure that provides a framework for managing and automating computational workflows while ensuring reproducibility through full provenance tracking. Through a collection of open-source user-friendly applications, AiiDAlab enables scientists to set up, execute, and analyze complex computational workflows without interacting directly with the underlying technical details, allowing them to focus on their research questions. In this paper, we discuss how AiiDAlab has matured over the past few years, expanding beyond computational materials science and its AiiDA origins. We present recent developments towards integrating with electronic laboratory notebooks (ELNs) for FAIR-compliant data management, adoption in large-scale facilities for secure access to experimental data and analytical tools, and applications in educational settings. Together with community-driven efforts to simplify onboarding, improve access to computational resources, and support large-scale data workflows, these advancements position AiiDAlab as a powerful platform for accelerating scientific discovery and fostering collaboration across disciplines.

cs.DC

Implementing a Scalable, Redeployable and Multitiered Repository for FAIR and Secure Scientific Data Sharing: The BIG-MAP Archive

Data sharing in large consortia, such as research collaborations or industry partnerships, requires addressing both organizational and technical challenges. A common platform is essential to promote collaboration, facilitate exchange of findings, and ensure secure access to sensitive data. Key technical challenges include creating a scalable architecture, a user-friendly interface, and robust security and access control. The BIG-MAP Archive is a cloud-based, disciplinary, private repository designed to address these challenges. Built on InvenioRDM, it leverages platform functionalities to meet consortium-specific needs, providing a tailored solution compared to general repositories. Access can be restricted to members of specific communities or open to the entire consortium, such as the BATTERY 2030+, a consortium accelerating advanced battery technologies. Uploaded data and metadata are controlled via fine grained permissions, allowing access to individual project members or the full initiative. The formalized upload process ensures data are formatted and ready for publication in open repositories when needed. This paper reviews the repository's key features, showing how the BIG-MAP Archive enables secure, controlled data sharing within large consortia. It ensures data confidentiality while supporting flexible, permissions-based access and can be easily redeployed for other consortia, including MaterialsCommons4.eu and RAISE (Resource for AI Science in Europe).

cs.DB

FirecREST v2: lessons learned from redesigning an API for scalable HPC resource access

Introducing FirecREST v2, the next generation of our open-source RESTful API for programmatic access to HPC resources. FirecREST v2 delivers a 100x performance improvement over its predecessor. This paper explores the lessons learned from redesigning FirecREST from the ground up, with a focus on integrating enhanced security and high throughput as core requirements. We provide a detailed account of our systematic performance testing methodology, highlighting common bottlenecks in proxy-based APIs with intensive I/O operations. Key design and architectural changes that enabled these performance gains are presented. Finally, we demonstrate the impact of these improvements, supported by independent peer validation, and discuss opportunities for further improvements.

cs.DC

MC3D: The Materials Cloud computational database of experimentally known stoichiometric inorganics

DFT is a widely used method to compute properties of materials, which are often collected in databases and serve as valuable starting points for further studies. In this article, we present the Materials Cloud Three-Dimensional Structure Database (MC3D), an online database of computed three-dimensional (3D) inorganic crystal structures. Close to a million experimentally reported structures were imported from the COD, ICSD and MPDS databases; these were parsed and filtered to yield a collection of 72589 unique and stoichiometric structures, of which 95% are, to date, classified as experimentally known. The geometries of structures with up to 64 atoms were then optimized using density-functional theory (DFT) with automated workflows and curated input protocols. The procedure was repeated for different functionals (and computational protocols), with the latest version (MC3D PBEsol-v2) comprising 32013 unique structures. All versions of the MC3D are made available on the Materials Cloud portal, which provides a graphical interface to explore and download the data. The database includes the full provenance graph of all the calculations driven by the automated workflows, thus establishing full reproducibility of the results and more-than-FAIR procedures.

cond-mat.mtrl-sci

Making atomistic materials calculations accessible with the AiiDAlab Quantum ESPRESSO app

Despite the wide availability of density functional theory (DFT) codes, their adoption by the broader materials science community remains limited due to challenges such as software installation, input preparation, high-performance computing setup, and output analysis. To overcome these barriers, we introduce the Quantum ESPRESSO app, an intuitive, web-based platform built on AiiDAlab that integrates user-friendly graphical interfaces with automated DFT workflows. The app employs a modular Input-Process-Output model and a plugin-based architecture, providing predefined computational protocols, automated error handling, and interactive results visualization. We demonstrate the app's capabilities through plugins for electronic band structures, projected density of states, phonon, infrared/Raman, X-ray and muon spectroscopies, Hubbard parameters (DFT+$U$+$V$), Wannier functions, and post-processing tools. By extending the FAIR principles to simulations, workflows, and analyses, the app enhances the accessibility and reproducibility of advanced DFT calculations and provides a general template to interface with other first-principles calculation codes.

cond-mat.mtrl-sci

Robust Wannierization including magnetization and spin-orbit coupling via projectability disentanglement

Maximally-localized Wannier functions (MLWFs) are widely employed as an essential tool for calculating the physical properties of materials due to their localized nature and computational efficiency. Projectability-disentangled Wannier functions (PDWFs) have recently emerged as a reliable and efficient approach for automatically constructing MLWFs that span both occupied and lowest unoccupied bands. Here, we extend the applicability of PDWFs to magnetic systems and/or those including spin-orbit coupling, and implement such extensions in automated workflows. Furthermore, we enhance the robustness and reliability of constructing PDWFs by defining an extended protocol that automatically expands the projectors manifold, when required, by introducing additional appropriate hydrogenic atomic orbitals. We benchmark our extended protocol on a set of 200 chemically diverse materials, as well as on the 40 systems with the largest band distance obtained with the standard PDWF approach, showing that on our test set the present approach delivers a 100% success rate in obtaining accurate Wannier-function interpolations, i.e., an average band distance below 15 meV between the DFT and Wannier-interpolated bands, up to 2 eV above the Fermi level.

cond-mat.mtrl-sci

scicode-widgets: Bringing Computational Experiments to the Classroom with Jupyter Widgets

"Computational experiments" use code and interactive visualizations to convey mathematical and physical concepts in an intuitive way, and are increasingly used to support ex cathedra lecturing in scientific and engineering disciplines. Jupyter notebooks are particularly well-suited to implement them, but involve large amounts of ancillary code to process data and generate illustrations, which can distract students from the core learning outcomes. For a more engaging learning experience that only exposes relevant code to students, allowing them to focus on the interplay between code, theory and physical insights, we developed scicode-widgets (released as scwidgets), a Python package to build Jupyter-based applications. The package facilitates the creation of interactive exercises and demonstrations for students in any discipline in science, technology and engineering. Students are asked to provide pedagogically meaningful contributions in terms of theoretical understanding, coding ability, and analytical skills. The library provides the tools to connect custom pre- and post-processing of students' code, which runs seamlessly "behind the scenes", with the ability to test and verify the solution, as well as to convert it into live interactive visualizations driven by Jupyter widgets.

physics.ed-ph

Massive Atomic Diversity: a compact universal dataset for atomistic machine learning

The development of machine-learning models for atomic-scale simulations has benefited tremendously from the large databases of materials and molecular properties computed in the past two decades using electronic-structure calculations. More recently, these databases have made it possible to train universal models that aim at making accurate predictions for arbitrary atomic geometries and compositions. The construction of many of these databases was however in itself aimed at materials discovery, and therefore targeted primarily to sample stable, or at least plausible, structures and to make the most accurate predictions for each compound - e.g. adjusting the calculation details to the material at hand. Here we introduce a dataset designed specifically to train machine learning models that can provide reasonable predictions for arbitrary structures, and that therefore follows a different philosophy. Starting from relatively small sets of stable structures, the dataset is built to contain massive atomic diversity (MAD) by aggressively distorting these configurations, with near-complete disregard for the stability of the resulting configurations. The electronic structure details, on the other hand, are chosen to maximize consistency rather than to obtain the most accurate prediction for a given structure, or to minimize computational effort. The MAD dataset we present here, despite containing fewer than 100k structures, has already been shown to enable training universal interatomic potentials that are competitive with models trained on traditional datasets with two to three orders of magnitude more structures. We describe in detail the philosophy and details of the construction of the MAD dataset. We also introduce a low-dimensional structural latent space that allows us to compare it with other popular datasets and that can be used as a general-purpose materials cartography tool.

cond-mat.mtrl-sci

A Terminology for Scientific Workflow Systems

The term scientific workflow has evolved over the last two decades to encompass a broad range of compositions of interdependent compute tasks and data movements. It has also become an umbrella term for processing in modern scientific applications. Today, many scientific applications can be considered as workflows made of multiple dependent steps, and hundreds of workflow management systems (WMSs) have been developed to manage and run these workflows. However, no turnkey solution has emerged to address the diversity of scientific processes and the infrastructure on which they are implemented. Instead, new research problems requiring the execution of scientific workflows with some novel feature often lead to the development of an entirely new WMS. A direct consequence is that many existing WMSs share some salient features, offer similar functionalities, and can manage the same categories of workflows but also have some distinct capabilities. This situation makes researchers who develop workflows face the complex question of selecting a WMS. This selection can be driven by technical considerations, to find the system that is the most appropriate for their application and for the resources available to them, or other factors such as reputation, adoption, strong community support, or long-term sustainability. To address this problem, a group of WMS developers and practitioners joined their efforts to produce a community-based terminology of WMSs. This paper summarizes their findings and introduces this new terminology to characterize WMSs. This terminology is composed of fives axes: workflow characteristics, composition, orchestration, data management, and metadata capture. Each axis comprises several concepts that capture the prominent features of WMSs. Based on this terminology, this paper also presents a classification of 23 existing WMSs according to the proposed axes and terms.

cs.DC

A Python workflow definition for computational materials design

Numerous Workflow Management Systems (WfMS) have been developed in the field of computational materials science with different workflow formats, hindering interoperability and reproducibility of workflows in the field. To address this challenge, we introduce here the Python Workflow Definition (PWD) as a workflow exchange format to share workflows between Python-based WfMS, currently AiiDA, jobflow, and pyiron. This development is motivated by the similarity of these three Python-based WfMS, that represent the different workflow steps and data transferred between them as nodes and edges in a graph. With the PWD, we aim at fostering the interoperability and reproducibility between the different WfMS in the context of Findable, Accessible, Interoperable, Reusable (FAIR) workflows. To separate the scientific from the technical complexity, the PWD consists of three components: (1) a conda environment that specifies the software dependencies, (2) a Python module that contains the Python functions represented as nodes in the workflow graph, and (3) a workflow graph stored in the JavaScript Object Notation (JSON). The first version of the PWD supports directed acyclic graph (DAG)-based workflows. Thus, any DAG-based workflow defined in one of the three WfMS can be exported to the PWD and afterwards imported from the PWD to one of the other WfMS. After the import, the input parameters of the workflow can be adjusted and computing resources can be assigned to the workflow, before it is executed with the selected WfMS. This import from and export to the PWD is enabled by the PWD Python library that implements the PWD in AiiDA, jobflow, and pyiron.

cs.SE

Accurate and efficient protocols for high-throughput first-principles materials simulations

Advancements in theoretical and algorithmic approaches, workflow engines, and an ever-increasing computational power have enabled a novel paradigm for materials discovery through first-principles high-throughput simulations. A major challenge in these efforts is to automate the selection of parameters used by simulation codes to deliver numerical precision and computational efficiency. Here, we propose a rigorous methodology to assess the quality of self-consistent DFT calculations with respect to smearing and $k$-point sampling across a wide range of crystalline materials. For this goal, we develop criteria to reliably estimate average errors on total energies, forces, and other properties as a function of the desired computational efficiency, while consistently controlling $k$-point sampling errors. The present results provide automated protocols (named standard solid-state protocols or SSSP) for selecting optimized parameters based on different choices of precision and efficiency tradeoffs. These are available through open-source tools that range from interactive input generators for DFT codes to high-throughput workflows.

cond-mat.mtrl-sci

Charting the landscape of Bardeen-Cooper-Schrieffer superconductors in experimentally known compounds

We perform a high-throughput computational search for novel phonon-mediated superconductors, starting from the Materials Cloud 3-dimensional structure database of experimentally known inorganic stoichiometric compounds. We first compute the Allen-Dynes critical temperature (T$_c$) for 4533 non-magnetic metals using a direct and progressively finer sampling of the electron-phonon couplings. For the candidates with the largest T$_c$, we use automated Wannierizations and electron-phonon interpolations to obtain a high-quality dataset for the most promising 250 dynamically stable structures, for which we calculate spectral functions, superconducting bandgaps, and isotropic Migdal-Eliashberg critical temperatures. For 140 of these, we also provide anisotropic Migdal-Eliashberg superconducting gaps and critical temperatures. The approach is remarkably successful in finding known superconductors, and we find 24 unknown ones with a predicted anisotropic T$_{\rm c}$ above 10~K. Among them, we identify a possible double gap superconductor (p-doped BaB$_2$), a non-magnetic half-Heusler ZrRuSb, and the perovskite TaRu$_3$C, all exhibiting significant T$_{\rm c}$. Finally, we introduce a sensitivity analysis to estimate the robustness of the predictions.

cond-mat.supr-con

The Wannier Function Software Ecosystem for Materials Simulations

Over the last two decades, following the early developments on maximally localized Wannier functions, an ecosystem of electronic-structure simulation techniques and software packages leveraging the Wannier representation has flourished. This environment includes codes to obtain Wannier functions and interfaces with first-principles simulation software, as well as an increasing number of related post-processing packages. Wannier functions can be obtained for isolated or extended systems (both crystalline and disordered), and can be used to understand chemical bonding, to characterize electric polarization, magnetization, and topology, or as an optimal basis set, providing very accurate interpolations in reciprocal space or large-scale Hamiltonians in real space. In this review, we summarize the current landscape of techniques, materials properties and simulation codes based on Wannier functions that have been made accessible to the research community, and that are now well integrated into what we term a \emph{Wannier function software ecosystem}. First, we introduce the theory and practicalities of Wannier functions, starting from their broad domains of applicability to advanced minimization methods using alternative approaches beyond maximal localization. Then we define the concept of a Wannier ecosystem and its interactions and interoperability with many quantum simulations engines and post-processing packages. We focus on some of the key properties and capabilities that are empowered by such ecosystem\textemdash from band interpolations and large-scale simulations to electronic transport, Berryology, topology, electron-phonon couplings, dynamical mean-field theory, embedding, and Koopmans functionals\textemdash concluding with the current status of interoperability and automation. [...]

cond-mat.mtrl-sci

Jupyter widgets and extensions for education and research in computational physics and chemistry

Interactive notebooks are a precious tool for creating graphical user interfaces and teaching materials. Python and Jupyter are becoming increasingly popular in this context, with Jupyter widgets at the core of the interactive functionalities. However, while packages and libraries which offer a broad range of general-purpose widgets exist, there is limited development of specialized widgets for computational physics, chemistry and materials science. This deficiency implies significant time investments for the development of effective Jupyter notebooks for research and education in these domains. Here, we present custom Jupyter widgets that we have developed to target the needs of these communities. These widgets constitute high-quality interactive graphical components and can be employed, for example, to visualize and manipulate data, or to explore different visual representations of concepts, clarifying the relationships existing between them. In addition, we discuss with one example how similar functionality can be exposed in the form of JupyterLab extensions, modifying the JupyterLab interface for an enhanced user experience when working with applications within the targeted scientific domains.

physics.ed-ph

Automated computational workflows for muon spin spectroscopy

Positive muon spin rotation and relaxation spectroscopy is a well established experimental technique for studying materials. It provides a local probe that generally complements scattering techniques in the study of magnetic systems and represents a valuable alternative for materials that display strong incoherent scattering or neutron absorption. Computational methods can effectively quantify the microscopic interactions underlying the experimentally observed signal, thus substantially boosting the predictive power of this technique. Here, we present an efficient set of algorithms and workflows devoted to the automation of this task. In particular, we adopt the so-called DFT+μ procedure, where the system is characterised in the density functional theory (DFT) framework with the muon modeled as a hydrogen impurity. We devise an automated strategy to obtain candidate muon stopping sites, their dipolar interaction with the nuclei, and hyperfine interactions with the electronic ground state. We validate the implementation on well-studied compounds, showing the effectiveness of our protocol in terms of accuracy and simplicity of use

physics.comp-ph