SearcharxivSearch

arXiv subjects

Martin Kuban

Publications and source records attributed to Martin Kuban.

7 recordsLinked to original sources

FUCrIMODo: structure recovery from atomistic descriptors via multi-stage genetic algorithms

Data-driven approaches to materials discovery rely on numerical representations of atomic structures as input for machine learning models. Inverting these descriptors - recovering atomic structures from their representations - is essential for most generative material design pipelines, yet it remains challenging, particularly for periodic systems. Existing inversion methods are either tailored to specific invertible descriptors or require candidate structures with similar atomic arrangements and compositions, limiting the exploration of novel regions in chemical and configurational space. Here, we propose a generalizable, similarity-driven sampling approach, powered by a novel stage-wise optimization strategy, to recover atom types, atomic positions, and unit cell shapes directly from a descriptor. Our approach requires only descriptor features and parameters as input without any prior structural knowledge. The capability of our method is demonstrated by the averaged Smooth Overlap of Atomic Positions (SOAP) descriptor.

cond-mat.mtrl-sci

An exciting approach to theoretical spectroscopy

Theoretical spectroscopy, and more generally, electronic-structure theory, are powerful concepts for describing the complex many-body interactions in materials. They comprise a variety of methods that can capture all aspects, from ground-state properties to lattice excitations to different types of light-matter interaction, including time-resolved variants. Modern electronic-structure codes implement either a few or several of these methods. Among them, exciting is an all-electron full-potential package that has a very rich portfolio of all levels of theory, with a particular focus on excitations. It implements the linearized augmented planewave plus local orbital (LAPW+LO) basis, which is known as the gold standard for solving the Kohn-Sham equations of density-functional theory (DFT). Based on this, it also offers benchmark-quality results for a wide range of excited-state methods. In this review, we provide a comprehensive overview of the features implemented in exciting in recent years, accompanied by short summaries on the state of the art of the underlying methodologies. They comprise DFT and time-dependent DFT (TDDFT), density-functional perturbation theory (DFPT) for phonons and electron-phonon coupling, and many-body perturbation theory in terms of the $GW$ approach and the Bethe-Salpeter equation (BSE). Moreover, exciting can handle resonant inelastic x-ray scattering (RIXS), pump-probe spectroscopy as well as exciton-phonon coupling (EXPC). Finally, we cover workflows and a view on data and machine learning (ML). All aspects are demonstrated with examples for scientifically relevant materials.

cond-mat.mtrl-sci

How big is Big Data?

Big data has ushered in a new wave of predictive power using machine learning models. In this work, we assess what {\it big} means in the context of typical materials-science machine-learning problems. This concerns not only data volume, but also data quality and veracity as much as infrastructure issues. With selected examples, we ask (i) how models generalize to similar datasets, (ii) how high-quality datasets can be gathered from heterogenous sources, (iii) how the feature set and complexity of a model can affect expressivity, and (iv) what infrastructure requirements are needed to create larger datasets and train models on them. In sum, we find that big data present unique challenges along very different aspects that should serve to motivate further work.

stat.ML

MADAS -- A Python framework for assessing similarity in materials-science data

Computational materials science produces large quantities of data, both in terms of high-throughput calculations and individual studies. Extracting knowledge from this large and heterogeneous pool of data is challenging due to the wide variety of computational methods and approximations, resulting in significant veracity in the sheer amount of available data. Here, we present MADAS, a Python framework for computing similarity relations between material properties. It can be used to automate the download of data from various sources, compute descriptors and similarities between materials, analyze the relationship between materials through their properties, and can incorporate a variety of existing machine learning methods. We explain the design of the package and demonstrate its power with representative examples.

cond-mat.mtrl-sci

CELL: a Python package for cluster expansion with a focus on complex alloys

We present the Python package CELL, which provides a modular approach to the cluster expansion (CE) method. CELL can treat a wide variety of substitutional systems, including one-, two-, and three-dimensional alloys, in a general multi-component and multi-sublattice framework. It is capable of dealing with complex materials comprising several atoms in their parent lattice. CELL uses state-of-the-art techniques for the construction of training data sets, model selection, and finite-temperature simulations. The user interface consists of well-documented Python classes and modules (http://sol.physik.hu-berlin.de/cell/). CELL also provides visualization utilities and can be interfaced with virtually any ab initio package, total-energy codes based on interatomic potentials, and more. The usage and capabilities of CELL are illustrated by a number of examples, comprising a Cu-Pt surface alloy with oxygen adsorption, featuring two coupled binary sublattices, and the thermodynamic analysis of its order-disorder transition; the demixing transition and lattice-constant bowing of the Si-Ge alloy; and an iterative CE approach for a complex clathrate compound with a parent lattice consisting of 54 atoms.

cond-mat.mtrl-sci

Similarity of materials and data-quality assessment by fingerprinting

Identifying similar materials, i.e., those sharing a certain property or feature, requires interoperable data of high quality. It also requires means to measure similarity. We demonstrate how a spectral fingerprint as a descriptor, combined with a similarity metric, can be used for establishing quantitative relationships between materials data, thereby serving multiple purposes. This concerns, for instance, the identification of materials exhibiting electronic properties similar to a chosen one. The same approach can be used for assessing uncertainty in data that potentially come from different sources. Selected examples show how to quantify differences between measured optical spectra or the impact of methodology and computational parameters on calculated properties, like the the density of states or excitonic spectra. Moreover, combining the same fingerprint with a clustering approach allows us to explore materials spaces in view of finding (un)expected trends or patterns. In all cases, we provide physical reasoning behind the findings of the automatized assessment of data.

cond-mat.mtrl-sci

Density-of-states similarity descriptor for unsupervised learning from materials data

We develop a materials descriptor based on the electronic density of states and investigate the similarity of materials based on it. As an application example, we study the Computational 2D Materials Database that hosts thousands of two-dimensional materials with their properties calculated by density-functional theory. Combining our descriptor with a clustering algorithm, we identify groups of materials with similar electronic structure. We characterize these clusters in terms of their crystal structure, their atomic composition, and the respective electronic configurations to rationalize the found (dis)similarities.

cond-mat.mtrl-sci