SearcharxivSearch

arXiv subjects

Lauri Himanen

Publications and source records attributed to Lauri Himanen.

7 recordsLinked to original sources

ML-guided screening of chalcogenide perovskites as solar energy materials

Chalcogenide perovskites have emerged as promising absorber materials for next-generation photovoltaic devices, yet their experimental realization remains limited by competing phases, structural polymorphism, and synthetic challenges. Here, we present a fully data-driven and experimentally grounded screening and ranking framework to assess the stability and experimental feasibility of chalcogenide perovskites, integrating interpretable analytical descriptors, machine-learning models, and sustainability metrics. Using a curated experimental dataset of halide and chalcogenide compounds, we derive a new tolerance factor via the SISSO (sure independence screening and sparsifying operator) algorithm that more accurately distinguishes perovskite-forming compositions than established tolerance-factor-based screening criteria. This descriptor is combined with generative crystal structure prediction, composition-based bandgap estimation, and machine-learning-based feasibility assessment to systematically explore a wide chemical space of hypothetical chalcogenide perovskites. The resulting candidates are further evaluated using sustainability indicators, enabling multi-objective ranking tailored to both single-junction and tandem photovoltaic architectures. Beyond identifying several promising and previously unexplored chalcogenide perovskites, this work demonstrates a transferable screening strategy for chemically constrained materials spaces that balances optoelectronic performance, experimental viability, and long-term sustainability.

cond-mat.mtrl-sci

An autonomous living database for perovskite photovoltaics

Scientific discovery is severely bottlenecked by the inability of manual curation to keep pace with exponential publication rates. This creates a widening knowledge gap. This is especially stark in photovoltaics, where the leading database for perovskite solar cells has been stagnant since 2021 despite massive ongoing research output. Here, we resolve this challenge by establishing an autonomous, self-updating living database (PERLA). Our pipeline integrates large language models with physics-aware validation to extract complex device data from the continuous literature stream, achieving human-level precision (>90%) and eliminating annotator variance. By employing this system on the previously inaccessible post-2021 literature, we uncover critical evolutionary trends hidden by data lag: the field has decisively shifted toward inverted architectures employing self-assembled monolayers and formamidinium-rich compositions, driving a clear trajectory of sustained voltage loss reduction. PERLA transforms static publications into dynamic knowledge resources that enable data-driven discovery to operate at the speed of publication.

cond-mat.mtrl-sci

Updates to the DScribe Library: New Descriptors and Derivatives

We present an update of the DScribe package, a Python library for atomistic descriptors. The update extends DScribe's descriptor selection with the Valle-Oganov materials fingerprint and provides descriptor derivatives to enable more advanced machine learning tasks, such as force prediction and structure optimization. For all descriptors, numeric derivatives are now available in DSribe. For the many-body tensor representation (MBTR) and the Smooth Overlap of Atomic Positions (SOAP), we have also implemented analytic derivatives. We demonstrate the effectiveness of the descriptor derivatives for machine learning models of Cu clusters and perovskite alloys.

cond-mat.mtrl-sci

Data-driven materials science: status, challenges and perspectives

Data-driven science is heralded as a new paradigm in materials science. In this field, data is the new resource, and knowledge is extracted from materials data sets that are too big or complex for traditional human reasoning - typically with the intent to discover new or improved materials or materials phenomena. Multiple factors, including the open science movement, national funding, and progress in information technology, have fueled its development. Such related tools as materials databases, machine learning, and high-throughput methods are now established as parts of the materials research toolset. However, there are a variety of challenges that impede progress in data-driven materials science: data veracity, integration of experimental and computational data, data longevity, standardization, and the gap between industrial interests and academic efforts. In this perspective article, we discuss the historical development and current state of data-driven materials science, building from the early evolution of open science to the rapid expansion of materials data infrastructures. We also review key successes and challenges so far, providing a perspective on the future development of the field.

physics.comp-ph

DScribe: Library of Descriptors for Machine Learning in Materials Science

DScribe is a software package for machine learning that provides popular feature transformations ("descriptors") for atomistic materials simulations. DScribe accelerates the application of machine learning for atomistic property prediction by providing user-friendly, off-the-shelf descriptor implementations. The package currently contains implementations for Coulomb matrix, Ewald sum matrix, sine matrix, Many-body Tensor Representation (MBTR), Atom-centered Symmetry Function (ACSF) and Smooth Overlap of Atomic Positions (SOAP). Usage of the package is illustrated for two different applications: formation energy prediction for solids and ionic charge prediction for atoms in organic molecules. The package is freely available under the open-source Apache License 2.0.

cond-mat.mtrl-sci

Chemical diversity in molecular orbital energy predictions with kernel ridge regression

Instant machine learning predictions of molecular properties are desirable for materials design, but the predictive power of the methodology is mainly tested on well-known benchmark datasets. Here, we investigate the performance of machine learning with kernel ridge regression (KRR) for the prediction of molecular orbital energies on three large datasets: the standard QM9 small organic molecules set, amino acid and dipeptide conformers, and organic crystal-forming molecules extracted from the Cambridge Structural Database. We focus on prediction of highest occupied molecular orbital (HOMO) energies, computed at density-functional level of theory. Two different representations that encode molecular structure are compared: the Coulomb matrix (CM) and the many-body tensor representation (MBTR). We find that KRR performance depends significantly on the chemistry of the underlying dataset and that the MBTR is superior to the CM, predicting HOMO energies with a mean absolute error as low as 0.09 eV. To demonstrate the power of our machine learning method, we apply our model to structures of 10k previously unseen molecules. We gain instant energy predictions that allow us to identify interesting molecules for future applications.

physics.chem-ph

Database-driven high-throughput study for hybrid perovskite coating materials

We developed a high-throughput screening scheme to acquire candidate coating materials for hybrid perovskites. From more than 1.8 million entries of an inorganic compound database, we collected 93 binary and ternary materials with promising properties for protectively coating halide-perovskite photoabsorbers in perovskite solar cells. These candidates fulfill a series of criteria, including wide band gaps, abundant and non-toxic elements, water-insoluble, and small lattice mismatch with surface models of halide perovskites.

physics.app-ph