SearcharxivSearch

arXiv subjects

Janine George

Publications and source records attributed to Janine George.

At least 19 recordsLinked to original sources

Representations from Pretrained Machine-Learning Interatomic Potentials as Coarse Coordinates for Material Generation and Evaluation

Generative machine learning is increasingly used for inorganic crystal structure generation. Most models and the corresponding evaluation approaches rely on simple forms of crystal structure representation. In this paper, we showcase the power of atom-averaged features from pretrained Machine-Learning Interatomic Potentials (MLIPs), such as MACE, for such tasks. We first introduce a distance measure that assesses the output of material generative models by capturing both quality and novelty in a single distribution-based evaluation framework. In particular, we introduce the Coarse-Fine Transport Distance (CFTD) using two different featurizers, where the quality component is based on coarse MACE features. We showcase CFTD's versatility in capturing crystal-structure quality while also detecting memorization, and compare it with the recently introduced continuous SUN metrics. We further show that coarse MACE features can be used as guidance for a material generative model.

cs.LG

Spin-Polarized Electronic Structure and Chemical Bonding Data for 2,500+ Halide Double Perovskites

Halide double perovskites (A$_2$BB'X$_6$) are a long-known class of materials that has recently been rediscovered for diverse applications, including photovoltaics, photocatalysis, and radiation detection. Their doubled unit cell provides immense chemical tunability, allowing the incorporation of magnetic ions and enabling access to a wide range of electronic-structure features, including different band-edge characters, alignments, and symmetries. Magnetic elements may further introduce spin degrees of freedom and magnetic behaviour, thereby broadening the functional landscape of these compounds. Here, we present the first comprehensive database of spin-polarised electronic-structure data for all halide double perovskites predicted to be stable by the recently introduced $\tau$ tolerance factor by Bartel et al. The dataset focuses on the Cs$_2$BB'X$_6$ family, with X = I, Br, Cl, and F, and includes density of states (DOS) for $>$2,500 compounds, calculated using hybrid-functional density functional theory. Among these, 719 compounds exhibit band gaps in the visible range and 118 display half-metallic character. In addition, we provide chemical-bonding analysis using \textsc{lobster}, which provides insights into orbital interactions across the dataset. To facilitate exploration, we further offer UMAP-based visualisations and an interactive app for systematic investigation of chemical composition, electronic structure, and magnetic properties.

cond-mat.mtrl-sci

Parameter-Efficient Fine-Tuning of Machine-Learning Interatomic Potentials for Phonon and Thermal Properties

Machine-learning interatomic potentials are widely used as computationally efficient surrogates for density functional theory in atomistic simulations, enabling large-scale, long-time modeling of materials systems. We investigate how different fine-tuning strategies influence the prediction of harmonic phonon band structures, thermal properties, and the potential energy surface along imaginary phonon modes. We achieve substantial accuracy improvements with minimal additional data, with as few as 10 additional training structures already yielding significant gains. In addition to existing approaches, we introduce Equitrain, a finetuning framework that implements LoRA-based adaptation. Across 53 materials systems, we show that fine-tuned models consistently outperform both the underlying pretrained model and models trained from scratch. Equitrain achieves the best overall performance, and our results demonstrate that fine-tuning enables accurate phonon predictions.

cond-mat.mtrl-sci

Perspective: Towards sustainable exploration of chemical spaces with machine learning

Artificial intelligence is transforming molecular and materials science, but its growing computational and data demands raise critical sustainability challenges. In this Perspective, we examine resource considerations across the AI-driven discovery pipeline--from quantum-mechanical (QM) data generation and model training to automated, self-driving research workflows--building on discussions from the ``SusML workshop: Towards sustainable exploration of chemical spaces with machine learning'' held in Dresden, Germany. In this context, the availability of large quantum datasets has enabled rigorous benchmarking and rapid methodological progress, while also incurring substantial energy and infrastructure costs. We highlight emerging strategies to enhance efficiency, including general-purpose machine learning (ML) models, multi-fidelity approaches, model distillation, and active learning. Moreover, incorporating physics-based constraints within hierarchical workflows, where fast ML surrogates are applied broadly and high-accuracy QM methods are used selectively, can further optimize resource use without compromising reliability. Equally important is bridging the gap between idealized computational predictions and real-world conditions by accounting for synthesizability and multi-objective design criteria, which is essential for practical impact. Finally, we argue that sustainable progress will rely on open data and models, reusable workflows, and domain-specific AI systems that maximize scientific value per unit of computation, enabling efficient and responsible discovery of technological materials and therapeutics.

cs.LG

A critical assessment of bonding descriptors for predicting materials properties

Most machine learning models for materials science rely on descriptors based on materials compositions and structures, even though the chemical bond has been proven to be a valuable concept for predicting materials properties. Over the years, various theoretical frameworks have been developed to characterize bonding in solid-state materials. However, integrating bonding information from these frameworks into machine learning pipelines at scale has been limited by the lack of a systematically generated and validated database. Recent advances in high-throughput bonding analysis workflows have addressed this issue, and our previously computed Quantum-Chemical Bonding Database for Solid-State Materials was extended to include approximately 13,000 materials. This database is then used to derive a new set of quantum-chemical bonding descriptors. A systematic assessment is performed using statistical significance tests to evaluate how the inclusion of these descriptors influences the performance of machine-learning models that otherwise rely solely on structure- and composition-derived features. Models are built to predict elastic, vibrational, and thermodynamic properties typically associated with chemical bonding in materials. The results demonstrate that incorporating quantum-chemical bonding descriptors not only improves predictive performance but also helps identify intuitive expressions for properties such as the projected force constant and lattice thermal conductivity via symbolic regression.

cond-mat.mtrl-sci

Transport Novelty Distance: A Distributional Metric for Evaluating Material Generative Models

Recent advances in generative machine learning have opened new possibilities for the discovery and design of novel materials. However, as these models become more sophisticated, the need for rigorous and meaningful evaluation metrics has grown. Existing evaluation approaches often fail to capture both the quality and novelty of generated structures, limiting our ability to assess true generative performance. In this paper, we introduce the Transport Novelty Distance (TNovD) to judge generative models used for materials discovery jointly by the quality and novelty of the generated materials. Based on ideas from Optimal Transport theory, TNovD uses a coupling between the features of the training and generated sets, which is refined into a quality and memorization regime by a threshold. The features are generated from crystal structures using a graph neural network that is trained to distinguish between materials, their augmented counterparts, and differently sized supercells using contrastive learning. We evaluate our proposed metric on typical toy experiments relevant for crystal structure prediction, including memorization, noise injection and lattice deformations. Additionally, we validate the TNovD on the MP20 validation set and the WBM substitution dataset, demonstrating that it is capable of detecting both memorized and low-quality material data. We also benchmark the performance of several popular material generative models. While introduced for materials, our TNovD framework is domain-agnostic and can be adapted for other areas, such as images and molecules.

cond-mat.mtrl-sci

Thermal Transport in Ag8TS6 (T= Si, Ge, Sn) Argyrodites: An Integrated Experimental, Quantum-Chemical, and Computational Modelling Study

Argyrodite-type Ag-based sulfides combine exceptionally low lattice thermal and high ionic conductivity, making them promising candidates for thermoelectric and solid-state energy applications. In this work, we studied Ag8TS6 (T= Si, Ge, Sn) argyrodite family by combining chemical-bonding analysis, lattice vibrational properties simulation, and experimental measurements to investigate their structural and thermal transport properties. Furthermore, we propose a two-channel lattice-dynamics model based on Gr\"uneisen-derived phonon lifetimes and compare it to an approach using machine-learned interatomic potentials. Both approaches are able to predict thermal conductivity in agreement with experimental lattice thermal conductivities along the whole temperature range, highlighting their potential suitability for future high-throughput predictions. Our findings also reveal a relationship between bond heterogeneity arising from weakly bonded Ag+ ions and occupied antibonding states in Ag-S and Ag-Ag interactions and strong anharmonicity, including large Gr\"uneisen parameters, and low sound velocities, which are responsible for the low lattice thermal conductivity of Ag8SnS6, Ag8GeS6, and Ag8SiS6. We furthermore show that thermal and ionic conductivities in all three compounds are independent of each other and can likely be tuned individually.

cond-mat.mtrl-sci

Medium-range structural order in amorphous arsenic

Medium-range order (MRO) is a key structural feature of amorphous materials, but its origin and nature remain elusive. Here, we reveal the MRO in amorphous arsenic (a-As) using advanced atomistic simulations, based on machine-learned potentials derived using automated workflows. Our simulations accurately reproduce the experimental structure factor of a-As, especially the first sharp diffraction peak (FSDP), which is a signature of MRO. We compare and contrast the structure of a-As with that of its lighter homologue, red amorphous phosphorus (a-P), identifying the dihedral-angle distribution as a key factor differentiating the MRO in both. The pressure-dependent structural behaviors of a-As and a-P differ as well, which we link to the interplay of ring topology and structural entropy. We finally show that the origin of the FSDP is closely correlated with the size and spatial distribution of voids in the amorphous networks. Our work provides fundamental insights into MRO in an amorphous elemental system, and more widely it illustrates the usefulness of automation for machine-learning-driven atomistic simulations.

cond-mat.mtrl-sci

A Python workflow definition for computational materials design

Numerous Workflow Management Systems (WfMS) have been developed in the field of computational materials science with different workflow formats, hindering interoperability and reproducibility of workflows in the field. To address this challenge, we introduce here the Python Workflow Definition (PWD) as a workflow exchange format to share workflows between Python-based WfMS, currently AiiDA, jobflow, and pyiron. This development is motivated by the similarity of these three Python-based WfMS, that represent the different workflow steps and data transferred between them as nodes and edges in a graph. With the PWD, we aim at fostering the interoperability and reproducibility between the different WfMS in the context of Findable, Accessible, Interoperable, Reusable (FAIR) workflows. To separate the scientific from the technical complexity, the PWD consists of three components: (1) a conda environment that specifies the software dependencies, (2) a Python module that contains the Python functions represented as nodes in the workflow graph, and (3) a workflow graph stored in the JavaScript Object Notation (JSON). The first version of the PWD supports directed acyclic graph (DAG)-based workflows. Thus, any DAG-based workflow defined in one of the three WfMS can be exported to the PWD and afterwards imported from the PWD to one of the other WfMS. After the import, the input parameters of the workflow can be adjusted and computing resources can be assigned to the workflow, before it is executed with the selected WfMS. This import from and export to the PWD is enabled by the PWD Python library that implements the PWD in AiiDA, jobflow, and pyiron.

cs.SE

34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery

Large Language Models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 34 total projects developed during the second annual Large Language Model Hackathon for Applications in Materials Science and Chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.

cs.LG

An automated framework for exploring and learning potential-energy surfaces

Machine learning has become ubiquitous in materials modelling and now routinely enables large-scale atomistic simulations with quantum-mechanical accuracy. However, developing machine-learned interatomic potentials requires high-quality training data, and the manual generation and curation of such data can be a major bottleneck. Here, we introduce an automated framework for the exploration and fitting of potential-energy surfaces, implemented in an openly available software package that we call autoplex (`automatic potential-landscape explorer'). We discuss design choices, particularly the interoperability with existing software architectures, and the ability for the end user to easily use the computational workflows provided. We show wide-ranging capability demonstrations: for the titanium-oxygen system, SiO2, crystalline and liquid water, as well as phase-change memory materials. More generally, our study illustrates how automation can speed up atomistic machine learning -- with a long-term vision of making it a genuine mainstream tool in physics, chemistry, and materials science.

physics.comp-ph

Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

Here, we present the outcomes from the second Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry, which engaged participants across global hybrid locations, resulting in 34 team submissions. The submissions spanned seven key application areas and demonstrated the diverse utility of LLMs for applications in (1) molecular and material property prediction; (2) molecular and material design; (3) automation and novel interfaces; (4) scientific communication and education; (5) research data management and automation; (6) hypothesis generation and evaluation; and (7) knowledge extraction and reasoning from scientific literature. Each team submission is presented in a summary table with links to the code and as brief papers in the appendix. Beyond team results, we discuss the hackathon event and its hybrid format, which included physical hubs in Toronto, Montreal, San Francisco, Berlin, Lausanne, and Tokyo, alongside a global online hub to enable local and virtual collaboration. Overall, the event highlighted significant improvements in LLM capabilities since the previous year's hackathon, suggesting continued expansion of LLMs for applications in materials science and chemistry research. These outcomes demonstrate the dual utility of LLMs as both multipurpose models for diverse machine learning tasks and platforms for rapid prototyping custom applications in scientific research.

cs.LG

SynCoTrain: A Dual Classifier PU-learning Framework for Synthesizability Prediction

Material discovery is a cornerstone of modern science, driving advancements in diverse disciplines from biomedical technology to climate solutions. Predicting synthesizability, a critical factor in realizing novel materials, remains a complex challenge due to the limitations of traditional heuristics and thermodynamic proxies. While stability metrics such as formation energy offer partial insights, they fail to account for kinetic factors and technological constraints that influence synthesis outcomes. These challenges are further compounded by the scarcity of negative data, as failed synthesis attempts are often unpublished or context-specific. We present SynCoTrain, a semi-supervised machine learning model designed to predict the synthesizability of materials. SynCoTrain employs a co-training framework leveraging two complementary graph convolutional neural networks: SchNet and ALIGNN. By iteratively exchanging predictions between classifiers, SynCoTrain mitigates model bias and enhances generalizability. Our approach uses Positive and Unlabeled (PU) Learning to address the absence of explicit negative data, iteratively refining predictions through collaborative learning. The model demonstrates robust performance, achieving high recall on internal and leave-out test sets. By focusing on oxide crystals, a well-characterized material family with extensive experimental data, we establish SynCoTrain as a reliable tool for predicting synthesizability while balancing dataset variability and computational efficiency. This work highlights the potential of co-training to advance high-throughput materials discovery and generative research, offering a scalable solution to the challenge of synthesizability prediction.

cond-mat.mtrl-sci

A foundation model for atomistic materials chemistry

Atomistic simulations of matter, especially those that leverage first-principles (ab initio) electronic structure theory, provide a microscopic view of the world, underpinning much of our understanding of chemistry and materials science. Over the last decade or so, machine-learned force fields have transformed atomistic modeling by enabling simulations of ab initio quality over unprecedented time and length scales. However, early ML force fields have largely been limited by: (i) the substantial computational and human effort of developing and validating potentials for each particular system of interest; and (ii) a general lack of transferability from one chemical system to the next. Here we show that it is possible to create a general-purpose atomistic ML model, trained on a public dataset of moderate size, that is capable of running stable molecular dynamics for a wide range of molecules and materials. We demonstrate the power of the MACE-MP-0 model - and its qualitative and at times quantitative accuracy - on a diverse set of problems in the physical sciences, including properties of solids, liquids, gases, chemical reactions, interfaces and even the dynamics of a small protein. The model can be applied out of the box as a starting or "foundation" model for any atomistic system of interest and, when desired, can be fine-tuned on just a handful of application-specific data points to reach ab initio accuracy. Establishing that a stable force-field model can cover almost all materials changes atomistic modeling in a fundamental way: experienced users get reliable results much faster, and beginners face a lower barrier to entry. Foundation models thus represent a step towards democratising the revolution in atomic-scale modeling that has been brought about by ML force fields.

physics.chem-ph

Glass fracture surface energy calculated from crystal structure and bond-energy data

We present a novel method to predict the fracture surface energy, γ, of isochemically crystallizing silicate glasses using readily available crystallographic structure data of their crystalline counterpart and tabled diatomic chemical bond energies, D0. The method assumes that γ equals the fracture surface energy of the most likely cleavage plane of the crystal. Calculated values were in excellent agreement with those calculated from glass density, network connectivity and D0 data in earlier work. This finding demonstrates a remarkable equivalence between crystal cleavage planes and glass fracture surfaces.

cond-mat.mtrl-sci

Determination of acoustic phonon anharmonicities via second-order Raman scattering in CuI

We demonstrate the determination of anharmonic acoustic phonon properties via second-order Raman scattering exemplarily on copper iodide single crystals. The origin of multi-phonon features from the second-order Raman spectra was assigned by the support of the calculated 2-phonon density of states. In this way, the temperature dependence of acoustic phonons was determined down to 10\,K. To determine independently the harmonic contributions of respective acoustic phonons, density functional theory (DFT) in quasi-harmonic approximation was used. Finally, the anharmonic contributions were determined. The results are in agreement with earlier publications and extend CuI's determined acoustic phonon properties to lower temperatures with higher accuracy. This approach demonstrates that it is possible to characterize the acoustic anharmonicities via Raman scattering down to zero-temperature renormalization constants of at least 0.1\,cm$^{-1}$.

cond-mat.mtrl-sci

A Quantum-Chemical Bonding Database for Solid-State Materials

An in-depth insight into the chemistry and nature of the individual chemical bonds is essential for understanding materials. Bonding analysis is thus expected to provide important features for large-scale data analysis and machine learning of material properties. Such chemical bonding information can be computed using the LOBSTER software package, which post-processes modern density functional theory data by projecting the plane wave-based wave functions onto a local, atomic orbital basis. With the help of a fully automatic workflow, the VASP and LOBSTER software packages are used to generate the data. We then perform bonding analyses on 1520 compounds (insulators and semiconductors) and provide the results as a database. The database structure of the bonding analysis database, which allows easy data retrieval, is also explained. The projected densities of states and bonding indicators are benchmarked on standard density-functional theory computations and available heuristics, respectively. Lastly, we illustrate the predictive power of bonding descriptors by constructing a machine-learning model for phononic properties, which shows an increase in prediction accuracies by 27 % (mean absolute errors) compared to a benchmark model differing only by not relying on any quantum-chemical bonding features.

cond-mat.mtrl-sci

"Ultima Ratio": Simulating wide-range X-ray scattering and diffraction

We demonstrate a strategy for simulating wide-range X-ray scattering patterns, which spans the small- and wide scattering angles as well as the scattering angles typically used for Pair Distribution Function (PDF) analysis. Such simulated patterns can be used to test holistic analysis models, and, since the diffraction intensity is on the same scale as the scattering intensity, may offer a novel pathway for determining the degree of crystallinity. The "Ultima Ratio" strategy is demonstrated on a 64-nm Metal Organic Framework (MOF) particle, calculated from Q < 0.01 1/nm up to Q < 150 1/nm, with a resolution of 0.16 Angstrom. The computations exploit a modified 3D Fast Fourier Transform (3D-FFT), whose modifications enable the transformations of matrices at least up to 8000^3 voxels in size. Multiple of these modified 3D-FFTs are combined to improve the low-Q behaviour. The resulting curve is compared to a wide-range scattering pattern measured on a polydisperse MOF powder. While computationally intensive, the approach is expected to be useful for simulating scattering from a wide range of realistic, complex structures, from (poly-)crystalline particles to hierarchical, multicomponent structures such as viruses and catalysts.

cond-mat.mtrl-sci