SearcharxivSearch

arXiv subjects

Martin Hoffmann Petersen

Publications and source records attributed to Martin Hoffmann Petersen.

8 recordsLinked to original sources

Virp: neural network-accelerated prediction of physical properties in site-disordered materials

Among metallic alloys, ceramics, and even common compounds such as water ice, it is usual to find materials in which crystalline order is expressed as a probability. In such cases, one or more sites within a crystal can be occupied by multiple elements or vacancies, according to a set of probabilities. These crystal structures remain inaccessible to common first-principles materials simulation methodologies, which assumes perfect crystal order. Workaround strategies to this limitation include quasirandom structures and cluster expansion. These methods are system-specific and computationally expensive as they rely on large scale Monte Carlo simulations of enlarged unit cells. To address these limitations, we propose a pipeline combining a permutation-based virtual cell generation algorithm, sampling regime, and thermodynamic post-processing which greatly improves the feasibility of computation analyses for site-disordered materials. We demonstrate that the massive configurational space can be adequately sampled with 400 virtual cells, as long as the supercell definition is sufficiently large.

cond-mat.mtrl-sci

Fast and Accurate Prediction of Lattice Thermal Conductivity via Machine Learning Surrogates

The appearance of generative models has opened vast chemical spaces in the design of functional materials. Although machine learning interatomic potentials (MLIPs) have substantially accelerated phonon calculations, high-fidelity prediction of lattice thermal conductivity \k{appa}lat still requires accurate treatment of anharmonic interactions, which remains a key challenge for existing potentials across novel chemical spaces. To address this challenge, we present a comprehensive benchmark of 15 surrogate models for predicting \k{appa}lat using the Phonix database, which contains 6,966 entries with anharmonic phonon properties derived from first-principles calculations. Firstly, We categorize these surrogate models into three distinct groups: Physical-informed feature descriptors combined with ML models, end-to-end deep neural networks, and pre-trained MLIP-embeddings combined with ML models. By evaluating model performance across random, space-group disjoint (testing generalization to unseen crystal symmetries), and Out-Of-Distribution splits (OOD dataset that testing extrapolation to property regimes beyond the training range) based on \k{appa}lat, we probe both interpolation and exploration capabilities. Our results reveal that MLIP-embedded models excel in interpolation within well-sampled regions, deep neural network models especially ALiEGNN demonstrate superior robustness in OOD regimes critical for discovering novel low-\k{appa}lat. Additionally, we find a systematic degradation in performance when the structural representation is reduced. Although surrogate models exhibit lower accuracy than direct simulations using first-principles calculation, they reduce computational costs by orders of magnitude, enabling efficient high-throughput screening of thermoelectric materials with minimal loss in generative design workflows.

cond-mat.mtrl-sci

Navigating Order-(Dis)Order Family Trees via Group-Subgroup Transitions

As closed-loop materials discovery systems scale to produce millions of candidate compounds, the credibility of the novelty they reward becomes a critical concern. Novelty is commonly assessed against databases of ordered crystal structures, in which atomic sites are fully occupied. Yet, a predicted ordered structure may simply correspond to a particular ordering of a known disordered phase, whose sites are occupied by multiple species in the statistical average structure; we refer to such a structure as an ordered child of a disordered parent. Here, we introduce order-(dis)order family trees, a symmetry-based framework that organizes ordered and disordered structures through group-subgroup relations and enables novelty to be explicitly evaluated. We develop a high-throughput family matching procedure, to identify possible disordered parents and symmetry-related ordered relatives for a given ordered structure. As validation, we test our framework on synthesis-facing case studies (A-Lab), where it correctly recovers existing disordered parents for the targeted ordered structures. Extending this family-tree-based benchmark to experimental structure databases (ICSD), computational datasets (MP-20, Alex-MP-20, and GNoME), and crystal generative models further reveals that many ordered structures that appear novel as individual entries are, in fact, better understood as members of experimentally known order-(dis)order family trees. We also show that this is particularly evident in symmetry-agnostic all-atom generative models, which more frequently produce ordered structures derived from known disordered parents, whereas symmetry-constrained models are 2-4x less prone to this behavior. Our results establish order-(dis)order family trees as a key requirement for achieving genuine novelty in data-driven materials discovery.

cond-mat.mtrl-sci

SWORD: Symmetry and Wyckoff-sequence of Ordered and Disordered crystals

Novelty in materials discovery requires candidates to be distinct, non-redundant, and thermodynamically plausible. While crystallographic databases continue to expand in both size and complexity, making efficient and reliable novelty assessment has become increasingly difficult. This becomes particularly acute when crystallographic disorder is involved, as partial occupancies greatly enlarge the structure-composition space and obscure the identification of genuinely distinct structures. Here, we introduce SWORD, a symmetry-aware, Wyckoff-based string representation compatible with both ordered and disordered crystals. SWORD provides (i) standardization of symmetry-equivalent structural descriptions into a consistent label, (ii) explicitly represents co-occupying species on partially occupied sites, and (iii) quantifies complex disorder through a degree of mixing descriptor that captures continuous variation in site stoichiometry. These features enable efficient structure grouping, duplicate identification, and finer refinement of disordered structures. Benchmarking against existing fingerprint and structure-matching methods shows that SWORD remains invariant under identity-preserving transformations while retaining interpretable sensitivity to structural perturbations. In addition, SWORD shows competitive performance in associating unrelaxed and intermediate configurations with their final relaxed states along relaxation trajectories. This feature could enable more reliable novelty assessment directly from partially relaxed or even unrelaxed generated structures. Finally, SWORD was used to showcase its capability of disorder-aware database-scale deduplication and curation for the Inorganic Crystal Structure Database (ICSD). The curated ICSD would serve as the basis for the materials informatics and data-driven materials design in the era of artificial intelligence.

cond-mat.mtrl-sci

Importance of Electronic Entropy for Machine Learning Interatomic Potentials

Machine learning interatomic potentials (MLIPs) enable large-scale atomistic simulations but remain challenged in describing mixed-valence materials where charge ordering strongly influences thermodynamic stability. Here we investigate the role of electronic entropy in MLIP structural optimization of the battery cathode material \ce{NaFePO4}. We show that conventional MLIPs fail to reproduce the correct stability of intermediate \ce{Na} concentrations because structural optimization leads to incorrect \ce{Fe^{2+}}/\ce{Fe^{3+}} charge assignments, resulting in erroneous energy ordering and convex-hull predictions. Analysis of magnetic moments during structural optimization reveals that MLIPs are unable to capture electronic entropy associated with charge ordering. To address this limitation, we introduce an approach that embeds charge-state information directly into the MLIP representation by distinguishing between \ce{Fe^{2+}} and \ce{Fe^{3+}} environments during training. Retraining CHGNet, cPaiNN, and MACE with this representation enables accurate structural optimization, correct identification of charge ordering, and improved agreement with density functional theory convex hulls. Our results demonstrate that incorporating electronic entropy into MLIP representations is essential for modeling charge-disordered materials and provide a practical framework for extending MLIP simulations to mixed-valence transition-metal systems.

cond-mat.mtrl-sci

Energy Underprediction from Symmetry in Machine-Learning Interatomic Potentials

Machine learning interatomic potentials (MLIAPs) have emerged as powerful tools for accelerating materials simulations with near-density functional theory (DFT) accuracy. However, despite significant advances, we identify a critical yet overlooked issue undermining their reliability: a systematic energy underprediction. This problem becomes starkly evident in large-scale thermodynamic stability assessments. By performing over 12 million calculations using nine MLIAPs for over 150,000 inorganic crystals in the Materials Project, we demonstrate that most frontier models consistently underpredict energy above hull (Ehull), a key metric for thermodynamic stability, total energy, and formation energy, despite the fact that over 90\% of test structures (DFT-relaxed) are in the training data. The mean absolute errors (MAE) for Ehull exceed ~30 meV/atom even by the best model, directly challenging claims of achieving ``DFT accuracy'' for property predictions central to materials discovery, especially related to (meta-)stability. Crucially, we trace this underprediction to insufficient handling of symmetry degrees of freedom (DOF), constituting both lattice symmetry and Wyckoff site symmetries for the space group. MLIAPs exhibit pronounced errors (MAE for Ehull $>$ ~40 meV/atom) in structures with high symmetry DOF, where subtle atomic displacements significantly impact energy landscapes. Further analysis also indicates that the MLIAPs show severe energy underprediction for a large proportion of near-hull materials. We argue for improvements on symmetry-aware models such as explicit DOF encoding or symmetry-regularized loss functions, and more robust MLIAPs for predicting crystal properties where the preservation and breaking of symmetry are pivotal.

cond-mat.mtrl-sci

Dis-GEN: Disordered crystal structure generation

A wide range of synthesized crystalline inorganic materials exhibit compositional disorder, where multiple atomic species partially occupy the same crystallographic site. As a result, the physical and chemical properties of such materials are dependent on how the atomic species are distributed among the corresponding symmetrical sites, making them exceptionally challenging to model using computational methods. For this reason, existing generative models cannot handle the complexities of disordered inorganic crystals. To address this gap, we introduce Dis-GEN, a generative model based on an empirical equivariant representation, derived from theoretical crystallography methodology. Dis-GEN is capable of generating symmetry-consistent structures that accommodate both compositional disorder and vacancies. The model is uniquely trained on experimental structures from the Inorganic Crystal Structure Database (ICSD) - the world's largest database of identified inorganic crystal structures. We demonstrate that Dis-GEN can effectively generate disordered inorganic materials while preserving crystallographic symmetry throughout the generation process. This approach provides a critical check point for the systematic exploration and discovery of disordered functional materials, expanding the scope of generative modeling in materials science.

cond-mat.mtrl-sci

Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

Here, we present the outcomes from the second Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry, which engaged participants across global hybrid locations, resulting in 34 team submissions. The submissions spanned seven key application areas and demonstrated the diverse utility of LLMs for applications in (1) molecular and material property prediction; (2) molecular and material design; (3) automation and novel interfaces; (4) scientific communication and education; (5) research data management and automation; (6) hypothesis generation and evaluation; and (7) knowledge extraction and reasoning from scientific literature. Each team submission is presented in a summary table with links to the code and as brief papers in the appendix. Beyond team results, we discuss the hackathon event and its hybrid format, which included physical hubs in Toronto, Montreal, San Francisco, Berlin, Lausanne, and Tokyo, alongside a global online hub to enable local and virtual collaboration. Overall, the event highlighted significant improvements in LLM capabilities since the previous year's hackathon, suggesting continued expansion of LLMs for applications in materials science and chemistry research. These outcomes demonstrate the dual utility of LLMs as both multipurpose models for diverse machine learning tasks and platforms for rapid prototyping custom applications in scientific research.

cs.LG