SearcharxivSearch

arXiv subjects

Weike Ye

Publications and source records attributed to Weike Ye.

11 recordsLinked to original sources

Spectra-Scope : A toolkit for automated and interpretable characterization of material properties from spectral data

Spectroscopy is a central pillar of materials characterization, providing useful information on properties like structure, composition, or excited state dynamics of a system. However, many spectroscopic techniques present challenges in the development of interpretable, performant, and reliable supervised learning models due to the wide range of possible nonlinear correlations that can exist between the signal and the response variable (target) of interest. Here, we present Spectra-Scope, an open-source AutoML framework for automatic characterization of material properties from spectroscopy data using interpretable machine learning (ML) models. The software is implemented in Python and a no-code web application. It comprises tools for data preprocessing, nonlinear feature extraction, machine learning model training, and feature selection. Users can easily train different types of simple, interpretable ML models on a set of feature transformations quickly and with modest computational resources. In this work, we outline the methods of Spectra-Scope and its effectiveness across diverse datasets, with applications to materials and agricultural spectroscopy data. We show that Spectra-Scope can reproduce performance of comparable models in the literature, and highlight how our emphasis on interpretability can be used to rationalize the behavior of individual models and understand the physical processes behind spectral features.

cond-mat.mtrl-sci

XxaCT-NN: Structure Agnostic Multimodal Learning for Materials Science

Recent advances in materials discovery have been driven by structure-based models, particularly those using crystal graphs. While effective for computational datasets, these models are impractical for real-world applications where atomic structures are often unknown or difficult to obtain. We propose a scalable multimodal framework that learns directly from elemental composition and X-ray diffraction (XRD) -- two of the more available modalities in experimental workflows without requiring crystal structure input. Our architecture integrates modality-specific encoders with a cross-attention fusion module and is trained on the 5-million-sample Alexandria dataset. We present masked XRD modeling (MXM), and apply MXM and contrastive alignment as self-supervised pretraining strategies. Pretraining yields faster convergence (up to 4.2x speedup) and improves both accuracy and representation quality. We further demonstrate that multimodal performance scales more favorably with dataset size than unimodal baselines, with gains compounding at larger data regimes. Our results establish a path toward structure-free, experimentally grounded foundation models for materials science.

cs.LG

UniMat: Unifying Materials Embeddings through Multi-modal Learning

Materials science datasets are inherently heterogeneous and are available in different modalities such as characterization spectra, atomic structures, microscopic images, and text-based synthesis conditions. The advancements in multi-modal learning, particularly in vision and language models, have opened new avenues for integrating data in different forms. In this work, we evaluate common techniques in multi-modal learning (alignment and fusion) in unifying some of the most important modalities in materials science: atomic structure, X-ray diffraction patterns (XRD), and composition. We show that structure graph modality can be enhanced by aligning with XRD patterns. Additionally, we show that aligning and fusing more experimentally accessible data formats, such as XRD patterns and compositions, can create more robust joint embeddings than individual modalities across various tasks. This lays the groundwork for future studies aiming to exploit the full potential of multi-modal data in materials science, facilitating more informed decision-making in materials design and discovery.

cs.LG

De novo Design of Polymer Electrolytes with High Conductivity using GPT-based and Diffusion-based Generative Models

Solid polymer electrolytes hold significant promise as materials for next-generation batteries due to their superior safety performance, enhanced specific energy, and extended lifespans compared to liquid electrolytes. However, the material's low ionic conductivity impedes its commercialization, and the vast polymer space poses significant challenges for the screening and design. In this study, we assess the capabilities of generative artificial intelligence (AI) for the de novo design of polymer electrolytes. To optimize the generation, we compare different deep learning architectures, including both GPT-based and diffusion-based models, and benchmark the results with hyperparameter tuning. We further employ various evaluation metrics and full-atom molecular dynamics simulations to assess the performance of different generative model architectures and to validate the top candidates produced by each model. Out of only 45 candidates being tested, we discovered 17 polymers that achieve superior ionic conductivity better than any other polymers in our database, with some of them doubling the conductivity value. In addition, by adopting a pretraining and fine-tuning methodology, we significantly improve the efficacy of our generative models, achieving quicker convergence, enhanced performance with limited data, and greater diversity. Using the proposed method, we can easily generate a large number of novel, diverse, and valid polymers, with a chance of synthesizability, enabling us to identify promising candidates with markedly improved efficiency.

physics.chem-ph

A Self-Improvable Polymer Discovery Framework Based on Conditional Generative Model

In this work, we introduce a polymer discovery platform to efficiently design polymers with tailored properties, exemplified by the discovery of high-performance polymer electrolytes. The platform integrates three core components: a conditioned generative model, a computational evaluation module, and a feedback mechanism, creating a self-improving system for material innovation. To demonstrate the efficacy of this platform, it is used to design polymer electrolyte materials with high ionic conductivity. A simple conditional generative model, based on the minGPT architecture, can effectively generate candidate polymers that exhibit a mean ionic conductivity that is significantly greater than those in the original training set. This approach, coupled with molecular dynamics simulations (MD) for testing and a specifically planned acquisition mechanism, allows the platform to refine its output iteratively. Notably, after the first iteration, we observed an increase in both the mean and the lower bound of the ionic conductivity of the new polymer candidates. The platform's effectiveness is underscored by the identification of 14 polymer repeating units, each displaying a computed ionic conductivity surpassing that of Polyethylene Oxide (PEO). The performance of these polymers in MD simulations verifies the platform's efficacy in generating potential polymer candidate materials. Acknowledging current limitations, future work will focus on enhancing modeling techniques, evaluation processes, and acquisition strategies, aiming for broader applicability in polymer science and machine learning.

physics.chem-ph

The Role of Reference Points in Machine-Learned Atomistic Simulation Models

This paper introduces the Chemical Environment Modeling Theory (CEMT), a novel, generalized framework designed to overcome the limitations inherent in traditional atom-centered Machine Learning Force Field (MLFF) models, widely used in atomistic simulations of chemical systems. CEMT demonstrated enhanced flexibility and adaptability by allowing reference points to exist anywhere within the modeled domain and thus, enabling the study of various model architectures. Utilizing Gaussian Multipole (GMP) featurization functions, several models with different reference point sets, including finite difference grid-centered and bond-centered models, were tested to analyze the variance in capabilities intrinsic to models built on distinct reference points. The results underscore the potential of non-atom-centered reference points in force training, revealing variations in prediction accuracy, inference speed and learning efficiency. Finally, a unique connection between CEMT and real-space orbital-free finite element Density Functional Theory (FE-DFT) is established, and the implications include the enhancement of data efficiency and robustness. It allows the leveraging of spatially-resolved energy densities and charge densities from FE-DFT calculations, as well as serving as a pivotal step towards integrating known quantum-mechanical laws into the architecture of ML models.

physics.chem-ph

A Universal Machine Learning Model for Elemental Grain Boundary Energies

The grain boundary (GB) energy has a profound influence on the grain growth and properties of polycrystalline metals. Here, we show that the energy of a GB, normalized by the bulk cohesive energy, can be described purely by four geometric features. By machine learning on a large computed database of 361 small $Σ$ ($Σ< 10$) GBs of more than 50 metals, we develop a model that can predict the grain boundary energies to within a mean absolute error of 0.13 J m$^{-2}$. More importantly, this universal GB energy model can be extrapolated to the energies of high $Σ$ GBs without loss in accuracy. These results highlight the importance of capturing fundamental scaling physics and domain knowledge in the design of interpretable, extrapolatable machine learning models for materials science.

cond-mat.mtrl-sci

Accelerating Materials Discovery with Bayesian Optimization and Graph Deep Learning

Machine learning (ML) models utilizing structure-based features provide an efficient means for accurate property predictions across diverse chemical spaces. However, obtaining equilibrium crystal structures typically requires expensive density functional theory (DFT) calculations, which limits ML-based exploration to either known crystals or a small number of hypothetical crystals. Here, we demonstrate that the application of Bayesian optimization with symmetry constraints using a graph deep learning energy model can be used to perform "DFT-free" relaxations of crystal structures. Using this approach to significantly improve the accuracy of ML-predicted formation energies and elastic moduli of hypothetical crystals, two novel ultra-incompressible hard materials MoWC2 (P63/mmc) and ReWB (Pca21) were identified and successfully synthesized via in-situ reactive spark plasma sintering from a screening of 399,960 transition metal borides and carbides. This work addresses a critical bottleneck to accurate property predictions for hypothetical materials, paving the way to ML-accelerated discovery of new materials with exceptional properties.

cond-mat.mtrl-sci

Learning Properties of Ordered and Disordered Materials from Multi-fidelity Data

Predicting the properties of a material from the arrangement of its atoms is a fundamental goal in materials science. While machine learning has emerged in recent years as a new paradigm to provide rapid predictions of materials properties, their practical utility is limited by the scarcity of high-fidelity data. Here, we develop multi-fidelity graph networks as a universal approach to achieve accurate predictions of materials properties with small data sizes. As a proof of concept, we show that the inclusion of low-fidelity Perdew-Burke-Ernzerhof band gaps greatly enhances the resolution of latent structural features in materials graphs, leading to a 22-45\% decrease in the mean absolute errors of experimental band gap predictions. We further demonstrate that learned elemental embeddings in materials graph networks provide a natural approach to model disorder in materials, addressing a fundamental gap in the computational prediction of materials properties.

cond-mat.mtrl-sci

Graph Networks as a Universal Machine Learning Framework for Molecules and Crystals

Graph networks are a new machine learning (ML) paradigm that supports both relational reasoning and combinatorial generalization. Here, we develop universal MatErials Graph Network (MEGNet) models for accurate property prediction in both molecules and crystals. We demonstrate that the MEGNet models outperform prior ML models such as the SchNet in 11 out of 13 properties of the QM9 molecule data set. Similarly, we show that MEGNet models trained on $\sim 60,000$ crystals in the Materials Project substantially outperform prior ML models in the prediction of the formation energies, band gaps and elastic moduli of crystals, achieving better than DFT accuracy over a much larger data set. We present two new strategies to address data limitations common in materials science and chemistry. First, we demonstrate a physically-intuitive approach to unify four separate molecular MEGNet models for the internal energy at 0 K and room temperature, enthalpy and Gibbs free energy into a single free energy MEGNet model by incorporating the temperature, pressure and entropy as global state inputs. Second, we show that the learned element embeddings in MEGNet models encode periodic chemical trends and can be transfer-learned from a property model trained on a larger data set (formation energies) to improve property models with smaller amounts of data (band gaps and elastic moduli).

cond-mat.mtrl-sci

Deep Neural Networks for Accurate Predictions of Garnet Stability

Predicting the stability of crystals is one of the central problems in materials science. Today, density functional theory (DFT) calculations are the computational tool of choice to obtain energies of crystals with quantitative accuracy. Despite algorithmic and computing advances, DFT calculations remain comparatively expensive and scale poorly with system size. Here we show that deep neural networks utilizing just two descriptors - the Pauling electronegativity and ionic radii - can predict the DFT formation energies of C3A2D3O12 garnets with extremely low mean absolute errors of 7-8 meV/atom, an order of magnitude improvement over previous machine learning models and well within the limits of DFT accuracy. Further extension to mixed garnets with little loss in accuracy can be achieved using a binary encoding scheme that introduces minimal increase in descriptor dimensionality. Our results demonstrate that generalizable deep-learning models for quantitative crystal stability prediction can be built on a small set of chemically-intuitive descriptors. Such models provide the means to rapidly transverse vast chemical spaces to accurately identify stable compositions, accelerating the discovery of novel materials with potentially superior properties.

cond-mat.mtrl-sci