SearcharxivSearch

arXiv subjects

Shih-Han Wang

Publications and source records attributed to Shih-Han Wang.

4 recordsLinked to original sources

Decoding the Stability of Transition-Metal Alloys with Theory-infused Deep Learning

We introduce an interpretable deep learning framework that predicts the cohesive energy of transition-metal alloys (TMAs) by embedding cohesion theory within graph neural networks (GNNs). Beyond accurate prediction of cohesive energy, a key indicator of thermodynamic stability, the model offers mechanistic insights by disentangling energy contributions into physically meaningful components. These data-driven interpretations reveal periodic trends and stability principles governing transition metals. We apply the model to single-atom alloys (SAAs) to assess their thermodynamic resilience against two destabilizing processes: agglomeration (adatom clustering) and segregation (migration into the subsurface). Our analysis shows that these phenomena are governed by distinct physical factors-agglomeration is primarily influenced by localized d-orbital coupling, while segregation is dictated by delocalized effects such as wavefunction renormalization. This model thus serves as an explainable AI tool for understanding and guiding the design of stable TMAs, with implications for catalysis and materials discovery.

cond-mat.mtrl-sci

JARVIS-Leaderboard: A Large Scale Benchmark of Materials Design Methods

Lack of rigorous reproducibility and validation are major hurdles for scientific development across many fields. Materials science in particular encompasses a variety of experimental and theoretical approaches that require careful benchmarking. Leaderboard efforts have been developed previously to mitigate these issues. However, a comprehensive comparison and benchmarking on an integrated platform with multiple data modalities with both perfect and defect materials data is still lacking. This work introduces JARVIS-Leaderboard, an open-source and community-driven platform that facilitates benchmarking and enhances reproducibility. The platform allows users to set up benchmarks with custom tasks and enables contributions in the form of dataset, code, and meta-data submissions. We cover the following materials design categories: Artificial Intelligence (AI), Electronic Structure (ES), Force-fields (FF), Quantum Computation (QC) and Experiments (EXP). For AI, we cover several types of input data, including atomic structures, atomistic images, spectra, and text. For ES, we consider multiple ES approaches, software packages, pseudopotentials, materials, and properties, comparing results to experiment. For FF, we compare multiple approaches for material property predictions. For QC, we benchmark Hamiltonian simulations using various quantum algorithms and circuits. Finally, for experiments, we use the inter-laboratory approach to establish benchmarks. There are 1281 contributions to 274 benchmarks using 152 methods with more than 8 million data-points, and the leaderboard is continuously expanding. The JARVIS-Leaderboard is available at the website: https://pages.nist.gov/jarvis_leaderboard

cond-mat.mtrl-sci

Federated Learning Enables Big Data for Rare Cancer Boundary Detection

Although machine learning (ML) has shown promise in numerous domains, there are concerns about generalizability to out-of-sample data. This is currently addressed by centrally sharing ample, and importantly diverse, data from multiple sites. However, such centralization is challenging to scale (or even not feasible) due to various limitations. Federated ML (FL) provides an alternative to train accurate and generalizable ML models, by only sharing numerical model updates. Here we present findings from the largest FL study to-date, involving data from 71 healthcare institutions across 6 continents, to generate an automatic tumor boundary detector for the rare disease of glioblastoma, utilizing the largest dataset of such patients ever used in the literature (25,256 MRI scans from 6,314 patients). We demonstrate a 33% improvement over a publicly trained model to delineate the surgically targetable tumor, and 23% improvement over the tumor's entire extent. We anticipate our study to: 1) enable more studies in healthcare informed by large and diverse data, ensuring meaningful results for rare diseases and underrepresented populations, 2) facilitate further quantitative analyses for glioblastoma via performance optimization of our consensus model for eventual public release, and 3) demonstrate the effectiveness of FL at such scale and task complexity as a paradigm shift for multi-site collaborations, alleviating the need for data sharing.

cs.LG

Infusing Theory into Deep Learning for Interpretable Reactivity Prediction

Despite recent advances of data acquisition and algorithms development, machine learning (ML) faces tremendous challenges to being adopted in practical catalyst design, largely due to its limited generalizability and poor explainability. Here, we develop a theory-infused neural network (TinNet) approach that integrates deep learning algorithms with the well-established $d$-band theory of chemisorption for reactivity prediction of transition-metal surfaces. With simple adsorbates (e.g., *OH, *O, and *N) at active site ensembles as representative descriptor species, we demonstrate that the TinNet is on par with purely data-driven ML methods in prediction performance, while being inherently interpretable. Incorporation of scientific knowledge of physical interactions into learning from data sheds further light on the nature of chemical bonding and opens up new avenues for ML discovery of novel motifs with desired catalytic properties.

physics.chem-ph