SearcharxivSearch

arXiv subjects

Maria K. Y. Chan

Publications and source records attributed to Maria K. Y. Chan.

At least 19 recordsLinked to original sources

Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature

X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through fragmented textual context in the literature. Here, we use multimodal (image and text) literature mining to transform this dispersed knowledge into an AI-ready experimental data resource. We developed a scalable spectroscopy data digitization pipeline that identifies XAS figures in full-text articles, digitizes spectral curves, and links each spectrum to accompanying metadata on the measured edge and material. Applying this pipeline to the battery literature produced an open dataset of 13,740 XAS spectra, spanning 66 absorbing elements and diverse battery chemistries, with expert validation confirming accurate extraction of spectral and metadata information. By converting literature-embedded spectra into structured numerical data, this dataset provides a foundation for large-scale XAS analysis, cross-laboratory comparison, high-throughput characterization, and autonomous discovery of advanced materials.

cond-mat.mtrl-sci

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research

AI co-scientists are increasingly used for scientific discovery, but current evaluations still do not test them on a key task: moving from a concrete scientific or technological problem to a plausible, mechanism-grounded solution hypothesis. This gap is especially important in materials science and, in particular, battery research, where a useful proposal must identify the relevant failure mode, propose a credible intervention, and explain why that intervention should improve the target property. We introduce Matter to Mechanism, a benchmark for evaluating AI co-scientists on problem-to-hypothesis reasoning in materials science, with a focus on battery materials research. The benchmark contains 2,645 instances derived from scientific publications. Each instance includes a structured problem statement, a candidate solution hypothesis, an explicit reasoning trace, and domain-grounded annotations such as material system, component, failure mode, intervention, mechanism, target property, and claimed outcome. We also introduce a metric suite that measures reasoning fidelity, problem alignment, mechanistic specificity, novelty, plausibility, and problem decomposition quality, and combine them into a composite score. Using this framework, we evaluate several AI co-scientist systems and show that Matter to Mechanism reveals interpretable system differences that are only partially recovered by standard text-similarity metrics. We further show through adversarial stress tests that the aggregate score is more stable than individual metric dimensions under superficial gaming attacks.

cs.CE

Robust interpretation of electrochemical impedance spectra using numerical complex analysis

Electrochemical Impedance Spectroscopy (EIS) is a non-invasive technique widely used for understanding charge transfer and charge transport processes in electrochemical systems and devices. Standard approaches for the interpretation of EIS data involve starting with a hypothetical circuit model for the physical processes in the device based on experience/intuition, and then fitting the EIS data to this circuit model. This work explores a mathematical approach for extracting key characteristic features from EIS data by relying on fundamental principles of complex analysis. These characteristic features can ascertain the presence of inductors and constant phase elements (non-ideal capacitors) in circuit models and enable us to answer questions about the identifiability and uniqueness of equivalent circuit models. In certain scenarios such as models with only resistors and capacitors, we are able to enumerate all possible families of circuit models. Finally, we apply the mathematical framework presented here to real-world electrochemical systems and highlight results using impedance measurements from a lithium-ion battery coin cell.

physics.app-ph

Search for Active and Inactive Ion Insertion Sites in Organic Crystalline Materials

The position of mobile active and inactive ions, specifically ion insertion sites, within organic crystals, significantly affects the properties of organic materials used for energy storage and ionic transport. Identifying the positions of these atom (and ion) sites in an organic crystal is difficult, especially when the element has a low X-ray scattering power, such as lithium (Li) and hydrogen, which are difficult to detect with powder X-ray diffraction (XRD) methods. First-principles calculations, exemplified by density functional theory (DFT), are very effective for confirming the relative stability of ion positions in materials. However, the lack of effective strategies to identify ion sites in these organic crystalline frameworks renders this task extremely challenging. This work presents two algorithms: (i) Efficient Location of Ion Insertion Sites from Extrema in electrostatic local potential and charge density (ELIISE), and (ii) ElectRostatic InsertioN (ERIN), which leverage charge density and electrostatic potential fields accessed from first-principles calculations, combined with the Simultaneous Ion Insertion and Evaluation (SIIE) workflow -- that inserts all ions simultaneously -- to determine ion positions in organic crystals. We demonstrate that these methods accurately reproduce known ion positions in 16 organic materials and also identify previously overlooked low-energy sites in tetralithium 2,6-naphthalenedicarboxylate (Li$_4$NDC), an organic electrode material, highlighting the importance of inserting all ions simultaneously as done in the SIIE workflow.

cond-mat.mtrl-sci

Robust Machine Learning Inference from X-ray Absorption Near Edge Spectra through Featurization

X-ray absorption spectroscopy (XAS) is a commonly-employed technique for characterizing functional materials. In particular, x-ray absorption near edge spectra (XANES) encodes local coordination and electronic information and machine learning approaches to extract this information is of significant interest. To date, most ML approaches for XANES have primarily focused on using the raw spectral intensities as input, overlooking the potential benefits of incorporating spectral transformations and dimensionality reduction techniques into ML predictions. In this work, we focused on systematically comparing the impact of different featurization methods on the performance of ML models for XAS analysis. We evaluated the classification and regression capabilities of these models on computed datasets and validated their performance on previously unseen experimental datasets. Our analysis revealed an intriguing discovery: the cumulative distribution function (CDF) feature achieves both high prediction accuracy and exceptional transferability. This remarkably robust performance can be attributed to its tolerance to horizontal shifts in spectra, which is crucial when validating models using experimental data. While this work exclusively focuses on XANES analysis, we anticipate that the methodology presented here will hold promise as a versatile asset to the broader spectroscopy community.

physics.comp-ph

Revealing Local Structures through Machine-Learning- Fused Multimodal Spectroscopy

Atomistic structures of materials offer valuable insights into their functionality. Determining these structures remains a fundamental challenge in materials science, especially for systems with defects. While both experimental and computational methods exist, each has limitations in resolving nanoscale structures. Core-level spectroscopies, such as x-ray absorption (XAS) or electron energy-loss spectroscopies (EELS), have been used to determine the local bonding environment and structure of materials. Recently, machine learning (ML) methods have been applied to extract structural and bonding information from XAS/EELS, but most of these frameworks rely on a single data stream, which is often insufficient. In this work, we address this challenge by integrating multimodal ab initio simulations, experimental data acquisition, and ML techniques for structure characterization. Our goal is to determine local structures and properties using EELS and XAS data from multiple elements and edges. To showcase our approach, we use various lithium nickel manganese cobalt (NMC) oxide compounds which are used for lithium ion batteries, including those with oxygen vacancies and antisite defects, as the sample material system. We successfully inferred local element content, ranging from lithium to transition metals, with quantitative agreement with experimental data. Beyond improving prediction accuracy, we find that ML model based on multimodal spectroscopic data is able to determine whether local defects such as oxygen vacancy and antisites are present, a task which is impossible for single mode spectra or other experimental techniques. Furthermore, our framework is able to provide physical interpretability, bridging spectroscopy with the local atomic and electronic structures.

cond-mat.mtrl-sci

Deterministic Creation of Identical Monochromatic Quantum Emitters in Hexagonal Boron Nitride

Deterministic creation of quantum emitters with high single-photon-purity and excellent indistinguishability is essential for practical applications in quantum information science. Many successful attempts have been carried out in hexagonal boron nitride showing its capability of hosting room temperature quantum emitters. However, most of the existing methods produce emitters with heterogeneous optical properties and unclear creation mechanisms. Here, the authors report a deterministic creation of identical room temperature quantum emitters using masked-carbon-ion implantation on freestanding hBN flakes. Quantum emitters fabricated by our approach showed thermally limited monochromaticity with an emission center wavelength distribution of 590.7 +- 2.7 nm, a narrow full width half maximum of 7.1 +- 1.7 nm, excellent brightness (1MHz emission rate), and extraordinary stability. Our method provides a reliable platform for characterization and fabrication research on hBN based quantum emitters, helping to reveal the origins of the single-photon-emission behavior in hBN and favoring practical applications, especially the industrial-scale production of quantum technology.

cond-mat.mtrl-sci

Data-driven discovery of dynamics from time-resolved coherent scattering

Coherent X-ray scattering (CXS) techniques are capable of interrogating dynamics of nano- to mesoscale materials systems at time scales spanning several orders of magnitude. However, obtaining accurate theoretical descriptions of complex dynamics is often limited by one or more factors -- the ability to visualize dynamics in real space, computational cost of high-fidelity simulations, and effectiveness of approximate or phenomenological models. In this work, we develop a data-driven framework to uncover mechanistic models of dynamics directly from time-resolved CXS measurements without solving the phase reconstruction problem for the entire time series of diffraction patterns. Our approach uses neural differential equations to parameterize unknown real-space dynamics and implements a computational scattering forward model to relate real-space predictions to reciprocal-space observations. This method is shown to recover the dynamics of several computational model systems under various simulated conditions of measurement resolution and noise. Moreover, the trained model enables estimation of long-term dynamics well beyond the maximum observation time, which can be used to inform and refine experimental parameters in practice. Finally, we demonstrate an experimental proof-of-concept by applying our framework to recover the probe trajectory from a ptychographic scan. Our proposed framework bridges the wide existing gap between approximate models and complex data.

cond-mat.mtrl-sci

Machine learning for classifying and interpreting coherent X-ray speckle patterns

Speckle patterns produced by coherent X-ray have a close relationship with the internal structure of materials but quantitative inversion of the relationship to determine structure from speckle patterns is challenging. Here, we investigate the link between coherent X-ray speckle patterns and sample structures using a model 2D disk system and explore the ability of machine learning to learn aspects of the relationship. Specifically, we train a deep neural network to classify the coherent X-ray speckle patterns according to the disk number density in the corresponding structure. It is demonstrated that the classification system is accurate for both non-disperse and disperse size distributions.

cond-mat.mtrl-sci

Ensemble of Pre-Trained Neural Networks for Segmentation and Quality Detection of Transmission Electron Microscopy Images

Automated analysis of electron microscopy datasets poses multiple challenges, such as limitation in the size of the training dataset, variation in data distribution induced by variation in sample quality and experiment conditions, etc. It is crucial for the trained model to continue to provide acceptable segmentation/classification performance on new data, and quantify the uncertainty associated with its predictions. Among the broad applications of machine learning, various approaches have been adopted to quantify uncertainty, such as Bayesian modeling, Monte Carlo dropout, ensembles, etc. With the aim of addressing the challenges specific to the data domain of electron microscopy, two different types of ensembles of pre-trained neural networks were implemented in this work. The ensembles performed semantic segmentation of ice crystal within a two-phase mixture, thereby tracking its phase transformation to water. The first ensemble (EA) is composed of U-net style networks having different underlying architectures, whereas the second series of ensembles (ER-i) are composed of randomly initialized U-net style networks, wherein each base learner has the same underlying architecture 'i'. The encoders of the base learners were pre-trained on the Imagenet dataset. The performance of EA and ER were evaluated on three different metrics: accuracy, calibration, and uncertainty. It is seen that EA exhibits a greater classification accuracy and is better calibrated, as compared to ER. While the uncertainty quantification of these two types of ensembles are comparable, the uncertainty scores exhibited by ER were found to be dependent on the specific architecture of its base member ('i') and not consistently better than EA. Thus, the challenges posed for the analysis of electron microscopy datasets appear to be better addressed by an ensemble design like EA, as compared to an ensemble design like ER.

cond-mat.mtrl-sci

Machine learning for impurity charge-state transition levels in semiconductors from elemental properties using multi-fidelity datasets

Quantifying charge-state transition energy levels of impurities in semiconductors is critical to understanding and engineering their optoelectronic properties for applications ranging from solar photovoltaics to infrared lasers. While these transition levels can be measured and calculated accurately, such efforts are time-consuming and more rapid prediction methods would be beneficial. Here, we significantly reduce the time typically required to predict impurity transition levels using multi-fidelity datasets and a machine learning approach employing features based on elemental properties and impurity positions. We use transition levels obtained from low-fidelity (i.e., local-density approximation or generalized gradient approximation) density functional theory (DFT) calculations, corrected using a recently proposed modified band alignment scheme, which well-approximates transition levels from high-fidelity DFT (i.e., hybrid HSE06). The model fit to the large multi-fidelity database shows improved accuracy compared to the models trained on the more limited high-fidelity values. Crucially, in our approach, when using the multi-fidelity data, high-fidelity values are not required for model training, significantly reducing the computational cost required for training the model. Our machine learning model of transition levels has a root mean squared (mean absolute) error of 0.36 (0.27) eV vs high-fidelity hybrid functional values when averaged over 14 semiconductor systems from the II-VI and III-V families. As a guide for use on other systems, we assessed the model on simulated data to show the expected accuracy level as a function of bandgap for new materials of interest. Finally, we use the model to predict a complete space of impurity charge-state transition levels in all zinc blende III-V and II-VI systems.

cond-mat.mtrl-sci

Ultrafast formation of transient 2D diamond-like structure in twisted bilayer graphene

Due to the absence of matching carbon atoms at honeycomb centers with carbon atoms in adjacent graphene sheets, theorists predicted that a sliding process is needed to form AA, AB, or ABC stacking when directly converting graphite into sp3 bonded diamond. Here, using twisted bilayer graphene, which naturally provides AA and AB stacking configurations, we report the ultrafast formation of a transient 2D diamond-like structure (which is not observed in aligned graphene) under femtosecond laser irradiation. This photo-induced phase transition is evidenced by the appearance of new bond lengths of 1.94A and 3.14A in the time-dependent differential pair distribution function using MeV ultrafast electron diffraction. Molecular dynamics and first principles calculation indicate that sp3 bonds nucleate at AA and AB stacked areas in moire pattern. This work sheds light on the direct graphite-to-diamond transformation mechanism, which has not been fully understood for more than 60 years.

cond-mat.mtrl-sci

A Two-stage Framework for Compound Figure Separation

Scientific literature contains large volumes of complex, unstructured figures that are compound in nature (i.e. composed of multiple images, graphs, and drawings). Separation of these compound figures is critical for information retrieval from these figures. In this paper, we propose a new strategy for compound figure separation, which decomposes the compound figures into constituent subfigures while preserving the association between the subfigures and their respective caption components. We propose a two-stage framework to address the proposed compound figure separation problem. In particular, the subfigure label detection module detects all subfigure labels in the first stage. Then, in the subfigure detection module, the detected subfigure labels help to detect the subfigures by optimizing the feature selection process and providing the global layout information as extra features. Extensive experiments are conducted to validate the effectiveness and superiority of the proposed framework, which improves the detection precision by 9%.

cs.CV

Data-Driven Design of Novel Halide Perovskite Alloys

The great tunability of the properties of halide perovskites presents new opportunities for optoelectronic applications as well as significant challenges associated with exploring combinatorial chemical spaces. In this work, we develop a framework powered by high-throughput computations and machine learning for the design and prediction of mixed cation halide perovskite alloys. In a chemical space of ABX$_{3}$ perovskites with a selected set of options for A, B, and X atoms, pseudo-cubic structures of compounds with B-site mixing are simulated using density functional theory (DFT) and several properties are computed, including stability, lattice constant, band gap, vacancy formation energy, refractive index, and optical absorption spectrum, using both semi-local and hybrid functionals. Neural networks (NN) are used to train predictive models for every property using tabulated elemental properties of A, B, and X site atoms as descriptors. Starting from the DFT dataset of 229 points, we use the trained NN models to predict the structural, energetic, electronic and optical properties of a complete dataset of 17,955 compounds, and perform high-throughput screening in terms of stability, band gap and defect tolerance, to obtain 574 promising compounds that are ranked as potential absorbers according to their photovoltaic figure of merit. Compositional trends in the screened set of attractive mixed cation halide perovskites are revealed and additional computations are performed on selected compounds. The data-driven design framework developed here is promising for designing novel mixed compositions and can be extended to a wider perovskite chemical space in terms of A, B, and X atoms, different kinds of mixing at the A, B, or X sites, non-cubic phases, and other properties of interest.

cond-mat.mtrl-sci

Plot2Spectra: an Automatic Spectra Extraction Tool

Different types of spectroscopies, such as X-ray absorption near edge structure (XANES) and Raman spectroscopy, play a very important role in analyzing the characteristics of different materials. In scientific literature, XANES/Raman data are usually plotted in line graphs which is a visually appropriate way to represent the information when the end-user is a human reader. However, such graphs are not conducive to direct programmatic analysis due to the lack of automatic tools. In this paper, we develop a plot digitizer, named Plot2Spectra, to extract data points from spectroscopy graph images in an automatic fashion, which makes it possible for large scale data acquisition and analysis. Specifically, the plot digitizer is a two-stage framework. In the first axis alignment stage, we adopt an anchor-free detector to detect the plot region and then refine the detected bounding boxes with an edge-based constraint to locate the position of two axes. We also apply scene text detector to extract and interpret all tick information below the x-axis. In the second plot data extraction stage, we first employ semantic segmentation to separate pixels belonging to plot lines from the background, and from there, incorporate optical flow constraints to the plot line pixels to assign them to the appropriate line (data instance) they encode. Extensive experiments are conducted to validate the effectiveness of the proposed plot digitizer, which shows that such a tool could help accelerate the discovery and machine learning of materials properties.

cs.CV

Ingrained -- An automated framework for fusing atomic-scale image simulations into experiments

To fully leverage the power of image simulation to corroborate and explain patterns and structures in atomic resolution microscopy (e.g., electron and scanning probe), an initial correspondence between the simulation and experimental image must be established at the outset of further high accuracy simulations or calculations. Furthermore, if simulation is to be used in context of highly automated processes or high-throughput optimization, the process of finding this correspondence itself must be automated. In this work, we introduce ingrained, an open-source automation framework which solves for this correspondence and fuses atomic resolution image simulations into the experimental images to which they correspond. We describe herein the overall ingrained workflow, focusing on its application to interface structure approximations, and the development of an experimentally rationalized forward model for scanning tunneling microscopy simulation.

cond-mat.mtrl-sci

EXSCLAIM! -- An automated pipeline for the construction of labeled materials imaging datasets from literature

Due to recent improvements in image resolution and acquisition speed, materials microscopy is experiencing an explosion of published imaging data. The standard publication format, while sufficient for traditional data ingestion scenarios where a select number of images can be critically examined and curated manually, is not conducive to large-scale data aggregation or analysis, hindering data sharing and reuse. Most images in publications are presented as components of a larger figure with their explicit context buried in the main body or caption text, so even if aggregated, collections of images with weak or no digitized contextual labels have limited value. To solve the problem of curating labeled microscopy data from literature, this work introduces the EXSCLAIM! Python toolkit for the automatic EXtraction, Separation, and Caption-based natural Language Annotation of IMages from scientific literature. We highlight the methodology behind the construction of EXSCLAIM! and demonstrate its ability to extract and label open-source scientific images at high volume.

cs.IR

Defect Physics of Pseudo-cubic Mixed Halide Lead Perovskites from First Principles

Owing to the increasing popularity of lead-based hybrid perovskites for photovoltaic (PV) applications, it is crucial to understand their defect physics and its influence on their optoelectronic properties. In this work, we simulate various point defects in pseudo-cubic structures of mixed iodide-bromide and bromide-chloride methylammonium lead perovskites with the general formula MAPbI_{3-y}Br_{y} or MAPbBr_{3-y}Cl_{y} (where y is between 0 and 3), and use first principles based density functional theory computations to study their relative formation energies and charge transition levels. We identify vacancy defects and Pb on MA anti-site defect as the lowest energy native defects in each perovskite. We observe that while the low energy defects in all MAPbI_{3-y}Br_{y} systems only create shallow transition levels, the Br or Cl vacancy defects in the Cl-containing pervoskites have low energy and form deep levels which become deeper for higher Cl content. Further, we study extrinsic substitution by different elements at the Pb site in MAPbBr_{3}, MAPbCl_{3} and the 50-50 mixed halide perovskite, MAPbBr_{1.5}Cl_{1.5}, and identify some transition metals that create lower energy defects than the dominant intrinsic defects and also create mid-gap charge transition levels.

physics.app-ph