SearcharxivSearch

arXiv subjects

Tong-yi Zhang

Publications and source records attributed to Tong-yi Zhang.

5 recordsLinked to original sources

From Literature to Lab: Closed-Loop Advancement of Perovskite Solar Cells via Domain Knowledge Guided LLM

Perovskite solar cells (PSCs) have been considered as a next-generation disruptive photovoltaic technology, yet their advancement is constrained by the complexity of perovskite recipe with high-dimensional material and process design space. Despite the impressive general reasoning of Large Language Models (LLMs), they struggle with two limitations for application in PSCs: an inability to align general semantics with the perovskite domain knowledge, and an inefficiency in navigating high-dimensional perovskite material and recipe design spaces. To address these limitations, we introduce a domain-knowledge-guided framework PVK-LLM, a specialized model to serve as an expert to bridge general semantics with perovskite domain knowledge. By integrating this domain knowledge into a hierarchical Bayesian Optimization workflow, our approach efficiently navigates the high-dimension design space on a solar cell simulator platform. The domain knowledge resolves cold-start problems while dynamically adapting to simulator feedback. Moreover, in an individual wet-lab experiment aimed at maximizing power conversion efficiency (PCE), our framework autonomously proposes a novel synergistic four-component recipe comprising specialized organic passivation recipe (3MTPAI, PDAI2, EDAI2, and PipDI) which has not been reported in existing literature. This AI-designed recipe effectively achieves a champion PCE value of over 26.0 %, approaching world records achieved through extensive expert trial-and-error. Our approach can effectively enable LLM comprehend the domain knowledge, which can efficiently navigate in a high-dimensional, capable to accelerate the advancement in real-world perovskite as well as other material science development.

cond-mat.mtrl-sci

Perovskite-LLM: Knowledge-Enhanced Large Language Models for Perovskite Solar Cell Research

The rapid advancement of perovskite solar cells (PSCs) has led to an exponential growth in research publications, creating an urgent need for efficient knowledge management and reasoning systems in this domain. We present a comprehensive knowledge-enhanced system for PSCs that integrates three key components. First, we develop Perovskite-KG, a domain-specific knowledge graph constructed from 1,517 research papers, containing 23,789 entities and 22,272 relationships. Second, we create two complementary datasets: Perovskite-Chat, comprising 55,101 high-quality question-answer pairs generated through a novel multi-agent framework, and Perovskite-Reasoning, containing 2,217 carefully curated materials science problems. Third, we introduce two specialized large language models: Perovskite-Chat-LLM for domain-specific knowledge assistance and Perovskite-Reasoning-LLM for scientific reasoning tasks. Experimental results demonstrate that our system significantly outperforms existing models in both domain-specific knowledge retrieval and scientific reasoning tasks, providing researchers with effective tools for literature review, experimental design, and complex problem-solving in PSC research.

cs.AI

Materials Generation in the Era of Artificial Intelligence: A Comprehensive Survey

Materials are the foundation of modern society, underpinning advancements in energy, electronics, healthcare, transportation, and infrastructure. The ability to discover and design new materials with tailored properties is critical to solving some of the most pressing global challenges. In recent years, the growing availability of high-quality materials data combined with rapid advances in Artificial Intelligence (AI) has opened new opportunities for accelerating materials discovery. Data-driven generative models provide a powerful tool for materials design by directly create novel materials that satisfy predefined property requirements. Despite the proliferation of related work, there remains a notable lack of up-to-date and systematic surveys in this area. To fill this gap, this paper provides a comprehensive overview of recent progress in AI-driven materials generation. We first organize various types of materials and illustrate multiple representations of crystalline materials. We then provide a detailed summary and taxonomy of current AI-driven materials generation approaches. Furthermore, we discuss the common evaluation metrics and summarize open-source codes and benchmark datasets. Finally, we conclude with potential future directions and challenges in this fast-growing field. The related sources can be found at https://github.com/ZhixunLEE/Awesome-AI-for-Materials-Generation.

cond-mat.mtrl-sci

opXRD: Open Experimental Powder X-ray Diffraction Database

Powder X-ray diffraction (pXRD) experiments are a cornerstone for materials structure characterization. Despite their widespread application, analyzing pXRD diffractograms still presents a significant challenge to automation and a bottleneck in high-throughput discovery in self-driving labs. Machine learning promises to resolve this bottleneck by enabling automated powder diffraction analysis. A notable difficulty in applying machine learning to this domain is the lack of sufficiently sized experimental datasets, which has constrained researchers to train primarily on simulated data. However, models trained on simulated pXRD patterns showed limited generalization to experimental patterns, particularly for low-quality experimental patterns with high noise levels and elevated backgrounds. With the Open Experimental Powder X-Ray Diffraction Database (opXRD), we provide an openly available and easily accessible dataset of labeled and unlabeled experimental powder diffractograms. Labeled opXRD data can be used to evaluate the performance of models on experimental data and unlabeled opXRD data can help improve the performance of models on experimental data, e.g. through transfer learning methods. We collected 92552 diffractograms, 2179 of them labeled, from a wide spectrum of materials classes. We hope this ongoing effort can guide machine learning research toward fully automated analysis of pXRD data and thus enable future self-driving materials labs.

cond-mat.mtrl-sci

SimXRD-4M: Big Simulated X-ray Diffraction Data Accelerate the Crystal Symmetry Classification

Spectroscopic data, particularly diffraction data, contain detailed crystal and microstructure information and thus are crucial for materials discovery. Powder X-ray diffraction (XRD) patterns are greatly effective in identifying crystals. Although machine learning (ML) has significantly advanced the analysis of powder XRD patterns, the progress is hindered by a lack of training data. To address this, we introduce SimXRD, the largest open-source simulated XRD pattern dataset so far, to accelerate the development of crystallographic informatics. SimXRD comprises 4,065,346 simulated powder X-ray diffraction patterns, representing 119,569 distinct crystal structures under 33 simulated conditions that mimic real-world variations. We find that the crystal symmetry inherently follows a long-tailed distribution and evaluate 21 sequence learning models on SimXRD. The results indicate that existing neural networks struggle with low-frequency crystal classifications. The present work highlights the academic significance and the engineering novelty of simulated XRD patterns in this interdisciplinary field.

cond-mat.mtrl-sci