SearcharxivSearch

arXiv subjects

Yoshihiro Hayashi

Publications and source records attributed to Yoshihiro Hayashi.

10 recordsLinked to original sources

Simulation-Supervised Foundation Models for Retention Time Prediction in High-Performance Liquid Chromatography beyond Experimental Data Coverage

Accurate prediction of high-performance liquid chromatography (HPLC) retention times (RTs) across diverse molecules and chromatographic methods remains challenging because experimental training data cover only a limited region of chemical and method spaces. Here, we develop FUSE-RT (Foundation model Unifying Simulation and Experimental supervision for Retention Time), a multitask foundation model that integrates RT data from 179 chromatographic methods and adapts to unseen molecules and methods using limited target-domain data. To extend transferability beyond experimental coverage, we introduce simulation-to-real (Sim2Real) transfer learning, in which molecular representations learned from large-scale computational data are transferred to experimental RT prediction. Specifically, we use PolyOmics, comprising 39 properties for approximately 21,400 molecules generated by molecular dynamics and density-functional theory calculations, as auxiliary supervision. We evaluate generalization under molecular, method, and joint molecular--method distribution shifts. Simulation-derived supervision substantially improves transfer beyond the experimental molecular domain, particularly under pronounced coverage gaps and few-shot adaptation. Moreover, RT-prediction error decreases systematically with increasing simulation-data size, following a significant power-law relationship. These results establish Sim2Real transfer as a scalable strategy for extending RT prediction beyond the finite coverage of experimental chromatographic data.

physics.chem-ph

Omics-scale polymer computational database transferable to real-world artificial intelligence applications

Developing large-scale foundational datasets is a critical milestone in advancing artificial intelligence (AI)-driven scientific innovation. However, unlike AI-mature fields such as natural language processing, materials science, particularly polymer research, has significantly lagged in developing extensive open datasets. This lag is primarily due to the high costs of polymer synthesis and property measurements, along with the vastness and complexity of the chemical space. This study presents PolyOmics, an omics-scale computational database generated through fully automated molecular dynamics simulation pipelines that provide diverse physical properties for over $10^5$ polymeric materials. The PolyOmics database is collaboratively developed by approximately 260 researchers from 48 institutions to bridge the gap between academia and industry. Machine learning models pretrained on PolyOmics can be efficiently fine-tuned for a wide range of real-world downstream tasks, even when only limited experimental data are available. Notably, the generalisation capability of these simulation-to-real transfer models improve significantly as the size of the PolyOmics database increases, exhibiting power-law scaling. The emergence of scaling laws supports the "more is better" principle, highlighting the significance of ultralarge-scale computational materials data for improving real-world prediction performance. This unprecedented omics-scale database reveals vast unexplored regions of polymer materials, providing a foundation for AI-driven polymer science.

physics.chem-ph

Machine Learning for Polymer Chemical Resistance to Organic Solvents

Predicting the chemical resistance of polymers to organic solvents is a longstanding challenge in materials science, with significant implications for sustainable materials design and industrial applications. Here, we address the need for interpretable and generalizable frameworks to understand and predict polymer chemical resistance beyond conventional solubility models. We systematically analyze a large dataset of polymer solvent combinations using a data-driven approach. Our study reveals that polymer crystallinity and density, as well as solvent polarity, are key factors governing chemical resistance, and that these trends are consistent with established theoretical models. These findings provide a foundation for rational screening and design of polymer materials with tailored chemical resistance, advancing both fundamental understanding and practical applications.

cond-mat.soft

SPACIER: On-Demand Polymer Design with Fully Automated All-Atom Classical Molecular Dynamics Integrated into Machine Learning Pipelines

Machine learning has rapidly advanced the design and discovery of new materials with targeted applications in various systems. First-principles calculations and other computer experiments have been integrated into material design pipelines to address the lack of experimental data and the limitations of interpolative machine learning predictors. However, the enormous computational costs and technical challenges of automating computer experiments for polymeric materials have limited the availability of open-source automated polymer design systems that integrate molecular simulations and machine learning. We developed SPACIER, an open-source software program that integrates RadonPy, a Python library for fully automated polymer property calculations based on all-atom classical molecular dynamics into a Bayesian optimization-based polymer design system to overcome these challenges. As a proof-of-concept study, we successfully synthesized optical polymers that surpass the Pareto boundary formed by the tradeoff between the refractive index and Abbe number.

cond-mat.mtrl-sci

Scaling Law of Sim2Real Transfer Learning in Expanding Computational Materials Databases for Real-World Predictions

To address the challenge of limited experimental materials data, extensive physical property databases are being developed based on high-throughput computational experiments, such as molecular dynamics simulations. Previous studies have shown that fine-tuning a predictor pretrained on a computational database to a real system can result in models with outstanding generalization capabilities compared to learning from scratch. This study demonstrates the scaling law of simulation-to-real (Sim2Real) transfer learning for several machine learning tasks in materials science. Case studies of three prediction tasks for polymers and inorganic materials reveal that the prediction error on real systems decreases according to a power-law as the size of the computational data increases. Observing the scaling behavior offers various insights for database development, such as determining the sample size necessary to achieve a desired performance, identifying equivalent sample sizes for physical and computational experiments, and guiding the design of data production protocols for downstream real-world tasks.

cond-mat.mtrl-sci

Advancing Extrapolative Predictions of Material Properties through Learning to Learn

Recent advancements in machine learning have showcased its potential to significantly accelerate the discovery of new materials. Central to this progress is the development of rapidly computable property predictors, enabling the identification of novel materials with desired properties from vast material spaces. However, the limited availability of data resources poses a significant challenge in data-driven materials research, particularly hindering the exploration of innovative materials beyond the boundaries of existing data. While machine learning predictors are inherently interpolative, establishing a general methodology to create an extrapolative predictor remains a fundamental challenge, limiting the search for innovative materials beyond existing data boundaries. In this study, we leverage an attention-based architecture of neural networks and meta-learning algorithms to acquire extrapolative generalization capability. The meta-learners, experienced repeatedly with arbitrarily generated extrapolative tasks, can acquire outstanding generalization capability in unexplored material spaces. Through the tasks of predicting the physical properties of polymeric materials and hybrid organic--inorganic perovskites, we highlight the potential of such extrapolatively trained models, particularly with their ability to rapidly adapt to unseen material domains in transfer learning scenarios.

cond-mat.mtrl-sci

Transfer learning with affine model transformation

Supervised transfer learning has received considerable attention due to its potential to boost the predictive power of machine learning in scenarios where data are scarce. Generally, a given set of source models and a dataset from a target domain are used to adapt the pre-trained models to a target domain by statistically learning domain shift and domain-specific factors. While such procedurally and intuitively plausible methods have achieved great success in a wide range of real-world applications, the lack of a theoretical basis hinders further methodological development. This paper presents a general class of transfer learning regression called affine model transfer, following the principle of expected-square loss minimization. It is shown that the affine model transfer broadly encompasses various existing methods, including the most common procedure based on neural feature extractors. Furthermore, the current paper clarifies theoretical properties of the affine model transfer such as generalization error and excess risk. Through several case studies, we demonstrate the practical benefits of modeling and estimating inter-domain commonality and domain-specific factors separately with the affine-type transfer models.

stat.ML

RadonPy: Automated Physical Property Calculation using All-atom Classical Molecular Dynamics Simulations for Polymer Informatics

The rapid growth of data-driven materials research has made it necessary to develop systematically designed, open databases of material properties. However, there are few open databases for polymeric materials compared to other material systems such as inorganic crystals. To this end, we developed RadonPy, the world-first open-source Python library for fully automated all-atom classical molecular dynamics (MD) simulations. For a given polymer repeating unit, the entire process of molecular modeling, equilibrium and nonequilibrium MD calculations, and property calculations can be conducted fully automatically. In this study, 15 different properties, including the thermal conductivity, density, specific heat capacity, thermal expansion coefficients, and refractive index, were calculated for more than 1,000 unique amorphous polymers. The calculated properties were compared and validated systematically with experimental values from PoLyInfo. During the high-throughput data production, eight amorphous polymers with extremely high thermal conductivities, exceeding 0.4 W/mK, were identified, including six polymers with unreported thermal conductivities. These polymers were found to have a high density of hydrogen bonding units or rigid backbones. A decomposition analysis of the heat conduction, which is implemented in RadonPy, revealed the underlying mechanisms that yield a high thermal conductivity of the amorphous polymers: heat transfer via hydrogen bonds and dipole-dipole interactions between the polymer chains with their hydrogen bonding units or via the covalent bonds of the polymer backbone with high rigidity. The creation of massive amounts of computational property data using RadonPy will facilitate the development of polymer informatics, similar to how the emergence of the first-principles computational database for inorganic crystals had significantly advanced materials informatics.

cond-mat.mtrl-sci

Machine Learning-Assisted Exploration of Thermally Conductive Polymers Based on High-Throughput Molecular Dynamics Simulations

Finding amorphous polymers with higher thermal conductivity is important, as they are ubiquitous in heat transfer applications. With recent progress in material informatics, machine learning approaches have been increasingly adopted for finding or designing materials with desired properties. However, relatively limited effort has been put into finding thermally conductive polymers using machine learning, mainly due to the lack of polymer thermal conductivity databases with reasonable data volume. In this work, we combine high-throughput molecular dynamics (MD) simulations and machine learning to explore polymers with relatively high thermal conductivity (> 0.300 W/m-K). We first randomly select 365 polymers from the existing PolyInfo database and calculate their thermal conductivity using MD simulations. The data are then employed to train a machine learning regression model to quantify the structure-thermal conductivity relation, which is further leveraged to screen polymer candidates in the PolyInfo database with thermal conductivity > 0.300 W/m-K. 133 polymers with MD-calculated thermal conductivity above this threshold are eventually identified. Polymers with a wide range of thermal conductivity values are selected for re-calculation under different simulation conditions, and those polymers found with thermal conductivity above 0.300 W/m-K are mostly calculated to maintain values above this threshold despite fluctuation in the exact values. A classification model is also constructed, and similar results were obtained compared to the regression model in predicting polymers with thermal conductivity above or below 0.300 W/m-K. The strategy and results from this work may contribute to automating the design of polymers with high thermal conductivity.

cond-mat.mtrl-sci

Potentials and challenges of polymer informatics: exploiting machine learning for polymer design

There has been rapidly growing demand of polymeric materials coming from different aspects of modern life because of the highly diverse physical and chemical properties of polymers. Polymer informatics is an interdisciplinary research field of polymer science, computer science, information science and machine learning that serves as a platform to exploit existing polymer data for efficient design of functional polymers. Despite many potential benefits of employing a data-driven approach to polymer design, there has been notable challenges of the development of polymer informatics attributed to the complex hierarchical structures of polymers, such as the lack of open databases and unified structural representation. In this study, we review and discuss the applications of machine learning on different aspects of the polymer design process through four perspectives: polymer databases, representation (descriptor) of polymers, predictive models for polymer properties, and polymer design strategy. We hope that this paper can serve as an entry point for researchers interested in the field of polymer informatics.

cond-mat.soft