Searcharxiv⌕ Search

arXiv subjects

Ryo Yoshida

Publications and source records attributed to Ryo Yoshida.

35 records · Page 2Linked to original sources

Tree-Planted Transformers: Unidirectional Transformer Language Models with Implicit Syntactic Supervision

Syntactic Language Models (SLMs) can be trained efficiently to reach relatively high performance; however, they have trouble with inference efficiency due to the explicit generation of syntactic structures. In this paper, we propose a new method dubbed tree-planting: instead of explicitly generating syntactic structures, we "plant" trees into attention weights of unidirectional Transformer LMs to implicitly reflect syntactic structures of natural language. Specifically, unidirectional Transformer LMs trained with tree-planting will be called Tree-Planted Transformers (TPT), which inherit the training efficiency from SLMs without changing the inference efficiency of their underlying Transformer LMs. Targeted syntactic evaluations on the SyntaxGym benchmark demonstrated that TPTs, despite the lack of explicit generation of syntactic structures, significantly outperformed not only vanilla Transformer LMs but also various SLMs that generate hundreds of syntactic structures in parallel. This result suggests that TPTs can learn human-like syntactic knowledge as data-efficiently as SLMs while maintaining the modeling space of Transformer LMs unchanged.

cs.CL↗

Advancing Extrapolative Predictions of Material Properties through Learning to Learn

Recent advancements in machine learning have showcased its potential to significantly accelerate the discovery of new materials. Central to this progress is the development of rapidly computable property predictors, enabling the identification of novel materials with desired properties from vast material spaces. However, the limited availability of data resources poses a significant challenge in data-driven materials research, particularly hindering the exploration of innovative materials beyond the boundaries of existing data. While machine learning predictors are inherently interpolative, establishing a general methodology to create an extrapolative predictor remains a fundamental challenge, limiting the search for innovative materials beyond existing data boundaries. In this study, we leverage an attention-based architecture of neural networks and meta-learning algorithms to acquire extrapolative generalization capability. The meta-learners, experienced repeatedly with arbitrarily generated extrapolative tasks, can acquire outstanding generalization capability in unexplored material spaces. Through the tasks of predicting the physical properties of polymeric materials and hybrid organic--inorganic perovskites, we highlight the potential of such extrapolatively trained models, particularly with their ability to rapidly adapt to unseen material domains in transfer learning scenarios.

cond-mat.mtrl-sci↗

Transfer learning with affine model transformation

Supervised transfer learning has received considerable attention due to its potential to boost the predictive power of machine learning in scenarios where data are scarce. Generally, a given set of source models and a dataset from a target domain are used to adapt the pre-trained models to a target domain by statistically learning domain shift and domain-specific factors. While such procedurally and intuitively plausible methods have achieved great success in a wide range of real-world applications, the lack of a theoretical basis hinders further methodological development. This paper presents a general class of transfer learning regression called affine model transfer, following the principle of expected-square loss minimization. It is shown that the affine model transfer broadly encompasses various existing methods, including the most common procedure based on neural feature extractors. Furthermore, the current paper clarifies theoretical properties of the affine model transfer such as generalization error and excess risk. Through several case studies, we demonstrate the practical benefits of modeling and estimating inter-domain commonality and domain-specific factors separately with the affine-type transfer models.

stat.ML↗

Modeling Human Sentence Processing with Left-Corner Recurrent Neural Network Grammars

In computational linguistics, it has been shown that hierarchical structures make language models (LMs) more human-like. However, the previous literature has been agnostic about a parsing strategy of the hierarchical models. In this paper, we investigated whether hierarchical structures make LMs more human-like, and if so, which parsing strategy is most cognitively plausible. In order to address this question, we evaluated three LMs against human reading times in Japanese with head-final left-branching structures: Long Short-Term Memory (LSTM) as a sequential model and Recurrent Neural Network Grammars (RNNGs) with top-down and left-corner parsing strategies as hierarchical models. Our computational modeling demonstrated that left-corner RNNGs outperformed top-down RNNGs and LSTM, suggesting that hierarchical and left-corner architectures are more cognitively plausible than top-down or sequential architectures. In addition, the relationships between the cognitive plausibility and (i) perplexity, (ii) parsing, and (iii) beam size will also be discussed.

cs.CL↗

Composition, Attention, or Both?

In this paper, we propose a novel architecture called Composition Attention Grammars (CAGs) that recursively compose subtrees into a single vector representation with a composition function, and selectively attend to previous structural information with a self-attention mechanism. We investigate whether these components -- the composition function and the self-attention mechanism -- can both induce human-like syntactic generalization. Specifically, we train language models (LMs) with and without these two components with the model sizes carefully controlled, and evaluate their syntactic generalization performance against six test circuits on the SyntaxGym benchmark. The results demonstrated that the composition function and the self-attention mechanism both play an important role to make LMs more human-like, and closer inspection of linguistic phenomenon implied that the composition function allowed syntactic features, but not semantic features, to percolate into subtree representations.

cs.CL↗

Lower Perplexity is Not Always Human-Like

In computational psycholinguistics, various language models have been evaluated against human reading behavior (e.g., eye movement) to build human-like computational models. However, most previous efforts have focused almost exclusively on English, despite the recent trend towards linguistic universal within the general community. In order to fill the gap, this paper investigates whether the established results in computational psycholinguistics can be generalized across languages. Specifically, we re-examine an established generalization -- the lower perplexity a language model has, the more human-like the language model is -- in Japanese with typologically different structures from English. Our experiments demonstrate that this established generalization exhibits a surprising lack of universality; namely, lower perplexity is not always human-like. Moreover, this discrepancy between English and Japanese is further explored from the perspective of (non-)uniform information density. Overall, our results suggest that a cross-lingual evaluation will be necessary to construct human-like computational models.

cs.CL↗

Crystal structure prediction with machine learning-based element substitution

The prediction of energetically stable crystal structures formed by a given chemical composition is a central problem in solid-state physics. In principle, the crystalline state of assembled atoms can be determined by optimizing the energy surface, which in turn can be evaluated using first-principles calculations. However, performing the iterative gradient descent on the potential energy surface using first-principles calculations is prohibitively expensive for complex systems, such as those with many atoms per unit cell. Here, we present a unique methodology for crystal structure prediction (CSP) that relies on a machine learning algorithm called metric learning. It is shown that a binary classifier, trained on a large number of already identified crystal structures, can determine the isomorphism of crystal structures formed by two given chemical compositions with an accuracy of approximately 96.4\%. For a given query composition with an unknown crystal structure, the model is used to automatically select from a crystal structure database a set of template crystals with nearly identical stable structures to which element substitution is to be applied. Apart from the local relaxation calculation of the identified templates, the proposed method does not use ab initio calculations. The potential of this substation-based CSP is demonstrated for a wide variety of crystal systems.

cond-mat.mtrl-sci↗

RadonPy: Automated Physical Property Calculation using All-atom Classical Molecular Dynamics Simulations for Polymer Informatics

The rapid growth of data-driven materials research has made it necessary to develop systematically designed, open databases of material properties. However, there are few open databases for polymeric materials compared to other material systems such as inorganic crystals. To this end, we developed RadonPy, the world-first open-source Python library for fully automated all-atom classical molecular dynamics (MD) simulations. For a given polymer repeating unit, the entire process of molecular modeling, equilibrium and nonequilibrium MD calculations, and property calculations can be conducted fully automatically. In this study, 15 different properties, including the thermal conductivity, density, specific heat capacity, thermal expansion coefficients, and refractive index, were calculated for more than 1,000 unique amorphous polymers. The calculated properties were compared and validated systematically with experimental values from PoLyInfo. During the high-throughput data production, eight amorphous polymers with extremely high thermal conductivities, exceeding 0.4 W/mK, were identified, including six polymers with unreported thermal conductivities. These polymers were found to have a high density of hydrogen bonding units or rigid backbones. A decomposition analysis of the heat conduction, which is implemented in RadonPy, revealed the underlying mechanisms that yield a high thermal conductivity of the amorphous polymers: heat transfer via hydrogen bonds and dipole-dipole interactions between the polymer chains with their hydrogen bonding units or via the covalent bonds of the polymer backbone with high rigidity. The creation of massive amounts of computational property data using RadonPy will facilitate the development of polymer informatics, similar to how the emergence of the first-principles computational database for inorganic crystals had significantly advanced materials informatics.

cond-mat.mtrl-sci↗

Bayesian Sequential Stacking Algorithm for Concurrently Designing Molecules and Synthetic Reaction Networks

In the last few years, de novo molecular design using machine learning has made great technical progress but its practical deployment has not been as successful. This is mostly owing to the cost and technical difficulty of synthesizing such computationally designed molecules. To overcome such barriers, various methods for synthetic route design using deep neural networks have been studied intensively in recent years. However, little progress has been made in designing molecules and their synthetic routes simultaneously. Here, we formulate the problem of simultaneously designing molecules with the desired set of properties and their synthetic routes within the framework of Bayesian inference. The design variables consist of a set of reactants in a reaction network and its network topology. The design space is extremely large because it consists of all combinations of purchasable reactants, often in the order of millions or more. In addition, the designed reaction networks can adopt any topology beyond simple multistep linear reaction routes. To solve this hard combinatorial problem, we present a powerful sequential Monte Carlo algorithm that recursively designs a synthetic reaction network by sequentially building up single-step reactions. In a case study of designing drug-like molecules based on commercially available compounds, compared with heuristic combinatorial search methods, the proposed method shows overwhelming performance in terms of computational efficiency and coverage and novelty with respect to existing compounds.

q-bio.BM↗

Descriptors of intrinsic hydrodynamic thermal transport: screening a phonon database in a machine learning approach

Machine learning techniques are used to explore the intrinsic origins of the hydrodynamic thermal transport and to find new materials interesting for science and engineering. The hydrodynamic thermal transport is governed intrinsically by the hydrodynamic scale and the thermal conductivity. The correlations between these intrinsic properties and harmonic and anharmonic properties, and a large number of compositional (290) and structural (1224) descriptors of 131 crystal compound materials are obtained, revealing some of the key descriptors that determines the magnitude of the intrinsic hydrodynamic effects, most of them related with the phonon relaxation times. Then, a trained black-box model is applied to screen more than 5000 materials. The results identify materials with potential technological applications. Understanding the properties correlated to hydrodynamic thermal transport can help to find new thermoelectric materials and on the design of new materials to ease the heat dissipation in electronic devices.

cond-mat.mtrl-sci↗

Machine Learning-Assisted Exploration of Thermally Conductive Polymers Based on High-Throughput Molecular Dynamics Simulations

Finding amorphous polymers with higher thermal conductivity is important, as they are ubiquitous in heat transfer applications. With recent progress in material informatics, machine learning approaches have been increasingly adopted for finding or designing materials with desired properties. However, relatively limited effort has been put into finding thermally conductive polymers using machine learning, mainly due to the lack of polymer thermal conductivity databases with reasonable data volume. In this work, we combine high-throughput molecular dynamics (MD) simulations and machine learning to explore polymers with relatively high thermal conductivity (> 0.300 W/m-K). We first randomly select 365 polymers from the existing PolyInfo database and calculate their thermal conductivity using MD simulations. The data are then employed to train a machine learning regression model to quantify the structure-thermal conductivity relation, which is further leveraged to screen polymer candidates in the PolyInfo database with thermal conductivity > 0.300 W/m-K. 133 polymers with MD-calculated thermal conductivity above this threshold are eventually identified. Polymers with a wide range of thermal conductivity values are selected for re-calculation under different simulation conditions, and those polymers found with thermal conductivity above 0.300 W/m-K are mostly calculated to maintain values above this threshold despite fluctuation in the exact values. A classification model is also constructed, and similar results were obtained compared to the regression model in predicting polymers with thermal conductivity above or below 0.300 W/m-K. The strategy and results from this work may contribute to automating the design of polymers with high thermal conductivity.

cond-mat.mtrl-sci↗

Recreation of the Periodic Table with an Unsupervised Machine Learning Algorithm

In 1869, the first draft of the periodic table was published by Russian chemist Dmitri Mendeleev. In terms of data science, his achievement can be viewed as a successful example of feature embedding based on human cognition: chemical properties of all known elements at that time were compressed onto the two-dimensional grid system for tabular display. In this study, we seek to answer the question of whether machine learning can reproduce or recreate the periodic table by using observed physicochemical properties of the elements. To achieve this goal, we developed a periodic table generator (PTG). The PTG is an unsupervised machine learning algorithm based on the generative topographic mapping (GTM), which can automate the translation of high-dimensional data into a tabular form with varying layouts on-demand. The PTG autonomously produced various arrangements of chemical symbols, which organized a two-dimensional array such as Mendeleev's periodic table or three-dimensional spiral table according to the underlying periodicity in the given data. We further showed what the PTG learned from the element data and how the element features, such as melting point and electronegativity, are compressed to the lower-dimensional latent spaces.

stat.ML↗

A General Class of Transfer Learning Regression without Implementation Cost

We propose a novel framework that unifies and extends existing methods of transfer learning (TL) for regression. To bridge a pretrained source model to the model on a target task, we introduce a density-ratio reweighting function, which is estimated through the Bayesian framework with a specific prior distribution. By changing two intrinsic hyperparameters and the choice of the density-ratio model, the proposed method can integrate three popular methods of TL: TL based on cross-domain similarity regularization, a probabilistic TL using the density-ratio estimation, and fine-tuning of pretrained neural networks. Moreover, the proposed method can benefit from its simple implementation without any additional cost; the regression model can be fully trained using off-the-shelf libraries for supervised learning in which the original output variable is simply transformed to a new output variable. We demonstrate its simplicity, generality, and applicability using various real data applications.

stat.ML↗

Potentials and challenges of polymer informatics: exploiting machine learning for polymer design

There has been rapidly growing demand of polymeric materials coming from different aspects of modern life because of the highly diverse physical and chemical properties of polymers. Polymer informatics is an interdisciplinary research field of polymer science, computer science, information science and machine learning that serves as a platform to exploit existing polymer data for efficient design of functional polymers. Despite many potential benefits of employing a data-driven approach to polymer design, there has been notable challenges of the development of polymer informatics attributed to the complex hierarchical structures of polymers, such as the lack of open databases and unified structural representation. In this study, we review and discuss the applications of machine learning on different aspects of the polymer design process through four perspectives: polymer databases, representation (descriptor) of polymers, predictive models for polymer properties, and polymer design strategy. We hope that this paper can serve as an entry point for researchers interested in the field of polymer informatics.

cond-mat.soft↗

A Bayesian algorithm for retrosynthesis

The identification of synthetic routes that end with a desired product has been an inherently time-consuming process that is largely dependent on expert knowledge regarding a limited fraction of the entire reaction space. At present, emerging machine-learning technologies are overturning the process of retrosynthetic planning. The objective of this study is to discover synthetic routes backwardly from a given desired molecule to commercially available compounds. The problem is reduced to a combinatorial optimization task with the solution space subject to the combinatorial complexity of all possible pairs of purchasable reactants. We address this issue within the framework of Bayesian inference and computation. The workflow consists of two steps: a deep neural network is trained that forwardly predicts a product of the given reactants with a high level of accuracy, following which this forward model is inverted into the backward one via Bayes' law of conditional probability. Using the backward model, a diverse set of highly probable reaction sequences ending with a given synthetic target is exhaustively explored using a Monte Carlo search algorithm. The Bayesian retrosynthesis algorithm could successfully rediscover 80.3% and 50.0% of known synthetic routes of single-step and two-step reactions within top-10 accuracy, respectively, thereby outperforming state-of-the-art algorithms in terms of the overall accuracy. Remarkably, the Monte Carlo method, which was specifically designed for the presence of diverse multiple routes, often revealed a ranked list of hundreds of reaction routes to the same synthetic target. We investigated the potential applicability of such diverse candidates based on expert knowledge from synthetic organic chemistry.

stat.ML↗

Exploring diamond-like lattice thermal conductivity crystals via feature-based transfer learning

Ultrahigh lattice thermal conductivity materials hold great importance since they play a critical role in the thermal management of electronic and optical devices. Models using machine learning can search for materials with outstanding higher-order properties like thermal conductivity. However, the lack of sufficient data to train a model is a serious hurdle. Herein we show that big data can complement small data for accurate predictions when lower-order feature properties available in big data are selected properly and applied to transfer learning. The connection between the crystal information and thermal conductivity is directly built with a neural network by transferring descriptors acquired through a pre-trained model for the feature property. Successful transfer learning shows the ability of extrapolative prediction and reveals descriptors for lattice anharmonicity. Transfer learning is employed to screen over 60000 compounds to identify novel crystals that can serve as alternatives to diamond.

cond-mat.mtrl-sci↗

SPF-CellTracker: Tracking multiple cells with strongly-correlated moves using a spatial particle filter

Tracking many cells in time-lapse 3D image sequences is an important challenging task of bioimage informatics. Motivated by a study of brain-wide 4D imaging of neural activity in C. elegans, we present a new method of multi-cell tracking. Data types to which the method is applicable are characterized as follows: (i) cells are imaged as globular-like objects, (ii) it is difficult to distinguish cells based only on shape and size, (iii) the number of imaged cells ranges in several hundreds, (iv) moves of nearly-located cells are strongly correlated and (v) cells do not divide. We developed a tracking software suite which we call SPF-CellTracker. Incorporating dependency on cells' moves into prediction model is the key to reduce the tracking errors: cell-switching and coalescence of tracked positions. We model target cells' correlated moves as a Markov random field and we also derive a fast computation algorithm, which we call spatial particle filter. With the live-imaging data of nuclei of C. elegans neurons in which approximately 120 nuclei of neurons are imaged, we demonstrate an advantage of the proposed method over the standard particle filter and a method developed by Tokunaga et al. (2014).

cs.CV↗