Searcharxiv⌕ Search

arXiv subjects

Tomoei Takahashi

Publications and source records attributed to Tomoei Takahashi.

4 recordsLinked to original sources

Dynamical Regimes of Discrete Diffusion Models

Diffusion models generate high-dimensional data such as images by learning a process that gradually removes noise from corrupted data. Recent studies have shown that the backward dynamics of diffusion models exhibit two characteristic transitions: the speciation transition, at which generated samples begin to capture the global structure of the training data, and the collapse transition, at which the generation dynamics starts committing to individual training samples. While these transitions have been theoretically analyzed for continuous data, the same theoretical criteria have not been applied for discrete diffusion models, which are diffusion models for discrete data. In this work, we propose a simple effective model for discrete diffusion models trained on two-class Ising variable data with a general mixture ratio and analyze its backward dynamics using methods from statistical mechanics. We show that, as in the previous study on continuous data, the speciation transition can be determined through a second-order phase transition analysis using high-temperature expansion, while the collapse transition corresponds to a condensation transition described by the Random Energy Model. An analytical expression for the speciation time is obtained, and we show that its scaling becomes consistent with that of the continuous case when the noise increases with time as in practical diffusion models. These theoretical predictions are confirmed by numerical simulations and experiments with trained discrete diffusion models on real datasets. These results suggest that the original theoretical framework for continuous data remain valid for discrete data, and may provide a useful starting point for the statistical-mechanics analysis of discrete generative diffusion in more realistic settings.

cond-mat.stat-mech↗

Alpha helices are more evolutionarily robust to environmental perturbations than beta sheets: Bayesian learning and statistical mechanics for protein evolution

How typical elements that shape organisms, such as protein secondary structures, have evolved, or how evolutionarily susceptible/resistant they are to environmental changes, are significant issues in evolutionary biology, structural biology, and biophysics. According to Darwinian evolution, natural selection and genetic mutations are the primary drivers of biological evolution. However, the concept of ``robustness of the phenotype to environmental perturbations across successive generations," which seems crucial from the perspective of natural selection, has not been formalized or analyzed. In this study, through Bayesian learning and statistical mechanics we formalize the stability of the free energy in the space of amino acid sequences that can design particular protein structure against perturbations of the chemical potential of water surrounding a protein as such robustness. This evolutionary stability is defined as a decreasing function of a quantity analogous to the susceptibility in the statistical mechanics of magnetic bodies specific to the amino acid sequence of a protein. Consequently, in a two-dimensional square lattice protein model composed of 36 residues, we found that as we increase the stability of the free energy against perturbations in environmental conditions, the structural space shows a steep step-like reduction. Furthermore, lattice protein structures with higher stability against perturbations in environmental conditions tend to have a higher proportion of $α$-helices and a lower proportion of $β$-sheets. This result is qualitatively confirmed by comparing the histograms of the percentage of secondary structures of evolutionarily robust proteins and randomly selected proteins through an empirical validation using a protein database.

physics.bio-ph↗

The cavity method to protein design problem

In this study, we propose an analytic statistical mechanics approach to solve a fundamental problem in biological physics called protein design. Protein design is an inverse problem of protein structure prediction, and its solution is the amino acid sequence that best stabilizes a given conformation. Despite recent rapid progress in protein design using deep learning, the challenge of exploring protein design principles remains. Contrary to previous computational physics studies, we used the cavity method, an extension of the mean-field approximation that becomes rigorous when the interaction network is a tree. We found that for small two-dimensional (2D) lattice hydrophobic-polar (HP) protein models, the design by the cavity method yields results almost equivalent to those from the Markov chain Monte Carlo method with lower computational cost.

cond-mat.stat-mech↗

Lattice protein design using Bayesian learning

Protein design is the inverse approach of the three-dimensional (3D) structure prediction for elucidating the relationship between the 3D structures and amino acid sequences. In general, the computation of the protein design involves a double loop: a loop for amino acid sequence changes and a loop for an exhaustive conformational search for each amino acid sequence. Herein, we propose a novel statistical mechanical design method using Bayesian learning, which can design lattice proteins without the exhaustive conformational search. We consider a thermodynamic hypothesis of the evolution of proteins and apply it to the prior distribution of amino acid sequences. Furthermore, we take the water effect into account in view of the grand canonical picture. As a result, on applying the 2D lattice hydrophobic-polar (HP) model, our design method successfully finds an amino acid sequence for which the target conformation has a unique ground state. However, the performance was not as good for the 3D lattice HP models compared to the 2D models. The performance of the 3D model improves on using a 20-letter lattice proteins. Furthermore, we find a strong linearity between the chemical potential of water and the number of surface residues, thereby revealing the relationship between protein structure and the effect of water molecules. The advantage of our method is that it greatly reduces computation time, because it does not require long calculations for the partition function corresponding to an exhaustive conformational search. As our method uses a general form of Bayesian learning and statistical mechanics and is not limited to lattice proteins, the results presented here elucidate some heuristics used successfully in previous protein design methods.

physics.bio-ph↗