SearcharxivSearch

arXiv subjects

Akira R. Kinjo

Publications and source records attributed to Akira R. Kinjo.

At least 19 recordsLinked to original sources

Hybrid Fuzzy-Crisp Clustering Algorithm: Theory and Experiments

With the membership function being strictly positive, the conventional fuzzy c-means clustering method sometimes causes imbalanced influence when clusters of vastly different sizes exist. That is, an outstandingly large cluster drags to its center all the other clusters, however far they are separated. To solve this problem, we propose a hybrid fuzzy-crisp clustering algorithm based on a target function combining linear and quadratic terms of the membership function. In this algorithm, the membership of a data point to a cluster is automatically set to exactly zero if the data point is ``sufficiently'' far from the cluster center. In this paper, we present a new algorithm for hybrid fuzzy-crisp clustering along with its geometric interpretation. The algorithm is tested on twenty simulated data generated and five real-world datasets from the UCI repository and compared with conventional fuzzy and crisp clustering methods. The proposed algorithm is demonstrated to outperform the conventional methods on imbalanced datasets and can be competitive on more balanced datasets.

stat.ML

Computational Modelling of Plasticity-Led Evolution

Plasticity-led evolution is a form of evolution where a change in the environment induces novel traits via phenotypic plasticity, after which the novel traits are genetically accommodated over generations under the novel environment. This mode of evolution is expected to resolve the problem of gradualism (i.e., evolution by the slow accumulation of mutations that induce phenotypic variation) implied by the Modern Evolutionary Synthesis, in the face of a large environmental change. While experimental works are essential for validating that plasticity-led evolution indeed happened, we need computational models to gain insight into its underlying mechanisms and make qualitative predictions. Such computational models should include the developmental process and gene-environment interactions in addition to genetics and natural selection. We point out that gene regulatory network models can incorporate all the above notions. In this review, we highlight results from computational modelling of gene regulatory networks that consolidate the criteria of plasticity-led evolution. Since gene regulatory networks are mathematically equivalent to artificial recurrent neural networks, we also discuss their analogies and discrepancies, which may help further understand the mechanisms underlying plasticity-led evolution.

q-bio.PE

Cooperative "folding transition" in the sequence space facilitates function-driven evolution of protein families

In the protein sequence space, natural proteins form clusters of families which are characterized by their unique native folds whereas the great majority of random polypeptides are neither clustered nor foldable to unique structures. Since a given polypeptide can be either foldable or unfoldable, a kind of "folding transition" is expected at the boundary of a protein family in the sequence space. By Monte Carlo simulations of a statistical mechanical model of protein sequence alignment that coherently incorporates both short-range and long-range interactions as well as variable-length insertions to reproduce the statistics of the multiple sequence alignment of a given protein family, we demonstrate the existence of such transition between natural-like sequences and random sequences in the sequence subspaces for 15 domain families of various folds. The transition was found to be highly cooperative and two-state-like. Furthermore, enforcing or suppressing consensus residues on a few of the well-conserved sites enhanced or diminished, respectively, the natural-like pattern formation over the entire sequence. In most families, the key sites included ligand binding sites. These results suggest some selective pressure on the key residues, such as ligand binding activity, may cooperatively facilitate the emergence of a protein family during evolution. From a more practical aspect, the present results highlight an essential role of long-range effects in precisely defining protein families, which are absent in conventional sequence models.

q-bio.BM

Monte Carlo simulation of a statistical mechanical model of multiple protein sequence alignment

A grand canonical Monte Carlo (MC) algorithm is presented for studying the lattice gas model (LGM) of multiple protein sequence alignment, which coherently combines long-range interactions and variable-length insertions. MC simulations are used for both parameter optimization of the model and production runs to explore the sequence subspace around a given protein family. In this Note, I describe the details of the MC algorithm as well as some preliminary results of MC simulations with various temperatures and chemical potentials, and compare them with the mean-field approximation. The existence of a two-state transition in the sequence space is suggested for the SH3 domain family, and inappropriateness of the mean-field approximation for the LGM is demonstrated.

q-bio.BM

Essential role of long non-coding RNAs in de novo chromatin modifications: The genomic address code hypothesis

The epigenome, i.e. the whole of chromatin modifications, is transferred from mother to daughter cells during cell differentiation. When de novo chromatin modifications (establishment or erasure of, respectively, new or pre-existing DNA methylations and/or histone modifications) are made in a daughter cell, however, it has a different epigenome than its mother cell. Although de novo chromatin modifications are an important event that comprises elementary processes of cell differentiation, its molecular mechanism remains poorly understood. We argue in this Letter that a key to solving this problem lies in understanding the role of long non-coding RNAs (lncRNAs)- a type of RNA that is becoming increasingly prominent in epigenetic studies. Many studies show that lncRNAs form ribonucleo-protein complexes in the nucleus and are involved in chromatin modifications. However, chromatin-modifying enzymes lack the information about genomic positions on which they act. It is known, on the other hand, that a single-stranded RNA in general can bind to a double-stranded DNA to form a triple helix. If each lncRNA forms a ribonucleo-protein complex with chromatin-modifying enzymes on one hand and, at the same time, a triple helix with a genomic region based on its specific nucleotide sequence on the other hand, it can induce de novo chromatin modifications at specific sites. Thus, the great variety of lncRNAs can be explained by the requirement for the diversity of "genomic address codes" specific to their cognate genomic regions where de novo chromatin modifications take place.

q-bio.GN

A unified statistical model of protein multiple sequence alignment integrating direct coupling and insertions

The multiple sequence alignment (MSA) of a protein family provides a wealth of information in terms of the conservation pattern of amino acid residues not only at each alignment site but also between distant sites. In order to statistically model the MSA incorporating both short-range and long-range correlations as well as insertions, I have derived a lattice gas model of the MSA based on the principle of maximum entropy. The partition function, obtained by the transfer matrix method with a mean-field approximation, accounts for all possible alignments with all possible sequences. The model parameters for short-range and long-range interactions were determined by a self-consistent condition and by a Gaussian approximation, respectively. Using this model with and without long-range interactions, I analyzed the globin and V-set domains by increasing the "temperature" and by "mutating" a site. The correlations between residue conservation and various measures of the system's stability indicate that the long-range interactions make the conservation pattern more specific to the structure, and increasingly stabilize better conserved residues.

q-bio.BM

Liquid-theory analogy of direct-coupling analysis of multiple-sequence alignment and its implications for protein structure prediction

The direct-coupling analysis is a powerful method for protein contact prediction, and enables us to extract "direct" correlations between distant sites that are latent in "indirect" correlations observed in a protein multiple-sequence alignment. I show that the direct correlation can be obtained by using a formulation analogous to the Ornstein-Zernike integral equation in liquid theory. This formulation intuitively illustrates how the indirect or apparent correlation arises from an infinite series of direct correlations, and provides interesting insights into protein structure prediction.

q-bio.BM

Cooperation between genetic mutations and phenotypic plasticity can bypass the Weismann barrier: The cooperative model of evolution

The Weismann barrier, or the impossibility of inheritance of acquired traits, comprises a foundation of modern biology, and it has been a major obstacle in establishing the connection between evolution and ontogenesis. We propose the cooperative model based on the assumption that evolution is achieved by a cooperation between genetic mutations and acquired changes (phenotypic plasticity). It is also assumed in this model that natural selection operates on phenotypes, rather than genotypes, of individuals, and that the relationship between phenotypes and genotypes is one-to-many. In the simulations based on these assumptions, individuals exhibited phenotypic changes in response to an environmental change, corresponding multiple genetic mutations were increasingly accumulated in individuals in the population, and phenotypic plasticity was gradually replaced with genetic mutations. This result suggests that Lamarck's law of use and disuse can effectively hold without conflicting the Weismann barrier, and thus evolution can be logically connected with ontogenesis.

q-bio.PE

Composite structural motifs of binding sites for delineating biological functions of proteins

Most biological processes are described as a series of interactions between proteins and other molecules, and interactions are in turn described in terms of atomic structures. To annotate protein functions as sets of interaction states at atomic resolution, and thereby to better understand the relation between protein interactions and biological functions, we conducted exhaustive all-against-all atomic structure comparisons of all known binding sites for ligands including small molecules, proteins and nucleic acids, and identified recurring elementary motifs. By integrating the elementary motifs associated with each subunit, we defined composite motifs which represent context-dependent combinations of elementary motifs. It is demonstrated that function similarity can be better inferred from composite motif similarity compared to the similarity of protein sequences or of individual binding sites. By integrating the composite motifs associated with each protein function, we define meta-composite motifs each of which is regarded as a time-independent diagrammatic representation of a biological process. It is shown that meta-composite motifs provide richer annotations of biological processes than sequence clusters. The present results serve as a basis for bridging atomic structures to higher-order biological phenomena by classification and integration of binding site structures.

q-bio.BM

Properties of contact matrices induced by pairwise interactions in proteins

The total conformational energy is assumed to consist of pairwise interaction energies between atoms or residues, each of which is expressed as a product of a conformation-dependent function (an element of a contact matrix, C-matrix) and a sequence-dependent energy parameter (an element of a contact energy matrix, E-matrix). Such pairwise interactions in proteins force native C-matrices to be in a relationship as if the interactions are a Go-like potential [N. Go, Annu. Rev. Biophys. Bioeng. 12. 183 (1983)] for the native C-matrix, because the lowest bound of the total energy function is equal to the total energy of the native conformation interacting in a Go-like pairwise potential. This relationship between C- and E-matrices corresponds to (a) a parallel relationship between the eigenvectors of the C- and E-matrices and a linear relationship between their eigenvalues, and (b) a parallel relationship between a contact number vector and the principal eigenvectors of the C- and E-matrices; the E-matrix is expanded in a series of eigenspaces with an additional constant term, which corresponds to a threshold of contact energy that approximately separates native contacts from non-native ones. These relationships are confirmed in 182 representatives from each family of the SCOP database by examining inner products between the principal eigenvector of the C-matrix, that of the E-matrix evaluated with a statistical contact potential, and a contact number vector. In addition, the spectral representation of C- and E-matrices reveals that pairwise residue-residue interactions, which depends only on the types of interacting amino acids but not on other residues in a protein, are insufficient and other interactions including residue connectivities and steric hindrance are needed to make native structures the unique lowest energy conformations.

q-bio.BM

Profile Conditional Random Fields for Modeling Protein Families with Structural Information

A statistical model of protein families, called profile conditional random fields (CRFs), is proposed. This model may be regarded as an integration of the profile hidden Markov model (HMM) and the Finkelstein-Reva (FR) theory of protein folding. While the model structure of the profile CRF is almost identical to the profile HMM, it can incorporate arbitrary correlations in the sequences to be aligned to the model. In addition, like in the FR theory, the profile CRF can incorporate long-range pairwise interactions between model states via mean-field-like approximations. We give the detailed formulation of the model, self-consistent approximations for treating long-range interactions, and algorithms for computing partition functions and marginal probabilities. We also outline the methods for the global optimization of model parameters as well as a Bayesian framework for parameter learning and selection of optimal alignments.

q-bio.BM

Comprehensive structural classification of ligand binding motifs in proteins

Comprehensive knowledge of protein-ligand interactions should provide a useful basis for annotating protein functions, studying protein evolution, engineering enzymatic activity, and designing drugs. To investigate the diversity and universality of ligand binding sites in protein structures, we conducted the all-against-all atomic-level structural comparison of over 180,000 ligand binding sites found in all the known structures in the Protein Data Bank by using a recently developed database search and alignment algorithm. By applying a hybrid top-down-bottom-up clustering analysis to the comparison results, we determined approximately 3000 well-defined structural motifs of ligand binding sites. Apart from a handful of exceptions, most structural motifs were found to be confined within single families or superfamilies, and to be associated with particular ligands. Furthermore, we analyzed the components of the similarity network and enumerated more than 4000 pairs of ligand binding sites that were shared across different protein folds.

q-bio.BM

On the optimal contact potential of proteins

We analytically derive the lower bound of the total conformational energy of a protein structure by assuming that the total conformational energy is well approximated by the sum of sequence-dependent pairwise contact energies. The condition for the native structure achieving the lower bound leads to the contact energy matrix that is a scalar multiple of the native contact matrix, i.e., the so-called Go potential. We also derive spectral relations between contact matrix and energy matrix, and approximations related to one-dimensional protein structures. Implications for protein structure prediction are discussed.

q-bio.BM

Nature of protein family signatures: Insights from singular value analysis of position-specific scoring matrices

Position-specific scoring matrices (PSSMs) are useful for detecting weak homology in protein sequence analysis, and they are thought to contain some essential signatures of the protein families. In order to elucidate what kind of ingredients constitute such family-specific signatures, we apply singular value decomposition to a set of PSSMs and examine the properties of dominant right and left singular vectors. The first right singular vectors were correlated with various amino acid indices including relative mutability, amino acid composition in protein interior, hydropathy, or turn propensity, depending on proteins. A significant correlation between the first left singular vector and a measure of site conservation was observed. It is shown that the contribution of the first singular component to the PSSMs act to disfavor potentially but falsely functionally important residues at conserved sites. The second right singular vectors were highly correlated with hydrophobicity scales, and the corresponding left singular vectors with contact numbers of protein structures. It is suggested that sequence alignment with a PSSM is essentially equivalent to threading supplemented with functional information. The presented method may be used to separate functionally important sites from structurally important ones, and thus it may be a useful tool for predicting protein functions.

q-bio.BM

Similarity search for local protein structures at atomic resolution by exploiting a database management system

A method to search for local structural similarities in proteins at atomic resolution is presented. It is demonstrated that a huge amount of structural data can be handled within a reasonable CPU time by using a conventional relational database management system with appropriate indexing of geometric data. This method, which we call geometric indexing, can enumerate ligand binding sites that are structurally similar to sub-structures of a query protein among more than 160,000 possible candidates within a few hours of CPU time on an ordinary desktop computer. After detecting a set of high scoring ligand binding sites by the geometric indexing search, structural alignments at atomic resolution are constructed by iteratively applying the Hungarian algorithm, and the statistical significance of the final score is estimated from an empirical model based on a gamma distribution. Applications of this method to several protein structures clearly shows that significant similarities can be detected between local structures of non-homologous as well as homologous proteins.

q-bio.BM

Wang-Landau molecular dynamics technique to search for low-energy conformational space of proteins

Multicanonical molecular dynamics (MD) is a powerful technique for sampling conformations on rugged potential surfaces such as protein. However, it is notoriously difficult to estimate the multicanonical temperature effectively. Wang and Landau developed a convenient method for estimating the density of states based on a multicanonical Monte Carlo method. In their method, the density of states is calculated autonomously during a simulation. In this paper we develop a set of techniques to effectively apply the Wang-Landau method to MD simulations. In the multicanonical MD, the estimation of the derivative of the density of states is critical. In order to estimate it accurately, we devise two original improvements. First, the correction for the density of states is made smooth by using the Gaussian distribution obtained by a short canonical simulation. Second, an approximation is applied to the derivative, which is based on the Gaussian distribution and the multiple weighted histogram technique. A test of this method was performed with small polypeptides, Met-enkephalin and Trp-cage, and it is demonstrated that Wang-Landau MD is consistent with replica exchange MD but can sample much larger conformational space.

physics.comp-ph

Recoverable One-dimensional Encoding of Three-dimensional Protein Structures

Protein one-dimensional (1D) structures such as secondary structure and contact number provide intuitive pictures to understand how the native three-dimensional (3D) structure of a protein is encoded in the amino acid sequence. However, it has not been clear whether a given set of 1D structures contains sufficient information for recovering the underlying 3D structure. Here we show that the 3D structure of a protein can be recovered from a set of three types of 1D structures, namely, secondary structure, contact number and residue-wise contact order which is introduced here for the first time. Using simulated annealing molecular dynamics simulations, the structures satisfying the given native 1D structural restraints were sought for 16 proteins of various structural classes and of sizes ranging from 56 to 146 residues. By selecting the structures best satisfying the restraints, all the proteins showed a coordinate RMS deviation of less than 4Å from the native structure, and for most of them, the deviation was even less than 2Å. The present result opens a new possibility to protein structure prediction and our understanding of the sequence-structure relationship.

q-bio.BM

CRNPRED: Highly Accurate Prediction of One-dimensional Protein Structures by Large-scale Critical Random Networks

Background: One-dimensional protein structures such as secondary structures or contact numbers are useful for three-dimensional structure prediction and helpful for intuitive understanding of the sequence-structure relationship. Accurate prediction methods will serve as a basis for these and other purposes. Results: We implemented a program CRNPRED which predicts secondary structures, contact numbers and residue-wise contact orders. This program is based on a novel machine learning scheme called critical random networks. Unlike most conventional one-dimensional structure prediction methods which are based on local windows of an amino acid sequence, CRNPRED takes into account the whole sequence. CRNPRED achieves, on average per chain, Q3 = 81% for secondary structure prediction, and correlation coefficients of 0.75 and 0.61 for contact number and residue-wise contact order predictions, respectively. Conclusion: CRNPRED will be a useful tool for computational as well as experimental biologists who need accurate one-dimensional protein structure predictions.

q-bio.BM