SearcharxivSearch

arXiv subjects

Frank Potthast

Publications and source records attributed to Frank Potthast.

9 recordsLinked to original sources

Design of Sequences with Good Folding Properties in Coarse-Grained Protein Models

Background: Designing amino acid sequences that are stable in a given target structure amounts to maximizing a conditional probability. A straightforward approach to accomplish this is a nested Monte Carlo where the conformation space is explored over and over again for different fixed sequences, which requires excessive computational demand. Several approximate attempts to remedy this situation, based on energy minimization for fixed structure or high-$T$ expansions, have been proposed. These methods are fast but often not accurate since folding occurs at low $T$. Results: We develop a multisequence Monte Carlo procedure, where both sequence and conformation space are simultaneously probed with efficient prescriptions for pruning sequence space. The method is explored on hydrophobic/polar models. We first discuss short lattice chains, in order to compare with exact data and with other methods. The method is then successfully applied to lattice chains with up to 50 monomers, and to off-lattice 20-mers. Conclusions: The multisequence Monte Carlo method offers a new approach to sequence design in coarse-grained models. It is much more efficient than previous Monte Carlo methods, and is, as it stands, applicable to a fairly wide range of two-letter models.

cond-mat.soft

Monte Carlo Procedure for Protein Design

A new method for sequence optimization in protein models is presented. The approach, which has inherited its basic philosophy from recent work by Deutsch and Kurosky [Phys. Rev. Lett. 76, 323 (1996)] by maximizing conditional probabilities rather than minimizing energy functions, is based upon a novel and very efficient multisequence Monte Carlo scheme. By construction, the method ensures that the designed sequences represent good folders thermodynamically. A bootstrap procedure for the sequence space search is devised making very large chains feasible. The algorithm is successfully explored on the two-dimensional HP model with chain lengths N=16, 18 and 32.

cond-mat.soft

A Minimal Off-Lattice Model for Alpha-helical Proteins

A minimal off-lattice model for alpha-helical proteins is presented. It is based on hydrophobicity forces and sequence independent local interactions. The latter are chosen so as to favor the formation of alpha-helical structure. They model chirality and alpha-helical hydrogen bonding. The global structures resulting from the competition between these forces are studied by means of an efficient Monte Carlo method. The model is tested on two sequences of length N=21 and 33 which are intended to form 2- and 3-helix bundles, respectively. The local structure of our model proteins is compared to that of real alpha-helical proteins, and is found to be very similar. The two sequences display the desired numbers of helices in the folded phase. Only a few different relative orientations of the helices are thermodynamically allowed. Our ability to investigate the thermodynamics relies heavily upon the efficiency of the used algorithm, simulated tempering; in this Monte Carlo approach, the temperature becomes a fluctuating variable, enabling the crossing of free-energy barriers.

cond-mat.stat-mech

Binary Assignments of Amino Acids from Pattern Conservation

We develop a simple optimization procedure for assigning binary values to the amino acids. The binary values are determined by a maximization of the degree of pattern conservation in groups of closely related protein sequences. The maximization is carried out at fixed composition. For compositions approximately corresponding to an equipartition of the residues, the optimal encoding is found to be strongly correlated with hydrophobicity. The stability of the procedure is demonstrated. Our calculations are based upon sequences in the SWISS-PROT database.

chem-ph

Identification of Amino Acid Sequences with Good Folding Properties in an Off-Lattice Model

Folding properties of a two-dimensional toy protein model containing only two amino-acid types, hydrophobic and hydrophilic, respectively, are analyzed. An efficient Monte Carlo procedure is employed to ensure that the ground states are found. The thermodynamic properties are found to be strongly sequence dependent in contrast to the kinetic ones. Hence, criteria for good folders are defined entirely in terms of thermodynamic fluctuations. With these criteria sequence patterns that fold well are isolated. For 300 chains with 20 randomly chosen binary residues approximately 10% meet these criteria. Also, an analysis is performed by means of statistical and artificial neural network methods from which it is concluded that the folding properties can be predicted to a certain degree given the binary numbers characterizing the sequences.

chem-ph

Studies of an Off-Lattice Model for Protein Folding: Sequence Dependence and Improved Sampling at Finite Temperature

We study the thermodynamic behavior of a simple off-lattice model for protein folding. The model is two-dimensional and has two different ``amino acids''. Using numerical simulations of all chains containing eight or ten monomers, we examine the sequence dependence at a fixed temperature. It is shown that only a few of the chains exist in unique folded state at this temperature, and the energy level spectra of chains with different types of behavior are compared. Furthermore, we use this model as a testbed for two improved Monte Carlo algorithms. Both algorithms are based on letting some parameter of the model become a dynamical variable; one of the algorithms uses a fluctuating temperature and the other a fluctuating monomer sequence. We find that by these algorithms one gains large factors in efficiency in comparison with conventional methods.

chem-ph

Evidence for Non-Random Hydrophobicity Structures in Protein Chains

The question of whether proteins originate from random sequences of amino acids is addressed. A statistical analysis is performed in terms of blocked and random walk values formed by binary hydrophobic assignments of the amino acids along the protein chains. Theoretical expectations of these variables from random distributions of hydrophobicities are compared with those obtained from functional proteins. The results, which are based upon proteins in the SWISS-PROT data base, convincingly show that the amino acid sequences in proteins differ from what is expected from random sequences in a statistical significant way. By performing Fourier transforms on the random walks one obtains additional evidence for non-randomness of the distributions. We have also analyzed results from a synthetic model containing only two amino-acid types, hydrophobic and hydrophilic. With reasonable criteria on good folding properties in terms of thermodynamical and kinetic behavior, sequences that fold well are isolated. Performing the same statistical analysis on the sequences that fold well indicates similar deviations from randomness as for the functional proteins. The deviations from randomness can be interpreted as originating from anticorrelations in terms of an Ising spin model for the hydrophobicities. Our results, which differ from previous investigations using other methods, might have impact on how permissive with respect to sequence specificity the protein folding process is -- only sequences with non-random hydrophobicity distributions fold well. Other distributions give rise to energy landscapes with poor folding properties and hence did not survive the evolution.

chem-ph

Local Interactions and Protein Folding: A 3D Off-Lattice Approach

The thermodynamic behavior of a three-dimensional off-lattice model for protein folding is probed. The model has only two types of residues, hydrophobic and hydrophilic. In absence of local interactions, native structure formation does not occur for the temperatures considered. By including sequence independent local interactions, which qualitatively reproduce local properties of functional proteins, the dominance of a native state for many sequences is observed. As in lattice model approaches, folding takes place by gradual compactification, followed by a sequence dependent folding transition. Our results differ from lattice approaches in that bimodal energy distributions are not observed and that high folding temperatures are accompanied by relatively low temperatures for the peak of the specific heat. Also, in contrast to earlier studies using lattice models, our results convincingly demonstrate that one does not need more than two types of residues to generate sequences with good thermodynamic folding properties in three dimensions.

physics.chem-ph