SearcharxivSearch

arXiv subjects

Susan Khor

Publications and source records attributed to Susan Khor.

17 recordsLinked to original sources

Towards protein folding pathways by reconstructing protein residue networks with a policy-driven model

A method that reconstructs protein residue networks using suitable node selection and edge recovery policies produced numerical observations that correlate strongly (Pearson's correlation coefficient < -0.83) with published folding rates for 52 two-state folders and 21 multi-state folders; correlations are also strong at the fold-family level. These results were obtained serendipitously with the ND model, which was introduced previously, but is here extended with policies that dictate actions according to feature states. This result points to the importance of both the starting search point and the prevailing condition (random seed) for the quick success of policy search by a simple hill-climber. The two conditions, suitable policies and random seed, which (evidenced by the strong correlation statistic) setup a conducive environment for modelling protein folding within ND, could be compared to appropriate physiological conditions required by proteins to fold naturally. Of interest is an examination of the sequence of restored edges for potential as plausible protein folding pathways. Towards this end, trajectory data is collected for analysis and further model evaluation and development.

q-bio.BM

Forming native shortcut networks to simulate protein folding

Native shortcut networks (SCN0) are sub-graphs of native protein residue networks (PRN0). In this paper, we propose the Network Dynamics (ND) model, which reconstructs a PRN0 by adding back its edges according to some recipe, while its nodes are fixed at native (PDB) locations. A PRN0 reconstruction will eventually reconstruct its SCN0, but only after exploring several non-native shortcut network (SCN) configurations. It is these other SCN configurations that are of interest to us, as they produce the statistics to evaluate the different recipes. The recipes vary from each other slightly to investigate the effect of different edge orderings on protein folding. An edge ordering is deemed more successful if it produces a stronger correlation with experimental folding rates. Over proteins of different chain fold types, this basic requirement is best satisfied when a recipe favours earlier restoration of edges with smaller sequence separation, and earlier restoration of edges incident on nodes with larger remaining degree within the 1-hop neighborhood of a previously selected node. Further, this recipe generated more route-like trajectories over SCN space and a wider range of simulated folding rates, both of which signal cooperative two-state folding behavior. It also produced better correspondence between ND calculated phi-values and experimental phi-values obtained from a set of ten transition state ensembles (TSE). Also introduced is the local centrality measure which uses centrality of initial fold substructures to yield calculated phi-values from PRN0s with even better overall correspondence than ND calculated phi-values. Preliminary investigation with mixed ND recipes by chain fold type suggests that concocting further context-sensitive recipes by chain fold type and combining them could improve ND. This search task may be better undertaken by machine learning.

q-bio.MN

Folding with a protein's native shortcut network

A complex network approach to protein folding is proposed. The graph object is the network of shortcut edges present in a native-state protein (SCN0). Although SCN0s are found via an intuitive message passing algorithm (S. Milgram, Psychology Today, May 1967 pp. 61-67), they are meaningful enough that the logarithm form of their contact order (SCN0_lnCO) correlates significantly with protein kinetic rates, regardless of protein size. Further, the clustering coefficient of a SCN0 (CSCN0) can be used to combine protein segments iteratively within the Restricted Binary Collision model to form the whole native structure. This simple yet surprisingly effective strategy identified reasonable folding pathways for 12 small single-domain two-state folders, and three non-canonical proteins: ACBP (non-two-state), Top7 (non-cooperative) and DHFR (non-single-domain, > 100 residues). For two-state folders, CSCN0 is relatable to folding rates, transition-state placement and stability. The influence of CSCN0 on folding extends to non-native structures. Moreover, SCN analysis of non-native structures could suggest three fold success factors for the fast folding Villin headpiece peptide. These results support the view of protein folding as a bottom-up hierarchical process guided from above by native-state topology, and could facilitate future constructive demonstrations of this long held hypothesis for larger proteins.

q-bio.MN

Comparing local search paths with global search paths on protein residue networks: allosteric communication

Although proteins have been recognized as small-world networks and their small-world network properties of clustering and short paths have been exploited computationally to produce biologically relevant information, they have not been truly explored as such, i.e. as navigable small-world networks in the original spirit of Milgram's work. This research seeks to fill this gap by exploring local search on a network representation of proteins and to probe the source of navigability in proteins. Previously, we confirmed that proteins are navigable small-world networks and observed that local search paths exhibit different characteristics from global search paths. In this paper, we investigate the biological relevance of the differences in path characteristics on a type III receptor tyrosine kinase (KIT). A chief difference that works in favour of local search paths as intra-protein communication pathways is their weaker proclivity, compared to global search paths, for traversing long-range edges. Long-range edges tend to be less stable and their inclusion tends to decrease the communication propensity of a path. The source of protein navigability is traced to clustering provided by short-range edges. The majority of a protein's short-range edges reside within structures deemed important for long-range energy transport and modulation of allosteric communication in proteins. Therefore, the disruption of intra-protein communication as a result of the destruction of these structures via random rewiring is expected. A local search perspective leads us to this expected conclusion while a global search perspective does not. These findings initiate the compilation of a list of path properties that are characteristic of intra-protein pathways and could suggest fresh avenues for evolving and regulating navigable (small-world) networks.

q-bio.MN

Protein residue networks from a local search perspective

We examined protein residue networks (PRNs) from a local search perspective to understand why PRNs are highly clustered when having short paths is important for protein functionality. We found that by adopting a local search perspective, this conflict between form and function is resolved as increased clustering actually helps to reduce path length in PRNs. Further, the paths found via our EDS local search algorithm are more congruent with the characteristics of intra-protein communication. EDS identifies a subset of PRN edges called short-cuts that are distinct, have high usage, impacts EDS path length, diversity and stretch, and are dominated by short-range contacts. The short-cuts form a network (SCN) that increases in size and transitivity as a protein folds. The structure of a SCN supports its function and formation, and the function of a SCN influences its formation. Several significant differences in terms of SCN structure, function and formation is found between successful and unsuccessful MD trajectories. By connecting the static and the dynamic aspects of PRNs, the protein folding process becomes a problem of graph formation with the purpose of forming suitable pathways within proteins.

q-bio.MN

The short-cut network within protein residue networks

A protein residue network (PRN) is a network of interacting amino acids within a protein. We describe characteristics of a sparser, highly central and more volatile sub-network of a PRN called the short-cut network (SCN), as a protein folds under molecular dynamics (MD) simulation with the goal of understanding how proteins form navigable small-world networks within themselves. The edges of an SCN are found via a local greedy search on a PRN. SCNs grow in size and transitivity strength as a protein folds, and SCNs from successful MD trajectories are better formed in these terms. Findings from an investigation on how to model the formation of SCNs using dynamic graph theory, and suggestions to move forward are presented. A SCN is enriched with short-range contacts and its formation correlates positively with secondary structure formation. Thus our approach to modeling PRN formation, in essence protein folding from a graph theoretic view point, is more in tune with the notion of increasing order to a random graph than the other way around, and this increase in order coincides with improved navigability of PRNs.

q-bio.MN

Piggy-backing protein domains with Formal Concept Analysis

Identifying reliable domain-domain interactions (DDIs) will increase our ability to predict novel protein-protein interactions (PPIs), to unravel interactions in protein complexes, and thus gain more information about the function and behavior of genes. One of the challenges of identifying reliable DDIs is domain promiscuity. Promiscuous domains are domains that can occur in many domain architectures and are therefore found in many proteins. This becomes a problem for a method where the score of a domain-pair is the ratio between observed and expected frequencies because the PPI network is sparse. As such, many protein-pairs will be non-interacting and domain-pairs with promiscuous domains will be penalized. This domain promiscuity challenge to the problem of inferring reliable DDIs from PPIs has been recognized, and a number of work-arounds have been proposed. In this paper, we report an application of Formal Concept Analysis (FCA) to this problem. We find that the relationship between formal concepts provide a natural way for rare domains to elevate the rank of promiscuous domains, and enrich highly ranked domain-pairs with reliable DDIs. This piggy-backing of promiscuous domains onto rare domains is possible due to the domain architecture of proteins which mixes promiscuous with rare domains.

q-bio.QM

Local search on protein residue networks

Self-avoiding random walks were performed on protein residue networks. Compared with protein residue networks with randomized links, the probability of a walk being successful is lower and the length of successful walks shorter in (non-randomized) protein residue networks. Fewer successful walks and shorter successful walks point to higher communication specificity between protein residues, a conceivably favourable attribute for proteins to have. The use of random walks instead of shortest paths also produced lower node centrality, lower edge betweeness and lower edge load for (non-randomized) protein residue networks than in their respective randomized counterparts. The implications of these properties for protein residue networks are discussed in terms of communication congestion and network vulnerability. The randomized protein residue networks have lower network clustering than the (non-randomized) protein residue networks. Hence, our findings also shed light on a hitherto neglected aspect: the importance of high network clustering in protein residue networks. High clustering increases navigability of a network for local search and the combination of a local search process on a highly clustered small-world network topology such as protein residue networks reduces communication congestion and network vulnerability.

q-bio.MN

Optimality of Moore neighborhoods in protein contact maps

A protein contact map is a binary symmetric adjacency matrix capturing the distance relationship between atoms of a protein. Each cell (i, j) of a protein contact map states whether the atoms (nodes) i and j are within some Euclidean distance from each other. We examined the radius one Moore neighborhood surrounding each cell (i, j) where j > (i + 2) in complete protein contact maps by mutating them one at a time. We found that the particular configuration of a neighborhood is generally (97%) optimal in the sense that no other configuration could maintain or improve upon existing local and global efficiencies of the nodes residing in a neighborhood. Local efficiency of a node is directly related to its clustering measure. Global efficiency of a node is inversely related to its distance to other nodes in the network. This feature of the Moore neighborhood in complete protein contact maps may explain how protein residue networks are able to form long-range links to reduce average path length while maintaining a high level of clustering throughout the process of small-world formation, and could suggest new approaches to protein contact map prediction. Effectively, the problem of protein contact map prediction is transformed to one of maximizing the number of optimal neighborhoods. By comparison, Moore neighborhoods in protein contact maps with randomized long-range links are less optimal.

q-bio.MN

Speeding up the construction of slow adaptive walks

An algorithm (bliss) is proposed to speed up the construction of slow adaptive walks. Slow adaptive walks are adaptive walks biased towards closer points or smaller move steps. They were previously introduced to explore a search space, e.g. to detect potential local optima or to assess the ruggedness of a fitness landscape. To avoid the quadratic cost of computing Hamming distance (HD) for all-pairs of strings in a set in order to find the set of closest strings for each string, strings are sorted and clustered by bliss such that similar strings are more likely to get paired off for HD computation. To efficiently arrange the strings by similarity, bliss employs the idea of shared non-overlapping position specific subsequences between strings which is inspired by an alignment-free protein sequence comparison algorithm. Tests are performed to evaluate the quality of b-walks, i.e. slow adaptive walks constructed from the output of bliss, on enumerated search spaces. Finally, b-walks are applied to explore larger search spaces with the help of Wang-Landau sampling.

cs.NE

Search space analysis with Wang-Landau sampling and slow adaptive walks

Two complementary techniques for analyzing search spaces are proposed: (i) an algorithm to detect search points with potential to be local optima; and (ii) a slightly adjusted Wang-Landau sampling algorithm to explore larger search spaces. The detection algorithm assumes that local optima are points which are easier to reach and harder to leave by a slow adaptive walker. A slow adaptive walker moves to a nearest fitter point. Thus, points with larger outgoing step sizes relative to incoming step sizes are marked using the local optima score formulae as potential local optima points (PLOPs). Defining local optima in these more general terms allows their detection within the closure of a subset of a search space, and the sampling of a search space unshackled by a particular move set. Tests are done with NK and HIFF problems to confirm that PLOPs detected in the manner proposed retain characteristics of local optima, and that the adjusted Wang-Landau samples are more representative of the search space than samples produced by choosing points uniformly at random. While our approach shows promise, more needs to be done to reduce its computation cost that it may pave a way toward analyzing larger search spaces of practical meaning.

cs.NE

What can we learn from slow self-avoiding adaptive walks by an infinite radius search algorithm?

Slow self-avoiding adaptive walks by an infinite radius search algorithm (Limax) are analyzed as themselves, and as the network they form. The study is conducted on several NK problems and two HIFF problems. We find that examination of such "slacker" walks and networks can indicate relative search difficulty within a family of problems, help identify potential local optima, and detect presence of structure in fitness landscapes. Hierarchical walks are used to differentiate rugged landscapes which are hierarchical (e.g. HIFF) from those which are anarchic (e.g. NK). The notion of node viscidity as a measure of local optimum potential is introduced and found quite successful although more work needs to be done to improve its accuracy on problems with larger K.

cs.NE

Why aren't the small worlds of protein contact networks smaller

Computer experiments are performed to investigate why protein contact networks (networks induced by spatial contacts between amino acid residues of a protein) do not have shorter average shortest path lengths in spite of their importance to protein folding. We find that shorter average inter-nodal distances is no guarantee of finding a global optimum more easily. Results from the experiments also led to observations which parallel an existing view that neither short-range nor long-range interactions dominate the protein folding process. Nonetheless, runs where there was a slight delay in the use of long-range interactions yielded the best search performance. We incorporate this finding into the optimization function by giving more weight to short-range links. This produced results showing that randomizing long-range links does not yield better search performance than protein contact networks au natural even though randomizing long-range links significantly reduces average path lengths and retains much of the clustering and positive degree-degree correlation inherent in protein contact networks. Hence there can be explanations, other than the excluded volume argument, beneath the topological limits of protein contact networks.

q-bio.MN

Static and dynamic characteristics of protein contact networks

The principles underlying protein folding remains one of Nature's puzzles with important practical consequences for Life. An approach that has gathered momentum since the late 1990's, looks at protein hetero-polymers and their folding process through the lens of complex network analysis. Consequently, there is now a body of empirical studies describing topological characteristics of protein macro-molecules through their contact networks and linking these topological characteristics to protein folding. The present paper is primarily a review of this rich area. But it delves deeper into certain aspects by emphasizing short-range and long-range links, and suggests unconventional places where "power-laws" may be lurking within protein contact networks. Further, it considers the dynamical view of protein contact networks. This closer scrutiny of protein contact networks raises new questions for further research, and identifies new regularities which may be useful to parameterize a network approach to protein folding. Preliminary experiments with such a model confirm that the regularities we identified cannot be easily reproduced through random effects. Indeed, the grand challenge of protein folding is to elucidate the process(es) which not only generates the specific and diverse linkage patterns of protein contact networks, but also reproduces the dynamic behavior of proteins as they fold. Keywords: network analysis, protein contact networks, protein folding

cs.CE

Application of Graph Coloring to Biological Networks

We explore the application of graph coloring to biological networks, specifically protein-protein interaction (PPI) networks. First, we find that given similar conditions (i.e. number of nodes, number of links, degree distribution and clustering), fewer colors are needed to color disassortative (high degree nodes tend to connect to low degree nodes and vice versa) than assortative networks. Fewer colors create fewer independent sets which in turn imply higher concurrency potential for a network. Since PPI networks tend to be disassortative, we suggest that in addition to functional specificity and stability proposed previously by Maslov and Sneppen (Science 296, 2002), the disassortative nature of PPI networks may promote the ability of cells to perform multiple, crucial and functionally diverse tasks concurrently. Second, since graph coloring is closely related to the presence of cliques in a graph, the significance of node coloring information to the problem of identifying protein complexes, i.e. dense subgraphs in a PPI network, is investigated. We find that for PPI networks where 1% to 11% of nodes participate in at least one identified protein complex, such as H. sapien (DIP20070219, DIP20081014 and HPRD070609), DSATUR (a well-known complete graph coloring algorithm) node coloring information can improve the quality (homogeneity and separation) of initial candidate complexes. This finding may help to improve existing protein complex detection methods, and/or suggest new methods.

cs.CE

Generating Hierarchically Modular Networks via Link Switching

This paper introduces a method to generate hierarchically modular networks with prescribed node degree list by link switching. Unlike many existing network generating models, our method does not use link probabilities to achieve modularity. Instead, it utilizes a user-specified topology to determine relatedness between pairs of nodes in terms of edge distances and links are switched to increase edge distances. To measure the modular-ness of a network as a whole, a new metric called Q2 is proposed. Comparisons are made between the Q [15] and Q2 measures. We also comment on the effect of our modularization method on other network characteristics such as clustering, hierarchy, average path length, small-worldness, degree correlation and centrality. An application of this method is reported elsewhere [12]. Briefly, the generated networks are used as test problems to explore the effect of modularity and degree distribution on evolutionary search algorithms.

cs.OH

Effect of Degree Distribution on Evolutionary Search

This paper introduces a method to generate hierarchically modular networks with prescribed node degree list and proposes a metric to measure network modularity based on the notion of edge distance. The generated networks are used as test problems to explore the effect of modularity and degree distribution on evolutionary algorithm performance. Results from the experiments (i) confirm a previous finding that modularity increases the performance advantage of genetic algorithms over hill climbers, and (ii) support a new conjecture that test problems with modularized constraint networks having heavy-tailed right-skewed degree distributions are more easily solved than test problems with modularized constraint networks having bell-shaped normal degree distributions.

cs.NE