SearcharxivSearch

arXiv subjects

Amedeo Caflisch

Publications and source records attributed to Amedeo Caflisch.

5 recordsLinked to original sources

Multistep greedy algorithm identifies community structure in real-world and computer-generated networks

We have recently introduced a multistep extension of the greedy algorithm for modularity optimization. The extension is based on the idea that merging l pairs of communities (l>1) at each iteration prevents premature condensation into few large communities. Here, an empirical formula is presented for the choice of the step width l that generates partitions with (close to) optimal modularity for 17 real-world and 1100 computer-generated networks. Furthermore, an in-depth analysis of the communities of two real-world networks (the metabolic network of the bacterium E. coli and the graph of coappearing words in the titles of papers coauthored by Martin Karplus) provides evidence that the partition obtained by the multistep greedy algorithm is superior to the one generated by the original greedy algorithm not only with respect to modularity but also according to objective criteria. In other words, the multistep extension of the greedy algorithm reduces the danger of getting trapped in local optima of modularity and generates more reasonable partitions.

cs.DS

Efficient modularity optimization by multistep greedy algorithm and vertex mover refinement

Identifying strongly connected substructures in large networks provides insight into their coarse-grained organization. Several approaches based on the optimization of a quality function, e.g., the modularity, have been proposed. We present here a multistep extension of the greedy algorithm (MSG) that allows the merging of more than one pair of communities at each iteration step. The essential idea is to prevent the premature condensation into few large communities. Upon convergence of the MSG a simple refinement procedure called "vertex mover" (VM) is used for reassigning vertices to neighboring communities to improve the final modularity value. With an appropriate choice of the step width, the combined MSG-VM algorithm is able to find solutions of higher modularity than those reported previously. The multistep extension does not alter the scaling of computational cost of the greedy algorithm.

cs.DS

Local modularity measure for network clusterizations

Many complex networks have an underlying modular structure, i.e., structural subunits (communities or clusters) characterized by highly interconnected nodes. The modularity $Q$ has been introduced as a measure to assess the quality of clusterizations. $Q$ has a global view, while in many real-world networks clusters are linked mainly \emph{locally} among each other (\emph{local cluster-connectivity}). Here, we introduce a new measure, localized modularity $LQ$, which reflects local cluster structure. Optimization of $Q$ and $LQ$ on the clusterization of two biological networks shows that the localized modularity identifies more cohesive clusters, yielding a complementary view of higher granularity.

cond-mat.dis-nn

Estimation of protein folding probability from equilibrium simulations

The assumption that similar structures have similar folding probabilities ($p_{fold}$) leads naturally to a procedure to evaluate $p_{fold}$ for every snapshot saved along an equilibrium folding-unfolding trajectory of a structured peptide or protein. The procedure utilizes a structurally homogeneous clustering and does not require any additional simulation. It can be used to detect multiple folding pathways as shown for a three-stranded antiparallel $β$-sheet peptide investigated by implicit solvent molecular dynamics simulations.

q-bio.BM

The protein folding network

The conformation space of a 20-residue antiparallel $β$-sheet peptide, sampled by molecular dynamics simulations, is mapped to a network. Conformations are nodes of the network, and the transitions between them are links. The conformation space network describes the significant free energy minima and their dynamic connectivity without projections into arbitrarily chosen reaction coordinates. As previously found for the Internet and the World-Wide Web as well as for social and biological networks, the conformation space network is scale-free and contains highly connected hubs like the native state which is the most populated free energy basin. Furthermore, the native basin exhibits a hierarchical organization which is not found for a random heteropolymer lacking a predominant free-energy minimum. The network topology is used to identify conformations in the folding transition state ensemble, and provides a basis for understanding the heterogeneity of the transition state and denaturated state ensemble as well as the existence of multiple pathway

q-bio.BM