SearcharxivSearch

arXiv subjects

Koon-Kiu Yan

Publications and source records attributed to Koon-Kiu Yan.

6 recordsLinked to original sources

Fluctuations in Mass-Action Equilibrium of Protein Binding Networks

We consider two types of fluctuations in the mass-action equilibrium in protein binding networks. The first type is driven by relatively slow changes in total concentrations (copy numbers) of interacting proteins. The second type, to which we refer to as spontaneous, is caused by quickly decaying thermodynamic deviations away from the equilibrium of the system. As such they are amenable to methods of equilibrium statistical mechanics used in our study. We investigate the effects of network connectivity on these fluctuations and compare them to their upper and lower bounds. The collective effects are shown to sometimes lead to large power-law distributed amplification of spontaneous fluctuations as compared to the expectation for isolated dimers. As a consequence of this, the strength of both types of fluctuations is positively correlated with the overall network connectivity of proteins forming the complex. On the other hand, the relative amplitude of fluctuations is negatively correlated with the abundance of the complex. Our general findings are illustrated using a real network of protein-protein interactions in baker's yeast with experimentally determined protein concentrations.

q-bio.MN

Prediction and verification of indirect interactions in densely interconnected regulatory networks

We develop a matrix-based approach to predict and verify indirect interactions in gene and protein regulatory networks. It is based on the approximate transitivity of indirect regulations (e.g. A regulates B and B regulates C often implies that A regulates C) and optimally takes into account the length of a cascade and signs of intermediate interactions. Our method is at its most powerful when applied to large and densely interconnected networks. It successfully predicts both the yet unknown indirect regulations, as well as the sign (activation or repression) of already known ones. The reliability of sign predictions was calibrated using the gold-standard sets of positive and negative interactions. We fine-tuned the parameters of our algorithm by maximizing the area under the Receiver Operating Characteristic (ROC) curve. We then applied the optimized algorithm to large literature-derived networks of all direct and indirect regulatory interactions in several model organisms (Homo sapiens, Saccharomyces cerevisiae, Arabidopsis thaliana and Drosophila melanogaster).

q-bio.QM

Parameters of proteome evolution from histograms of amino-acid sequence identities of paralogous proteins

The evolution of the full repertoire of proteins encoded in a given genome is mostly driven by gene duplications, deletions, and sequence modifications of existing proteins. Indirect information about relative rates and other intrinsic parameters of these three basic processes is contained in the proteome-wide distribution of sequence identities of pairs of paralogous proteins. We introduce a simple mathematical framework based on a stochastic birth-and-death model that allows one to extract some of this information and apply it to the set of all pairs of paralogous proteins in seven model organisms. It was found that the histogram of sequence identities p generated by an all-to-all alignment of all protein sequences encoded in a genome is well fitted with a power-law form ~p^(-gamma) with the value of the exponent gamma around 4 for the majority of organisms used in this study. This implies that the intra-protein variability of substitution rates is best described by the Gamma-distribution with the exponent alpha ~ 0.33. We separately measure the short-term (``raw'') duplication and deletion rates r*_dup, r*_del which include gene copies that will be removed soon after the duplication event and their dramatically reduced long-term counterparts r_dup, r_del. Systematic trends of each of the four duplication/deletion rates with the total number of genes in the genome were analyzed. All but the deletion rate of recent duplicates r*_del were shown to systematically increase with N_genes. Abnormally flat shapes of sequence identity histograms observed for yeast and human are consistent with lineages leading to these organisms undergoing one or more whole-genome duplications.

q-bio.GN

Ranking Scientific Publications Using a Simple Model of Network Traffic

To account for strong aging characteristics of citation networks, we modify Google's PageRank algorithm by initially distributing random surfers exponentially with age, in favor of more recent publications. The output of this algorithm, which we call CiteRank, is interpreted as approximate traffic to individual publications in a simple model of how researchers find new information. We develop an analytical understanding of traffic flow in terms of an RPA-like model and optimize parameters of our algorithm to achieve the best performance. The results are compared for two rather different citation networks: all American Physical Society publications and the set of high-energy physics theory (hep-th) preprints. Despite major differences between these two networks, we find that their optimal parameters for the CiteRank algorithm are remarkably similar.

physics.soc-ph

Optimal ranking in networks with community structure

The World-Wide Web (WWW) is characterized by a strong community structure in which groups of webpages (e.g. those devoted to a common topic or belonging to the same organization) are densely interconnected by hyperlinks. We study how such network architecture affects the average Google rank of individual communities. Using a mean-field approximation, we quantify how the average Google rank of community webpages depends on the degree to which it is isolated from the rest of the world in both incoming and outgoing directions, and $α$ -- the only intrinsic parameter of Google's PageRank algorithm. Based on this expression we introduce a concept of a web-community being decoupled or conversely coupled to the rest of the network. We proceed with empirical study of several internal web-communities within two US universities. The predictions of our mean-field treatment were qualitatively verified in those real-life networks. Furthermore, the value $α=0.15$ used by Google seems to be optimized for the degree of isolation of communities as they exist in the actual WWW.

physics.soc-ph

Effects of Community Structure on Search and Ranking in Information Network

The World-Wide Web (WWW) is characterized by a strong community structure in which communities of webpages (e.g. those sharing a common keyword) are densely interconnected by hyperlinks. We study how such network architecture affects the average Google ranking of individual webpages in the comunity. It is shown that the Google rank of community webpages could either increase or decrease with the density of inter-community links depending on the exact balance between average in- and out-degrees in the community. The magnitude of this effect is described by a simple analytical formula and subsequently verified by numerical simulations of random scale-free networks with a desired level of the community structure. A new algorithm allowing for generation of such networks is proposed and studied. The number of inter-community links in such networks is controlled by a temperature-like parameter with the strongest community structure realized in "low-temperature" networks.

cond-mat.other