Searcharxiv⌕ Search

arXiv subjects

Alessandro Flammini

Publications and source records attributed to Alessandro Flammini.

88 records · Page 5Linked to original sources

Characterizing and modeling the dynamics of online popularity

Online popularity has enormous impact on opinions, culture, policy, and profits. We provide a quantitative, large scale, temporal analysis of the dynamics of online content popularity in two massive model systems, the Wikipedia and an entire country's Web space. We find that the dynamics of popularity are characterized by bursts, displaying characteristic features of critical systems such as fat-tailed distributions of magnitude and inter-event time. We propose a minimal model combining the classic preferential popularity increase mechanism with the occurrence of random popularity shifts due to exogenous factors. The model recovers the critical features observed in the empirical analysis of the systems analyzed here, highlighting the key factors needed in the description of popularity dynamics.

physics.soc-ph↗

Agents, Bookmarks and Clicks: A topical model of Web traffic

Analysis of aggregate and individual Web traffic has shown that PageRank is a poor model of how people navigate the Web. Using the empirical traffic patterns generated by a thousand users, we characterize several properties of Web traffic that cannot be reproduced by Markovian models. We examine both aggregate statistics capturing collective behavior, such as page and link traffic, and individual statistics, such as entropy and session size. No model currently explains all of these empirical observations simultaneously. We show that all of these traffic patterns can be explained by an agent-based model that takes into account several realistic browsing behaviors. First, agents maintain individual lists of bookmarks (a non-Markovian memory mechanism) that are used as teleportation targets. Second, agents can retreat along visited links, a branching mechanism that also allows us to reproduce behaviors such as the use of a back button and tabbed browsing. Finally, agents are sustained by visiting novel pages of topical interest, with adjacent pages being more topically related to each other than distant ones. This modulates the probability that an agent continues to browse or starts a new session, allowing us to recreate heterogeneous session lengths. The resulting model is capable of reproducing the collective and individual behaviors we observe in the empirical data, reconciling the narrowly focused browsing patterns of individual users with the extreme heterogeneity of aggregate traffic measurements. This result allows us to identify a few salient features that are necessary and sufficient to interpret the browsing patterns observed in our data. In addition to the descriptive and explanatory power of such a model, our results may lead the way to more sophisticated, realistic, and effective ranking and crawling algorithms.

cs.NI↗

Beyond Zipf's law: Modeling the structure of human language

Human language, the most powerful communication system in history, is closely associated with cognition. Written text is one of the fundamental manifestations of language, and the study of its universal regularities can give clues about how our brains process information and how we, as a society, organize and share it. Still, only classical patterns such as Zipf's law have been explored in depth. In contrast, other basic properties like the existence of bursts of rare words in specific documents, the topical organization of collections, or the sublinear growth of vocabulary size with the length of a document, have only been studied one by one and mainly applying heuristic methodologies rather than basic principles and general mechanisms. As a consequence, there is a lack of understanding of linguistic processes as complex emergent phenomena. Beyond Zipf's law for word frequencies, here we focus on Heaps' law, burstiness, and the topicality of document collections, which encode correlations within and across documents absent in random null models. We introduce and validate a generative model that explains the simultaneous emergence of all these patterns from simple rules. As a result, we find a connection between the bursty nature of rare words and the topical organization of texts and identify dynamic word ranking and memory across documents as key mechanisms explaining the non trivial organization of written text. Our research can have broad implications and practical applications in computer science, cognitive science, and linguistics.

cs.CL↗

Remembering what we like: Toward an agent-based model of Web traffic

Analysis of aggregate Web traffic has shown that PageRank is a poor model of how people actually navigate the Web. Using the empirical traffic patterns generated by a thousand users over the course of two months, we characterize the properties of Web traffic that cannot be reproduced by Markovian models, in which destinations are independent of past decisions. In particular, we show that the diversity of sites visited by individual users is smaller and more broadly distributed than predicted by the PageRank model; that link traffic is more broadly distributed than predicted; and that the time between consecutive visits to the same site by a user is less broadly distributed than predicted. To account for these discrepancies, we introduce a more realistic navigation model in which agents maintain individual lists of bookmarks that are used as teleportation targets. The model can also account for branching, a traffic property caused by browser features such as tabs and the back button. The model reproduces aggregate traffic patterns such as site popularity, while also generating more accurate predictions of diversity, link traffic, and return time distributions. This model for the first time allows us to capture the extreme heterogeneity of aggregate traffic measurements while explaining the more narrowly focused browsing patterns of individual users.

cs.HC↗

Co-evolution of density and topology in a simple model of city formation

We study the influence that population density and the road network have on each others' growth and evolution. We use a simple model of formation and evolution of city roads which reproduces the most important empirical features of street networks in cities. Within this framework, we explicitely introduce the topology of the road network and analyze how it evolves and interact with the evolution of population density. We show that accessibility issues -pushing individuals to get closer to high centrality nodes- lead to high density regions and the appearance of densely populated centers. In particular, this model reproduces the empirical fact that the density profile decreases exponentially from a core district. In this simplified model, the size of the core district depends on the relative importance of transportation and rent costs.

physics.soc-ph↗

Modeling urban street patterns

Urban streets patterns form planar networks whose empirical properties cannot be accounted for by simple models such as regular grids or Voronoi tesselations. Striking statistical regularities across different cities have been recently empirically found, suggesting that a general and details-independent mechanism may be in action. We propose a simple model based on a local optimization process combined with ideas previously proposed in studies of leaf pattern formation. The statistical properties of this model are in good agreement with the observed empirical patterns. Our results thus suggests that in the absence of a global design strategy, the evolution of many different transportation networks indeed follow a simple universal mechanism.

physics.soc-ph↗

Random Walks on Directed Networks: the Case of PageRank

PageRank, the prestige measure for Web pages used by Google, is the stationary probability of a peculiar random walk on directed graphs, which interpolates between a pure random walk and a process where all nodes have the same probability of being visited. We give some exact results on the distribution of PageRank in the cases in which the damping factor q approaches the two limit values 0 and 1. When q -> 0 and for several classes of graphs the distribution is a power law with exponent 2, regardless of the in-degree distribution. When q -> 1 it can always be derived from the in-degree distribution of the underlying graph, if the out-degree is the same for all nodes.

physics.soc-ph↗

The egalitarian effect of search engines

Search engines have become key media for our scientific, economic, and social activities by enabling people to access information on the Web in spite of its size and complexity. On the down side, search engines bias the traffic of users according to their page-ranking strategies, and some have argued that they create a vicious cycle that amplifies the dominance of established and already popular sites. We show that, contrary to these prior claims and our own intuition, the use of search engines actually has an egalitarian effect. We reconcile theoretical arguments with empirical evidence showing that the combination of retrieval by search engines and search behavior by users mitigates the attraction of popular pages, directing more traffic toward less popular sites, even in comparison to what would be expected from users randomly surfing the Web.

cs.CY↗

Optimal Traffic Networks

Inspired by studies on the airports' network and the physical Internet, we propose a general model of weighted networks via an optimization principle. The topology of the optimal network turns out to be a spanning tree that minimizes a combination of topological and metric quantities. It is characterized by a strongly heterogeneous traffic, non-trivial correlations between distance and traffic and a broadly distributed centrality. A clear spatial hierarchical organization, with local hubs distributing traffic in smaller regions, emerges as a result of the optimization. Varying the parameters of the cost function, different classes of trees are recovered, including in particular the minimum spanning tree and the shortest path tree. These results suggest that a variational approach represents an alternative and possibly very meaningful path to the study of the structure of complex weighted networks.

physics.soc-ph↗

Scale-free network growth by ranking

Network growth is currently explained through mechanisms that rely on node prestige measures, such as degree or fitness. In many real networks those who create and connect nodes do not know the prestige values of existing nodes, but only their ranking by prestige. We propose a criterion of network growth that explicitly relies on the ranking of the nodes according to any prestige measure, be it topological or not. The resulting network has a scale-free degree distribution when the probability to link a target node is any power law function of its rank, even when one has only partial information of node ranks. Our criterion may explain the frequency and robustness of scale-free degree distributions in real networks, as illustrated by the special case of the Web graph.

cond-mat.dis-nn↗

Geometry of proteins: hydrogen bonding, sterics and marginally compact tubes

The functionality of proteins is related to their structure in the native state. Protein structures are made up of emergent building blocks of helices and almost planar sheets. A simple coarse-grained geometrical model of a flexible tube barely subject to compaction provides a unified framework for understanding the common character of globular proteins.We argue that a recent critique of the tube idea is not well founded.

q-bio.BM↗

Detecting rich-club ordering in complex networks

Uncovering the hidden regularities and organizational principles of networks arising in physical systems ranging from the molecular level to the scale of large communication infrastructures is the key issue for the understanding of their fabric and dynamical properties [1-5]. The ``rich-club'' phenomenon refers to the tendency of nodes with high centrality, the dominant elements of the system, to form tightly interconnected communities and it is one of the crucial properties accounting for the formation of dominant communities in both computer and social sciences [4-8]. Here we provide the analytical expression and the correct null models which allow for a quantitative discussion of the rich-club phenomenon. The presented analysis enables the measurement of the rich-club ordering and its relation with the function and dynamics of networks in examples drawn from the biological, social and technological domains.

physics.data-an↗

How to make the top ten: Approximating PageRank from in-degree

PageRank has become a key element in the success of search engines, allowing to rank the most important hits in the top screen of results. One key aspect that distinguishes PageRank from other prestige measures such as in-degree is its global nature. From the information provider perspective, this makes it difficult or impossible to predict how their pages will be ranked. Consequently a market has emerged for the optimization of search engine results. Here we study the accuracy with which PageRank can be approximated by in-degree, a local measure made freely available by search engines. Theoretical and empirical analyses lead to conclude that given the weak degree correlations in the Web link graph, the approximation can be relatively accurate, giving service and information providers an effective new marketing tool.

cs.IR↗

A Stochastic Model for the Species Abundance Problem in an Ecological Community

We propose a model based on coupled multiplicative stochastic processes to understand the dynamics of competing species in an ecosystem. This process can be conveniently described by a Fokker-Planck equation. We provide an analytical expression for the marginalized stationary distribution. Our solution is found in excellent agreement with numerical simulations and compares rather well with observational data from tropical forests.

q-bio.PE↗

Short period attractors and non-ergodic behavior in the deterministic fixed energy sandpile model

We study the asymptotic behaviour of the Bak, Tang, Wiesenfeld sandpile automata as a closed system with fixed energy. We explore the full range of energies characterizing the active phase. The model exhibits strong non-ergodic features by settling into limit-cycles whose period depends on the energy and initial conditions. The asymptotic activity $ρ_a$ (topplings density) shows, as a function of energy density $ζ$, a devil's staircase behaviour defining a symmetric energy interval-set over which also the period lengths remain constant. The properties of $ζ$-$ρ_a$ phase diagram can be traced back to the basic symmetries underlying the model's dynamics.

cond-mat.stat-mech↗

Crucial stages of protein folding through a solvable model: predicting target sites for enzyme-inhibiting drugs

An exactly solvable model based on the topology of a protein native state is applied to identify bottlenecks and key-sites for the folding of HIV-1 Protease. The predicted sites are found to correlate well with clinical data on resistance to FDA-approved drugs. It has been observed that the effects of drug therapy are to induce multiple mutations on the protease. The sites where such mutations occur correlate well with those involved in folding bottlenecks identified through the deterministic procedure proposed in this study. The high statistical significance of the observed correlations suggests that the approach may be promisingly used in conjunction with traditional techniques to identify candidate locations for drug attacks.

cond-mat.stat-mech↗