SearcharxivSearch

arXiv subjects

Nima Sarshar

Publications and source records attributed to Nima Sarshar.

12 recordsLinked to original sources

Experience versus Talent Shapes the Structure of the Web

We use sequential large-scale crawl data to empirically investigate and validate the dynamics that underlie the evolution of the structure of the web. We find that the overall structure of the web is defined by an intricate interplay between experience or entitlement of the pages (as measured by the number of inbound hyperlinks a page already has), inherent talent or fitness of the pages (as measured by the likelihood that someone visiting the page would give a hyperlink to it), and the continual high rates of birth and death of pages on the web. We find that the web is conservative in judging talent and the overall fitness distribution is exponential, showing low variability. The small variance in talent, however, is enough to lead to experience distributions with high variance: The preferential attachment mechanism amplifies these small biases and leads to heavy-tailed power-law (PL) inbound degree distributions over all pages, as well as over pages that are of the same age. The balancing act between experience and talent on the web allows newly introduced pages with novel and interesting content to grow quickly and surpass older pages. In this regard, it is much like what we observe in high-mobility and meritocratic societies: People with entitlement continue to have access to the best resources, but there is just enough screening for fitness that allows for talented winners to emerge and join the ranks of the leaders. Finally, we show that the fitness estimates have potential practical applications in ranking query results.

cs.CY

Comparison of Image Similarity Queries in P2P Systems

Given some of the recent advances in Distributed Hash Table (DHT) based Peer-To-Peer (P2P) systems we ask the following questions: Are there applications where unstructured queries are still necessary (i.e., the underlying queries do not efficiently map onto any structured framework), and are there unstructured P2P systems that can deliver the high bandwidth and computing performance necessary to support such applications. Toward this end, we consider an image search application which supports queries based on image similarity metrics, such as color histogram intersection, and discuss why in this setting, standard DHT approaches are not directly applicable. We then study the feasibility of implementing such an image search system on two different unstructured P2P systems: power-law topology with percolation search, and an optimized super-node topology using structured broadcasts. We examine the average and maximum values for node bandwidth, storage and processing requirements in the percolation and super-node models, and show that current high-end computers and high-speed links have sufficient resources to enable deployments of large-scale complex image search systems.

cs.DC

Low Latency Wireless Ad-Hoc Networking: Power and Bandwidth Challenges and a Hierarchical Solution

This paper is concerned with the scaling of the number of hops in a large scale wireless ad-hoc network (WANET), a quantity we call network latency. A large network latency affects all aspects of data communication in a WANET, including an increase in delay, packet loss, required processing power and memory. We consider network management and data routing challenges in WANETs with scalable network latency. On the physical side, reducing network latency imposes a significantly higher power and bandwidth demand on nodes, as is reflected in a set of new bounds. On the protocol front, designing distributed routing protocols that can guarantee the delivery of data packets within scalable number of hops is a challenging task. To solve this, we introduce multi-resolution randomized hierarchy (MRRH), a novel power and bandwidth efficient WANET protocol with scalable network latency. MRRH uses a randomized algorithm for building and maintaining a random hierarchical network topology, which together with the proposed routing algorithm can guarantee efficient delivery of data packets in the wireless network. For a network of size $N$, MRRH can provide an average latency of only $O(\log^{3} N)$. The power and bandwidth consumption of MRRH are shown to be \emph{nearly} optimal for the latency it provides. Therefore, MRRH, is a provably efficient candidate for truly large scale wireless ad-hoc networking.

cs.IT

Finite Percolation at a Multiple of the Threshold

Bond percolation on infinite heavy-tailed power-law random networks lacks a proper phase transition; or one may say, there is a phase transition at {\em zero percolation probability}. Nevertheless, a finite size percolation threshold $q_c(N)$, where $N$ is the network size, can be defined. For such heavy-tailed networks, one can choose a percolation probability $q(N)=ρq_c(N)$ such that $\displaystyle \lim_{N\to \infty}(q-q_c(N)) =0$, and yet $ρ$ is arbitrarily large (such a scenario does not exist for networks with non-zero percolation threshold). We find that the critical behavior of random power-law networks is best described in terms of $ρ$ as the order parameter, rather than $q$. This paper makes the notion of the phase transition of the size of the largest connected component at $ρ=1$ precise. In particular, using a generating function based approach, we show that for $ρ>1$, and the power-law exponent, $2\leq τ<3$, the largest connected component scales as $\sim N^{1-1/τ}$, while for $0<ρ<1$ the scaling is at most $\sim N^{\frac{2-τ}τ}$; here, the maximum degree of any node, $k_{max}$, has been assumed to scale as N^{1/τ}$. In general, our approach yields that for large $N$, $ρ\gg 1$, $2\leq τ<3$, and $k_{max} \sim N^{1/τ}$, the largest connected component scales as $\sim ρ^{1/(3-τ)}N^{1-1/τ}$.Thus, for any fixed but large N, we recover, and make it precise, a recent result that computed a scaling behavior of $q^{1/(3-τ)}$ for "small $q$". We also provide large-scale simulation results validating some of these scaling predictions, and discuss applications of these scaling results to supporting efficient unstructured queries in peer-to-peer networks.

cond-mat.dis-nn

A Practical Approach to Joint Network-Source Coding

We are interested in how to best communicate a real valued source to a number of destinations (sinks) over a network with capacity constraints in a collective fidelity metric over all the sinks, a problem which we call joint network-source coding. It is demonstrated that multiple description codes along with proper diversity routing provide a powerful solution to joint network-source coding. A systematic optimization approach is proposed. It consists of optimizing the network routing given a multiple description code and designing optimal multiple description code for the corresponding optimized routes.

cs.IT

Joint Network-Source Coding: An Achievable Region with Diversity Routing

We are interested in how to best communicate a (usually real valued) source to a number of destinations (sinks) over a network with capacity constraints in a collective fidelity metric over all the sinks, a problem which we call joint network-source coding. Unlike the lossless network coding problem, lossy reconstruction of the source at the sinks is permitted. We make a first attempt to characterize the set of all distortions achievable by a set of sinks in a given network. While the entire region of all achievable distortions remains largely an open problem, we find a large, non-trivial subset of it using ideas in multiple description coding. The achievable region is derived over all balanced multiple-description codes and over all network flows, while the network nodes are allowed to forward and duplicate data packets.

cs.IT

Disaster Management in Scale-Free Networks: Recovery from and Protection Against Intentional Attacks

Susceptibility of scale free Power Law (PL) networks to attacks has been traditionally studied in the context of what may be termed as {\em instantaneous attacks}, where a randomly selected set of nodes and edges are deleted while the network is kept {\em static}. In this paper, we shift the focus to the study of {\em progressive} and instantaneous attacks on {\em reactive} grown and random PL networks, which can respond to attacks and take remedial steps. In the process, we present several techniques that managed networks can adopt to minimize the damages during attacks, and also to efficiently recover from the aftermath of successful attacks. For example, we present (i) compensatory dynamics that minimize the damages inflicted by targeted progressive attacks, such as linear-preferential deletions of nodes in grown PL networks; the resulting dynamic naturally leads to the emergence of networks with PL degree distributions with exponential cutoffs; (ii) distributed healing algorithms that can scale the maximum degree of nodes in a PL network using only local decisions, and (iii) efficient means of creating giant connected components in a PL network that has been fragmented by attacks on a large number of high-degree nodes. Such targeted attacks are considered to be a major vulnerability of PL networks; however, our results show that the introduction of only a small number of random edges, through a {\em reverse percolation} process, can restore connectivity, which in turn allows restoration of other topological properties of the original network. Thus, the scale-free nature of the networks can itself be effectively utilized for protection and recovery purposes.

cond-mat.stat-mech

Let Your CyberAlter Ego Share Information and Manage Spam

Almost all of us have multiple cyberspace identities, and these {\em cyber}alter egos are networked together to form a vast cyberspace social network. This network is distinct from the world-wide-web (WWW), which is being queried and mined to the tune of billions of dollars everyday, and until recently, has gone largely unexplored. Empirically, the cyberspace social networks have been found to possess many of the same complex features that characterize its real counterparts, including scale-free degree distributions, low diameter, and extensive connectivity. We show that these topological features make the latent networks particularly suitable for explorations and management via local-only messaging protocols. {\em Cyber}alter egos can communicate via their direct links (i.e., using only their own address books) and set up a highly decentralized and scalable message passing network that can allow large-scale sharing of information and data. As one particular example of such collaborative systems, we provide a design of a spam filtering system, and our large-scale simulations show that the system achieves a spam detection rate close to 100%, while the false positive rate is kept around zero. This system has several advantages over other recent proposals (i) It uses an already existing network, created by the same social dynamics that govern our daily lives, and no dedicated peer-to-peer (P2P) systems or centralized server-based systems need be constructed; (ii) It utilizes a percolation search algorithm that makes the query-generated traffic scalable; (iii) The network has a built in trust system (just as in social networks) that can be used to thwart malicious attacks; iv) It can be implemented right now as a plugin to popular email programs, such as MS Outlook, Eudora, and Sendmail.

physics.soc-ph

Multiple Scale-Free Structures in Complex Ad-Hoc Networks

This paper develops a framework for analyzing and designing dynamic networks comprising different classes of nodes that coexist and interact in one shared environment. We consider {\em ad hoc} (i.e., nodes can leave the network unannounced, and no node has any global knowledge about the class identities of other nodes) {\em preferentially grown networks}, where different classes of nodes are characterized by different sets of local parameters used in the stochastic dynamics that all nodes in the network execute. We show that multiple scale-free structures, one within each class of nodes, and with tunable power-law exponents (as determined by the sets of parameters characterizing each class) emerge naturally in our model. Moreover, the coexistence of the scale-free structures of the different classes of nodes can be captured by succinct phase diagrams, which show a rich set of structures, including stable regions where different classes coexist in heavy-tailed and light-tailed states, and sharp phase transitions. Finally, we show how the dynamics formulated in this paper will serve as an essential part of {\em ad-hoc networking protocols}, which can lead to the formation of robust and efficiently searchable networks (including, the well-known Peer-To-Peer (P2P) networks) even under very dynamic conditions.

cond-mat.dis-nn

Scalable Percolation Search in Power Law Networks

We introduce a scalable searching algorithm for finding nodes and contents in random networks with Power-Law (PL) and heavy-tailed degree distributions. The network is searched using a probabilistic broadcast algorithm, where a query message is relayed on each edge with probability just above the bond percolation threshold of the network. We show that if each node caches its directory via a short random walk, then the total number of {\em accessible contents exhibits a first-order phase transition}, ensuring very high hit rates just above the percolation threshold. In any random PL network of size, $N$, and exponent, $2 \leq τ< 3$, the total traffic per query scales sub-linearly, while the search time scales as $O(\log N)$. In a PL network with exponent, $τ\approx 2$, {\em any content or node} can be located in the network with {\em probability approaching one} in time $O(\log N)$, while generating traffic that scales as $O(\log^2 N)$, if the maximum degree, $k_{max}$, is unconstrained, and as $O(N^{{1/2}+ε})$ (for any $ε>0$) if $ k_{max}=O(\sqrt{N})$. Extensive large-scale simulations show these scaling laws to be precise. We discuss how this percolation search algorithm can be directly adapted to solve the well-known scaling problem in unstructured Peer-to-Peer (P2P) networks. Simulations of the protocol on sample large-scale subnetworks of existing P2P services show that overall traffic can be reduced by almost two-orders of magnitude, without any significant loss in search performance.

cond-mat.dis-nn

Scale-Free and Stable Structures in Complex {\em Ad hoc} networks

Unlike the well-studied models of growing networks, where the dominant dynamics consist of insertions of new nodes and connections, and rewiring of existing links, we study {\em ad hoc} networks, where one also has to contend with rapid and random deletions of existing nodes (and, hence, the associated links). We first show that dynamics based {\em only} on the well-known preferential attachments of new nodes {\em do not} lead to a sufficiently heavy-tailed degree distribution in {\em ad hoc} networks. In particular, the magnitude of the power-law exponent increases rapidly (from 3) with the deletion rate, becoming $\infty$ in the limit of equal insertion and deletion rates. \iffalse ; thus, forcing the degree distribution to be essentially an exponential one.\fi We then introduce a {\em local} and {\em universal} {\em compensatory rewiring} dynamic, and show that even in the limit of equal insertion and deletion rates true scale-free structures emerge, where the degree distributions obey a power-law with a tunable exponent, which can be made arbitrarily close to -2. These results provide the first-known evidence of emergence of scale-free degree distributions purely due to dynamics, i.e., in networks of almost constant average size. The dynamics discovered in this paper can be used to craft protocols for designing highly dynamic Peer-to-Peer networks, and also to account for the power-law exponents observed in existing popular services.

cond-mat.dis-nn

A Random Structure for Optimum Cache Size Distributed hash table (DHT) Peer-to-Peer design

We propose a new and easily-realizable distributed hash table (DHT) peer-to-peer structure, incorporating a random caching strategy that allows for {\em polylogarithmic search time} while having only a {\em constant cache} size. We also show that a very large class of deterministic caching strategies, which covers almost all previously proposed DHT systems, can not achieve polylog search time with constant cache size. In general, the new scheme is the first known DHT structure with the following highly-desired properties: (a) Random caching strategy with constant cache size; (b) Average search time of $O(log^{2}(N))$; (c) Guaranteed search time of $O(log^{3}(N))$; (d) Truly local cache dynamics with constant overhead for node deletions and additions; (e) Self-organization from any initial network state towards the desired structure; and (f) Allows a seamless means for various trade-offs, e.g., search speed or anonymity at the expense of larger cache size.

cs.NI