SearcharxivSearch

arXiv · 1603.08273

Collective Influence Algorithm to find influencers via optimal percolation in massively large social media

Abstract

We elaborate on a linear time implementation of the Collective Influence (CI) algorithm introduced by Morone, Makse, Nature 524, 65 (2015) to find the minimal set of influencers in a network via optimal percolation. We show that the computational complexity of CI is O(N log N) when removing nodes one-by-one, with N the number of nodes. This is made possible by using an appropriate data structure to process the CI values, and by the finite radius l of the CI sphere. Furthermore, we introduce a simple extension of CI when l is infinite, the CI propagation (CI_P) algorithm, that considers the global optimization of influence via message passing in the whole network and identifies a slightly smaller fraction of influencers than CI. Remarkably, CI_P is able to reproduce the exact analytical optimal percolation threshold obtained by Bau, Wormald, Random Struct. Alg. 21, 397 (2002) for cubic random regular graphs, leaving little improvement left for random graphs. We also introduce the Collective Immunization Belief Propagation algorithm (CI_BP), a belief-propagation (BP) variant of CI based on optimal immunization, which has the same performance as CI_P. However, this small augmented performance of the order of 1-2 % in the low influencers tail comes at the expense of increasing the computational complexity from O(N log N) to O(N^2 log N), rendering both, CI_P and CI_BP, prohibitive for finding influencers in modern-day big-data. The same nonlinear running time drawback pertains to a recently introduced BP-decimation (BPD) algorithm by Mugisha, Zhou, arXiv:1603.05781. For instance, we show that for big-data social networks of typically 200 million users (eg, active Twitter users sending 500 million tweets per day), CI finds the influencers in less than 3 hours running on a single CPU, while the BP algorithms (CI_P, CI_BP and BDP) would take more than 3,000 years to accomplish the same task.

Explore related subjects

Keep this discovery

BibTeXRIS

Flaviano Morone, Byungjoon Min, Lin Bo, Romain Mari, Hernan A. Makse. 2016-03-28. Collective Influence Algorithm to find influencers via optimal percolation in massively large social media. https://arxiv.org/abs/1603.08273

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Energy pathway variety and the progress of the energy transition in European countries

The integration of new energy forms into existing energy infrastructure has emerged as a critical challenge in the context of the pursuit of a sustainable energy transition. One of the main challenges is understanding how this integration takes place not only from the introduction, but also as energy follows existing paths or creates new ones through which it is transformed and used by different activities. Here we introduce techniques from network science to analyse this process for the case of 29 European countries between 1992 and 2021. We study how new energy forms increase or decrease the variety (heterogeneity) of paths through the system of each country by establishing new ones and replacing or phasing out existing ones. We find that the transition to systems based on renewable energy is characterised by an initial increase in the variety of paths while the heterogeneity of paths decreases at the end of the transition, when the proportion of non-renewables in the system tends to zero. We then demonstrate that greater heterogeneity (complexity) is associated with larger annual fluctuations in the proportion of non-renewable sources in the system, establishing a direct relationship between the progress of the transition and the complexity of the energy system in which it occurs. This contributes to the understanding of general properties of the dynamics of the energy transition and effects that accelerate or deter it.

physics.soc-ph

Fundamental limits to identifying node and tie memory in temporal networks: marginal artefacts and spreading dynamics

Temporal-network models attribute memory in contact data to either node self-excitation (branching ratio n_node) or tie reinforcement (kappa), carrying major consequences for epidemic spreading. We prove that when event initiators are observed, the two mechanisms are orthogonal: the Fisher information is block-diagonal and neither trades off against the other. In undirected proximity data, where initiators are unobserved, marginalising over them couples the mechanisms into a structural confound that survives posterior smoothing. On empirical proximity, messaging, and email records, however, a cruder failure dominates: fitted node memory is pinned to the inter-event marginal law and remains virtually invariant across latent label posterior samples (coefficient of variation below 1%). An inter-event-order shuffle test and burstiness-memory diagnostics reveal that exponential-Hawkes node memory is recovered from none, while tie reinforcement remains identifiable throughout. This near-unidentifiability is intrinsic, not an artefact of the exponential kernel: refitting flexible scale-free (sum-of-exponentials) kernels on synthetic power-law self-exciting processes fails to distinguish genuine node memory from memoryless renewal controls, with identical collapses recurring on algorithmic networks (edit bots, cloud microservices) and cortical spiking. Downstream epidemic consequences are quantitative: simulations fitted to empirical contact records under-predict outbreak sizes by up to a factor of 2.5 and shift the epidemic threshold. We conclude that observational temporal networks face a two-fold identifiability boundary: contact directionality is essential to decouple tie reinforcement, whereas heavy-tailed node self-excitation is intrinsically unidentifiable from contact timings alone.

physics.soc-ph

Assessing extreme flood impacts on urban rail transit: A passenger-oriented, resilience-informed framework

Urban rail transit systems (URTSs) are increasingly exposed to extreme floods following heavy precipitation, yet passenger travel impacts are often assessed through delay-based indicators that overlook infeasible journeys under large-scale disruptions. This study develops a passenger-oriented, resilience-informed framework for assessing flood impacts on URTS journeys from disruption onset to recovery completion. The framework presents a novel six-category classification of journey impacts, explicitly considering rerouting, alternative station use, and a delay threshold. It is demonstrated through hourly dynamic simulations of 15 London URTS lines under 30-year, 100-year, and 1,000-year flood risk scenarios. Results indicate that severe flood disruptions lead to substantial unsatisfied demand, driven primarily by unavailable routes rather than unacceptable delays. Compared with finer behaviour adjustments, rerouting dominates travel impacts. These findings highlight the significance of moving beyond delay-based assessment and provide valuable evidence on essential behavioural mechanisms for strategic-level stress testing intended to inform URTS flood resilience intervention planning.

physics.soc-ph