SearcharxivSearch

arXiv subjects

Santo Fortunato

Publications and source records attributed to Santo Fortunato.

At least 19 recordsLinked to original sources

The backbone of science: analysis of citation networks between papers and their sources

The bibliography of scientific papers lists items with variable degree of relevance for the contents of the paper itself. If we could identify the sources, i.e., the works that actually inspired the paper, their citations can help us uncover the genesis of scientific projects and would be more representative of the actual importance of papers and authors than the standard citation counts, when all references are considered. Here we present an analysis of the \textit{backbone of science}, i.e., the network of citations between papers and their sources. The latter are extracted from the full body of papers via Large Language Models (LLMs), which are currently very capable of correctly identifying the context in which a paper is cited. Using two different but related prompts, we find that the LLMs select only a small set of references, not taken at random, and that the resulting backbone networks are quite similar to each other with respect to their in-degree distributions, modularity, transitivity, and degree correlations. Backbone networks have higher heterogeneity in their in-degree distributions, compared to the full network, but the most cited papers are usually the same, with some important exceptions. Citation rankings among authors are also remarkably stable. We conclude that the full citation network, despite its redundancy with respect to the backbones, presents a reliable picture of the relative citation impact of papers and authors.

cs.DL

Heuristic and exact modularity optimization with size-constrained communities

When searching for communities in networks, domain experts may have some prior expectations about the size of communities. Yet, community detection methods normally do not optimize communities under cluster size constraints. Multi-resolution techniques allow users to indirectly control the average community size through changing a resolution parameter, but this practice does not control the size of individual communities. We here study the problem of size-constrained community detection, where the size of all communities is limited to a user-specified range of values, in the context of modularity optimization. We propose a heuristic for modularity optimization under community size constraints. To demonstrate the reliability of our proposed heuristic, we also formulate an exact integer optimization model and use its results as a baseline. Our analysis based on synthetic benchmarks and real networks demonstrate the issues with the currently common practice of changing resolution parameters and reveal the advantages of the proposed methods as a principled way of obtaining size-constrained communities. The proposed method is publicly available in the Python Leiden algorithm package.

physics.soc-ph

Modeling plant disease spread via high-resolution human mobility networks

Human mobility plays a crucial role in the spread of human diseases, but is rarely quantified in plant disease epidemics. To address this gap, we integrate a unique, high-resolution network of human movements in New Zealand with a metapopulation model to mechanistically simulate pathogen transmission. We calibrate the model on the nationwide 2010 kiwifruit vine disease (Psa-V) outbreak, and show that it accurately reproduces the observed spatiotemporal spread, confirming that the human mobility network is a strong foundation for modeling transmission dynamics. By analyzing spatial infection trends, we find that most dispersal occurs locally, as often illustrated in the plant-outbreak literature. However, sporadic long-range connections are necessary to model a nationwide outbreak. Using the model as an in-silico laboratory, we demonstrate that enhanced surveillance accelerates detection and that outbreak severity is highly sensitive to the timing and location of initial disease importation. We observe a potential causal link between seasonal labor patterns and epidemic risk in high-traffic seasons. This study provides a robust, data-driven framework for modeling and predicting the spatiotemporal spread of agricultural pathogens. It underscores the importance of leveraging human mobility networks to design timely interventions and surveillance systems, protecting global food security.

physics.soc-ph

Robustness in sparse artificial neural networks trained with adaptive topology

We investigate the robustness of sparse artificial neural networks trained with adaptive topology. We focus on a simple yet effective architecture consisting of three sparse layers with 99% sparsity followed by a dense layer, applied to image classification tasks such as MNIST and Fashion MNIST. By updating the topology of the sparse layers between each epoch, we achieve competitive accuracy despite the significantly reduced number of weights. Our primary contribution is a detailed analysis of the robustness of these networks, exploring their performance under various perturbations including random link removal, adversarial attack, and link weight shuffling. Through extensive experiments, we demonstrate that adaptive topology not only enhances efficiency but also maintains robustness. This work highlights the potential of adaptive sparse networks as a promising direction for developing efficient and reliable deep learning models.

cs.LG

Graph energy as a measure of community detectability in networks

A key challenge in network science is the detection of communities, which are sets of nodes in a network that are densely connected internally but sparsely connected to the rest of the network. A fundamental result in community detection is the existence of a nontrivial threshold for community detectability on sparse graphs that are generated by the planted partition model (PPM). Below this so-called ``detectability limit'', no community-detection method can perform better than random chance. Spectral methods for community detection fail before this detectability limit because the eigenvalues corresponding to the eigenvectors that are relevant for community detection can be absorbed by the bulk of the spectrum. One can bypass the detectability problem by using special matrices, like the non-backtracking matrix, but this requires one to consider higher-dimensional matrices. In this paper, we show that the difference in graph energy between a PPM and an Erd\H{o}s--R\'enyi (ER) network has a distinct transition at the detectability threshold even for the adjacency matrices of the underlying networks. The graph energy is based on the full spectrum of an adjacency matrix, so our result suggests that standard graph matrices still allow one to separate the parameter regions with detectable and undetectable communities.

cs.SI

What defines a group of friends? Rethinking community structure in signed, directed networks

We study the structure of personal relationships among 1068 high school students using a dataset that contains the network of self-reported friendly and conflictive relationships, with information on their directionality and intensity. We analyse the resulting weighted, directed, and signed network using a Bayesian stochastic block model framework, which enables the inference of group structure without imposing prior assumptions on the role of negative or asymmetric ties. While a full model incorporating all edge attributes yields statistically coherent clusters, these do not align with socially meaningful communities. To address this, we focus first on the network backbone of mutual affinities, and we characterize its group organization. Many communities display an assortative structure, often embedded within larger cohesive configurations, but we also observe more diverse patterns such as core-periphery structure and isolated nodes. We then examine how relationship intensity, directionality, and conflict shape group structure. Asymmetric ties, though often occurring between communities, are frequently present within them, revealing the stabilizing effect of group membership on non-mutual relationships. Furthermore, the presence of asymmetric ties does not inherently imply a hierarchical structure, given that all groups both receive and report significant levels of non-reciprocal ties. More intense ties play a disproportionate role in shaping community structure. Finally, negative ties tend to bridge communities, but we find that groups feature a significant level of internal conflict. Our research offers a new perspective on the study of group organization when rich information about the directionality, the intensity and the sign of ties is considered, with implications for identifying social vulnerability and designing targeted interventions.

physics.soc-ph

Multilayer network science: theory, methods, and applications

Multilayer network science has emerged as a central framework for analysing interconnected and interdependent complex systems. Its relevance has grown substantially with the increasing availability of rich, heterogeneous data, which makes it possible to uncover and exploit the inherently multilayered organisation of many real-world networks. In this review, we summarise recent developments in the field. On the theoretical and methodological front, we outline core concepts and survey advances in community detection, dynamical processes, temporal networks, higher-order interactions, and machine-learning-based approaches. On the application side, we discuss progress across diverse domains, including interdependent infrastructures, spreading dynamics, computational social science, economic and financial systems, ecological and climate networks, science-of-science studies, network medicine, and network neuroscience. We conclude with a forward-looking perspective, emphasizing the need for standardised datasets and software, deeper integration of temporal and higher-order structures, and a transition toward genuinely predictive models of complex systems.

physics.soc-ph

The Effects of Remote Working on Scientific Collaboration and Impact

The COVID-19 pandemic shifted academic collaboration from in-person to remote interactions. This study explores, for the first time, the effects on scientific collaborations and impact of such a shift, comparing research output before, during, and after the pandemic. Using large-scale bibliometric data, we track the evolution of collaboration networks and the resulting impact of research over time. Our findings are twofold: first, the geographic distribution of collaborations significantly shifted, with a notable increase in cross-border partnerships after 2020, indicating a reduction in the constraints of geographic proximity. Second, despite the expansion of collaboration networks, there was a concerning decline in citation impact, suggesting that the absence of spontaneous in-person interactions-which traditionally foster deep discussions and idea exchange-negatively affected research quality. As hybrid work models in academia gain traction, this study highlights the need for universities and research organizations to carefully consider the balance between remote and in-person engagement.

physics.soc-ph

Detectability threshold in weighted modular networks

We study the necessary condition to detect, by means of spectral modularity optimization, the ground-truth partition in networks generated according to the weighted planted-partition model with two equally sized communities. We analytically derive a general expression for the maximum level of mixing tolerated by the algorithm to retrieve community structure, showing that the value of this detectability threshold depends on the first two moments of the distributions of node degree and edge weight. We focus on the standard case of Poisson-distributed node degrees and compare the detectability thresholds of five edge-weight distributions: Dirac, Poisson, exponential, geometric, and signed Bernoulli. We show that Dirac distributed weights yield the smallest detectability threshold, while exponentially distributed weights increase the threshold by a factor $\sqrt{2}$, with other distributions exhibiting distinct behaviors that depend, either or both, on the average values of the degree and weight distributions. Our results indicate that larger variability in edge weights can make communities less detectable. In cases where edge weights carry no information about community structure, incorporating weights in community detection is detrimental.

physics.soc-ph

Scale invariance and statistical significance in complex weighted networks

Most networks encountered in nature, society, and technology have weighted edges, representing the strength of the interaction/association between their vertices. Randomizing the structure of a network is a classic procedure used to estimate the statistical significance of properties of the network, such as transitivity, centrality and community structure. Randomization of weighted networks has traditionally been done via the weighted configuration model (WCM), a simple extension of the configuration model, where weights are interpreted as bundles of edges. It has previously been shown that the ensemble of randomizations provided by the WCM is affected by the specific scale used to compute the weights, but the consequences for statistical significance were unclear. Here we find that statistical significance based on the WCM is scale-dependent, whereas in most cases results should be independent of the choice of the scale. More generally, we find that designing a null model that does not violate scale invariance is challenging. A two-step approach, originally introduced for network reconstruction, in which one first randomizes the structure, then the weights, with a suitable distribution, restores scale invariance, and allows us to conduct unbiased assessments of significance on weighted networks.

physics.soc-ph

The Value of Patience in Online Grocery Shopping

Since the COVID-19 pandemic, online grocery shopping has rapidly reshaped consumer behavior worldwide, fueled by ever-faster delivery promises aimed at maximizing convenience. Yet, this growth has also substantially increased urban traffic congestion, emissions, and pollution. Despite extensive research on urban delivery optimization, little is known about the trade-off between individual convenience and these societal costs. In this study, we investigate the value of marginal extensions in delivery times, termed customer patience, in mitigating the traffic burden caused by grocery deliveries. We first conceptualize the problem and present a mathematical model that highlights a convex relationship between patience and traffic congestion. The theoretical predictions are confirmed by an extensive, network-science based analysis leveraging two large-scale datasets encompassing over 8 million grocery orders in Dubai. Our findings reveal that allowing just five additional minutes in delivery time reduces daily delivery mileage by approximately 30 percent and life-cycle CO2 emissions by 20 percent. Beyond ten minutes of added patience, however, marginal benefits diminish significantly. These results highlight that modest increases in consumer patience can deliver substantial gains in traffic reduction and sustainability, offering a scalable strategy to balance individual convenience with societal welfare in urban delivery systems.

cs.SI

Toward a Physics of Deep Learning and Brains

Deep neural networks and brains both learn and share superficial similarities: processing nodes are likened to neurons and adjustable weights are likened to modifiable synapses. But can a unified theoretical framework be found to underlie them both? Here we show that the equations used to describe neuronal avalanches in living brains can also be applied to cascades of activity in deep neural networks. These equations are derived from non-equilibrium statistical physics and show that deep neural networks learn best when poised between absorbing and active phases. Because these networks are strongly driven by inputs, however, they do not operate at a true critical point but within a quasi-critical regime -- one that still approximately satisfies crackling noise scaling relations. By training networks with different initializations, we show that maximal susceptibility is a more reliable predictor of learning than proximity to the critical point itself. This provides a blueprint for engineering improved network performance. Finally, using finite-size scaling we identify distinct universality classes, including Barkhausen noise and directed percolation. This theoretical framework demonstrates that universal features are shared by both biological and artificial neural networks.

cond-mat.dis-nn

Task complexity shapes internal representations and robustness in neural networks

Neural networks excel across a wide range of tasks, yet remain black boxes. In particular, how their internal representations are shaped by the complexity of the input data and the problems they solve remains obscure. In this work, we introduce a suite of five data-agnostic probes-pruning, binarization, noise injection, sign flipping, and bipartite network randomization-to quantify how task difficulty influences the topology and robustness of representations in multilayer perceptrons (MLPs). MLPs are represented as signed, weighted bipartite graphs from a network science perspective. We contrast easy and hard classification tasks on the MNIST and Fashion-MNIST datasets. We show that binarizing weights in hard-task models collapses accuracy to chance, whereas easy-task models remain robust. We also find that pruning low-magnitude edges in binarized hard-task models reveals a sharp phase-transition in performance. Moreover, moderate noise injection can enhance accuracy, resembling a stochastic-resonance effect linked to optimal sign flips of small-magnitude weights. Finally, preserving only the sign structure-instead of precise weight magnitudes-through bipartite network randomizations suffices to maintain high accuracy. These phenomena define a model- and modality-agnostic measure of task complexity: the performance gap between full-precision and binarized or shuffled neural network performance. Our findings highlight the crucial role of signed bipartite topology in learned representations and suggest practical strategies for model compression and interpretability that align with task complexity.

cs.LG

Functional mesoscale organization of complex networks

The network density matrix (NDM) framework, enabling an information-theoretic and multiscale treatment of network flow, has been gaining momentum over the last decade. Benefiting from the counterparts of physical functions such as free energy and entropy, NDM's applications range from estimating how nodes influence network flows across scales the centrality of nodes at the local level to explaining the emergence of structural and functional order. Here, we introduce a generalized notion of the network internal energy $E_\tau$, where $\tau$ denotes a temporal hyperparameter allowing for multi-resolution analysis, showing how it measures the leakage of dynamical correlations from arbitrary partitions, where the minimally leaky subsystems have minimal $E_\tau$. Moreover, we analytically demonstrate that $E_\tau$ reduces to the well-known modularity function at the smallest temporal scale $\tau = 0$. We investigate this peculiar resemblance by comparing the communities minimizing $E_\tau$, with those detected by widely used methods like multiscale modularity and Markov stability. Our work provides a detailed analytical and computational picture of network generalized internal energy, and explores its effectiveness in detecting communities in synthetic and empirical networks within a unifying framework.

physics.soc-ph

Modeling individual attention dynamics on online social media

In the attention economy, understanding how individuals manage limited attention is critical. We introduce a simple model describing the decay of a user's engagement when facing multiple inputs. We analytically show that individual attention decay is determined by the overall duration of interactions, not their number or user activity. Our model is validated using data from Reddit's Change My View subreddit, where the user's attention dynamics is explicitly traceable. Despite its simplicity, our model offers a crucial microscopic perspective complementing macroscopic studies.

physics.soc-ph

Network community detection via neural embeddings

Recent advances in machine learning research have produced powerful neural graph embedding methods, which learn useful, low-dimensional vector representations of network data. These neural methods for graph embedding excel in graph machine learning tasks and are now widely adopted. However, how and why these methods work -- particularly how network structure gets encoded in the embedding -- remain largely unexplained. Here, we show that node2vec -- shallow, linear neural network -- encodes communities into separable clusters better than random partitioning down to the information-theoretic detectability limit for the stochastic block models. We show that this is due to the equivalence between the embedding learned by node2vec and the spectral embedding via the eigenvectors of the symmetric normalized Laplacian matrix. Numerical simulations demonstrate that node2vec is capable of learning communities on sparse graphs generated by the stochastic blockmodel, as well as on sparse degree-heterogeneous networks. Our results highlight the features of graph neural networks that enable them to separate communities in embedding space.

physics.soc-ph

Epidemic spreading in group-structured populations

Individuals involved in common group activities/settings -- e.g., college students that are enrolled in the same class and/or live in the same dorm -- are exposed to recurrent contacts of physical proximity. These contacts are known to mediate the spread of an infectious disease, however, it is not obvious how the properties of the spreading process are determined by the structure of and the interrelation among the group settings that are at the root of those recurrent interactions. Here, we show that reshaping the organization of groups within a population can be used as an effective strategy to decrease the severity of an epidemic. Specifically, we show that when group structures are sufficiently correlated -- e.g., the likelihood for two students living in the same dorm to attend the same class is sufficiently high -- outbreaks are longer but milder than for uncorrelated group structures. Also, we show that the effectiveness of interventions for disease containment increases as the correlation among group structures increases. We demonstrate the practical relevance of our findings by taking advantage of data about housing and attendance of students at the Indiana University campus in Bloomington. By appropriately optimizing the assignment of students to dorms based on their enrollment, we are able to observe a two- to five-fold reduction in the severity of simulated epidemic processes.

physics.soc-ph

What we should learn from pandemic publishing

Authors of COVID-19 papers produced during the pandemic were overwhelmingly not subject matter experts. Such a massive inflow of scholars from different expertise areas is both an asset and a potential problem. Domain-informed scientific collaboration is the key to preparing for future crises.

physics.soc-ph