SearcharxivSearch

arXiv subjects

Nicola Perra

Publications and source records attributed to Nicola Perra.

At least 19 recordsLinked to original sources

A systematic review of COVID-19 epidemic models with endogenous human behaviour. What's next?

Human behaviour and epidemic dynamics are intertwined, yet accounting for this feedback remains one of the key challenges of epidemiological modelling. The COVID-19 pandemic was an opportunity to overcome the traditional limitations of the field, raising expectations that data-informed endogenous approaches to behaviour modelling would advance substantially. To quantify the progresses made, we conducted a systematic review of SARS-CoV-2 transmission models endogenously including human behaviour in response to epidemic dynamics. The COVID-19 pandemic saw great strides in terms of the expanded use of empirical data in epi-behavioural modelling. However, it also showed shortcomings with respect to limited use of behavioural empirical data, lack of innovation in model structure, and limited engagement with other disciplines and decision-makers. Overall, our results suggest that identifying priorities in model design and behavioural data, building an adequate data collection infrastructure, leveraging on AI advancements, and fostering interdisciplinarity are strategies of utmost importance for pandemic preparedness.

physics.soc-ph

Unequal changes in commuting patterns across socio-economic strata in response to pandemic restrictions

Commuting patterns are a central component of urban dynamics and many societal activities. Exogenous shocks, such as a pandemic, might drastically modify them inducing heterogeneous variations across socioeconomic strata. Here, we quantify changes in work commuting patterns in Bogot\'a, Colombia during three different periods of the COVID-19 pandemic: pre-pandemic (2019), COVID-19 restrictions (2020), and partial reopening (2021). To this end, we use anonymized mobile phone data to infer home and work locations from recurring nighttime and weekday connection patterns, and to build daily commuting metrics. We aggregate mobility flows by administrative boundaries and socioeconomic strata. Additionally, we enrich the dataset with a range of other variables such as territorial vocation (i.e., urban versus rural), demographic information (i.e., population density) and, as a proxy for digital infrastructure quality, geolocated Speedtest measurements from Ookla. We find a marked reduction of commuting during restrictions in 2020 and a strong recovery in 2021, but with persistent heterogeneity across socioeconomic strata. Indeed, while commuting declined similarly across income groups during restrictions, groups of the population in the lower-income bracket rebounded faster to pre-pandemic levels. On the contrary, we find that groups in the higher-income bracket managed to keep higher stay-at-home behavior. Regression analyses reveal that territorial characteristics and disparities in digital connectivity significantly contribute to these differences, suggesting that infrastructure investments could help mitigate mobility-based inequalities.

physics.soc-ph

Controlling the spread of deception-based cyber-threats on time-varying networks

We study the efficacy of strategies aimed at controlling the spread of deception-based cyber-threats unfolding on online social networks. We model directed and temporal interactions between users using a family of activity-driven networks featuring tunable homophily levels among gullibility classes. We simulate the spreading of cyber-threats using classic Susceptible-Infected-Susceptible (SIS) models. We explore and quantify the effectiveness of four control strategies. Akin to vaccination campaigns with a limited budget, each strategy selects a fraction of nodes with the aim to increase their awareness and provide protection from cyber-threats. The first strategy picks nodes randomly. The second assumes global knowledge of the system selecting nodes based on their activity. The third picks nodes via egocentric sampling. The fourth selects nodes based on the outcome of standard security awareness tests, customarily used by institutions to probe, estimate, and raise the awareness of their workforce. We quantify the impact of each strategy by deriving analytically how they affect the spreading threshold. Analytical expressions are validated via large-scale numerical simulations. Interestingly, we find that targeted strategies, focusing on key features of the population such as the activity, are extremely effective. Egocentric sampling strategies, though not as effective, emerge as clear second best despite not assuming any knowledge about the system. Interestingly, we find that networks characterized by highly homophilic interactions linked to gullibility might expand the range of transmissibility parameters that allows for macroscopic outbreaks. At the same time, they reduce the reach of these spreading events. Hence, rather isolated patches of the network formed by highly gullible individuals might provide fertile grounds for the propagation and survival of cyber-threats.

physics.soc-ph

Bridging the Digital Divide: Mapping Internet Connectivity Evolution, Inequalities, and Resilience in six Brazilian Cities

We investigate the evolution of Internet speed and its implications for access to key digital services, as well as the resilience of the network during crises, focusing on six major Brazilian cities: Belo Horizonte, Brasília, Fortaleza, Manaus, Rio de Janeiro, and São Paulo. Leveraging a unique dataset of Internet Speedtest results provided by Ookla, we analyze Internet speed trends from 2017 to 2023. Our findings reveal significant improvements in Internet speed across all cities. However, we find that prosperous areas generally exhibit better Internet access, and that the dependence of Internet quality on wealth have increased over time. Additionally, we investigate the impact of Internet quality on access to critical online services, focusing on e-learning. Our analysis shows that nearly 13% of catchment areas around educational facilities have Internet speeds below the threshold required for e-learning, with less rich areas experiencing more significant challenges. Moreover, we investigate the network's resilience during the COVID-19 pandemic, finding a sharp decline in network quality following the declaration of national emergency. We also find that less wealthy areas experience larger drops in network quality during crises. Overall, this study underscores the importance of addressing disparities in Internet access to ensure equitable access to digital services and enhance network resilience during crises

physics.soc-ph

The Concept of Decentralization Through Time and Disciplines: A Quantitative Exploration

Decentralization is a pervasive concept found across disciplines, including Economics, Political Science, and Computer Science, where it is used in distinct yet interrelated ways. Here, we develop and publicly release a general pipeline to investigate the scholarly history of the term, analysing 425,144 academic publications that refer to (de)centralization. We find that the fraction of papers on the topic has been exponentially increasing since the 1950s. In 2021, 1 author in 154 mentioned (de)centralization in the title or abstract of an article. Using both semantic information and citation patterns, we cluster papers in fields and characterize the knowledge flows between them. Our analysis reveals that the topic has independently emerged in the different fields, with small cross-disciplinary contamination. Moreover, we show how Blockchain has become the most influential field about 10 years ago, while Governance dominated before the 1990s. In summary, our findings provide a quantitative assessment of the evolution of a key yet elusive concept, which has undergone cycles of rise and fall within different fields. Our pipeline offers a powerful tool to analyze the evolution of any scholarly term in the academic literature, providing insights into the interplay between collective and independent discoveries in science.

physics.soc-ph

Modeling Self-Propagating Malware with Epidemiological Models

Self-propagating malware (SPM) has recently resulted in large financial losses and high social impact, with well-known campaigns such as WannaCry and Colonial Pipeline being able to propagate rapidly on the Internet and cause service disruptions. To date, the propagation behavior of SPM is still not well understood, resulting in the difficulty of defending against these cyber threats. To address this gap, in this paper we perform a comprehensive analysis of a newly proposed epidemiological model for SPM propagation, Susceptible-Infected-Infected Dormant-Recovered (SIIDR). We perform a theoretical analysis of the stability of the SIIDR model and derive its basic reproduction number by representing it as a system of Ordinary Differential Equations with continuous time. We obtain access to 15 WananCry attack traces generated under various conditions, derive the model's transition rates, and show that SIIDR fits best the real data. We find that the SIIDR model outperforms more established compartmental models from epidemiology, such as SI, SIS, and SIR, at modeling SPM propagation.

cs.CR

Generalized contact matrices for epidemic modeling

Contact matrices have become a key ingredient of modern epidemic models. They account for the stratification of contacts for the age of individuals and, in some cases, the context of their interactions. However, age and context are not the only factors shaping contact structures and affecting the spreading of infectious diseases. Socio-economic status (SES) variables such as wealth, ethnicity, and education play a major role as well. Here, we introduce generalized contact matrices capable of stratifying contacts across any number of dimensions including any SES variable. We derive an analytical expression for the basic reproductive number of an infectious disease unfolding on a population characterized by such generalized contact matrices. Our results, on both synthetic and real data, show that disregarding higher levels of stratification might lead to the under-estimation of the reproductive number and to a mis-estimation of the global epidemic dynamics. Furthermore, including generalized contact matrices allows for more expressive epidemic models able to capture heterogeneities in behaviours such as different levels of adoption of non-pharmaceutical interventions across different groups. Overall, our work contributes to the literature attempting to bring socio-economic, as well as other dimensions, to the forefront of epidemic modeling. Tackling this issue is crucial for developing more precise descriptions of epidemics, and thus to design better strategies to contain them.

physics.soc-ph

The structure of segregation in co-authorship networks and its impact on scientific production

Co-authorship networks, where nodes represent authors and edges represent co-authorship relations, are key to understanding the production and diffusion of knowledge in academia. Social constructs, biases (implicit and explicit), and constraints (e.g. spatial, temporal) affect who works with whom and cause co-authorship networks to organise into tight communities with different levels of segregation. We aim to look at aspects of the co-authorship network structure that lead to segregation and its impact on scientific production. We measure segregation using the Spectral Segregation Index (SSI) and find 4 ordered segregation categories: completely segregated, highly segregated, moderately segregated and non-segregated communities. We direct our attention to the non-segregated and highly segregated communities, quantifying and comparing their structural topologies and k-core positions. When considering communities of both categories (controlling for size), our results show no differences in density and clustering but substantial variability in core position. Larger non-segregated communities are more likely to occupy cores near the network nucleus, while the highly segregated ones tend to be closer to the network periphery. Finally, we analyse differences in citations gained by researchers within communities showing different segregation categories. Researchers in highly segregated communities get more citations from their community members in middle cores and gain more citations per publication in middle/periphery cores. Those in non-segregated communities get more citations per publication in the nucleus. To our knowledge, this work is the first to characterise community segregation in co-authorship networks and investigate the relationship between community segregation and author citations.

cs.SI

Cyber Network Resilience against Self-Propagating Malware Attacks

Self-propagating malware (SPM) has led to huge financial losses, major data breaches, and widespread service disruptions in recent years. In this paper, we explore the problem of developing cyber resilient systems capable of mitigating the spread of SPM attacks. We begin with an in-depth study of a well-known self-propagating malware, WannaCry, and present a compartmental model called SIIDR that accurately captures the behavior observed in real-world attack traces. Next, we investigate ten cyber defense techniques, including existing edge and node hardening strategies, as well as newly developed methods based on reconfiguring network communication (NodeSplit) and isolating communities. We evaluate all defense strategies in detail using six real-world communication graphs collected from a large retail network and compare their performance across a wide range of attacks and network topologies. We show that several of these defenses are able to efficiently reduce the spread of SPM attacks modeled with SIIDR. For instance, given a strong attack that infects 97% of nodes when no defense is employed, strategically securing a small number of nodes (0.08%) reduces the infection footprint in one of the networks down to 1%.

cs.CR

Modeling Teams Performance Using Deep Representational Learning on Graphs

The large majority of human activities require collaborations within and across formal or informal teams. Our understanding of how the collaborative efforts spent by teams relate to their performance is still a matter of debate. Teamwork results in a highly interconnected ecosystem of potentially overlapping components where tasks are performed in interaction with team members and across other teams. To tackle this problem, we propose a graph neural network model designed to predict a team's performance while identifying the drivers that determine such an outcome. In particular, the model is based on three architectural channels: topological, centrality, and contextual which capture different factors potentially shaping teams' success. We endow the model with two attention mechanisms to boost model performance and allow interpretability. A first mechanism allows pinpointing key members inside the team. A second mechanism allows us to quantify the contributions of the three driver effects in determining the outcome performance. We test model performance on a wide range of domains outperforming most of the classical and neural baselines considered. Moreover, we include synthetic datasets specifically designed to validate how the model disentangles the intended properties on which our model vastly outperforms baselines.

cs.SI

Macroscopic properties of buyer-seller networks in online marketplaces

Online marketplaces are the main engines of legal and illegal e-commerce, yet their empirical properties are poorly understood due to the absence of large-scale data. We analyze two comprehensive datasets containing 245M transactions (16B USD) that took place on online marketplaces between 2010 and 2021, covering 28 dark web marketplaces, i.e., unregulated markets whose main currency is Bitcoin, and 144 product markets of one popular regulated e-commerce platform. We show that transactions in online marketplaces exhibit strikingly similar patterns despite significant differences in language, lifetimes, products, regulation, and technology. Specifically, we find remarkable regularities in the distributions of transaction amounts, number of transactions, inter-event times and time between first and last transactions. We show that buyer behavior is affected by the memory of past interactions and use this insight to propose a model of network formation reproducing our main empirical observations. Our findings have implications for understanding market power on online marketplaces as well as inter-marketplace competition, and provide empirical foundation for theoretical economic models of online marketplaces.

physics.soc-ph

The adoption of non-pharmaceutical interventions and the role of digital infrastructure during the COVID-19 Pandemic in Colombia, Ecuador, and El Salvador

Adherence to the non-pharmaceutical interventions (NPIs) put in place to mitigate the spreading of infectious diseases is a multifaceted problem. Socio-demographic, socio-economic, and epidemiological factors can influence the perceived susceptibility and risk which are known to affect behavior. Furthermore, the adoption of NPIs is dependent upon the barriers, real or perceived, associated with their implementation. We study the determinants of NPIs adherence during the first wave of the COVID-19 Pandemic in Colombia, Ecuador, and El Salvador. Analyses are performed at the level of municipalities and include socio-economic, socio-demographic, and epidemiological indicators. Furthermore, by leveraging a unique dataset comprising tens of millions of internet Speedtest measurements from Ookla, we investigate the quality of the digital infrastructure as a possible barrier to adoption. We use publicly available data provided by Meta capturing aggregated mobility changes as a proxy of adherence to NPIs. Across the three countries considered, we find a significant correlation between mobility drops and digital infrastructure quality. The relationship remains significant after controlling for several factors including socio-economic status, population size, and reported COVID-19 cases. This finding suggests that municipalities with better connectivity were able to afford higher mobility reductions. The link between mobility drops and digital infrastructure quality is stronger at the peak of NPIs stringency. We also find that mobility reductions were more pronounced in larger, denser, and wealthier municipalities. Our work provides new insights on the significance of access to digital tools as an additional factor influencing the ability to follow social distancing guidelines during a health emergency

physics.soc-ph

Non-pharmaceutical interventions during the COVID-19 pandemic: a rapid review

Infectious diseases and human behavior are intertwined. On one side, our movements and interactions are the engines of transmission. On the other, the unfolding of viruses might induce changes to our daily activities. While intuitive, our understanding of such feedback loop is still limited. Before COVID-19 the literature on the subject was mainly theoretical and largely missed validation. The main issue was the lack of empirical data capturing behavioral change induced by diseases. Things have dramatically changed in 2020. Non-pharmaceutical interventions (NPIs) have been the key weapon against the SARS-CoV-2 virus and affected virtually any societal process. Travels bans, events cancellation, social distancing, curfews, and lockdowns have become unfortunately very familiar. The scale of the emergency, the ease of survey as well as crowdsourcing deployment guaranteed by the latest technology, several Data for Good programs developed by tech giants, major mobile phone providers, and other companies have allowed unprecedented access to data describing behavioral changes induced by the pandemic. Here, I aim to review some of the vast literature written on the subject of NPIs during the COVID-19 pandemic. In doing so, I analyze 347 articles written by more than 2518 of authors in the last $12$ months. While the large majority of the sample was obtained by querying PubMed, it includes also a hand-curated list. Considering the focus, and methodology I have classified the sample into seven main categories: epidemic models, surveys, comments/perspectives, papers aiming to quantify the effects of NPIs, reviews, articles using data proxies to measure NPIs, and publicly available datasets describing NPIs. I summarize the methodology, data used, findings of the articles in each category and provide an outlook highlighting future challenges as well as opportunities

physics.soc-ph

Self-initiated behavioural change and disease resurgence on activity-driven networks

We consider a population that experienced a first wave of infections, interrupted by strong, top-down, governmental restrictions and did not develop a significant immunity to prevent a second wave (i.e. resurgence). As restrictions are lifted, individuals adapt their social behaviour to minimize the risk of infection. We consider two scenarios. In the first, individuals reduce their overall social activity towards the rest of the population. In the second scenario, they maintain a normal social activity within a small community of peers (i.e., social bubble) while reducing social interactions with the rest of the population. In both cases, we consider possible correlations between social activity and behaviour change, reflecting for example the social dimension of certain occupations. We model these scenarios considering a Susceptible-Infected-Recovered epidemic model unfolding on activity-driven networks. Extensive analytical and numerical results show that i) a minority of very active individuals not changing behaviour may nullify the efforts of the large majority of the population, and ii) imperfect social bubbles of normal social activity may be less effective than an overall reduction of social interactions.

physics.soc-ph

Finding Patient Zero: Learning Contagion Source with Graph Neural Networks

Locating the source of an epidemic, or patient zero (P0), can provide critical insights into the infection's transmission course and allow efficient resource allocation. Existing methods use graph-theoretic centrality measures and expensive message-passing algorithms, requiring knowledge of the underlying dynamics and its parameters. In this paper, we revisit this problem using graph neural networks (GNNs) to learn P0. We establish a theoretical limit for the identification of P0 in a class of epidemic models. We evaluate our method against different epidemic models on both synthetic and a real-world contact network considering a disease with history and characteristics of COVID-19. % We observe that GNNs can identify P0 close to the theoretical bound on accuracy, without explicit input of dynamics or its parameters. In addition, GNN is over 100 times faster than classic methods for inference on arbitrary graph topologies. Our theoretical bound also shows that the epidemic is like a ticking clock, emphasizing the importance of early contact-tracing. We find a maximum time after which accurate recovery of the source becomes impossible, regardless of the algorithm used.

cs.SI

Collective response to the media coverage of COVID-19 Pandemic on Reddit and Wikipedia

The exposure and consumption of information during epidemic outbreaks may alter risk perception, trigger behavioural changes, and ultimately affect the evolution of the disease. It is thus of the uttermost importance to map information dissemination by mainstream media outlets and public response. However, our understanding of this exposure-response dynamic during COVID-19 pandemic is still limited. In this paper, we provide a characterization of media coverage and online collective attention to COVID-19 pandemic in four countries: Italy, United Kingdom, United States, and Canada. For this purpose, we collect an heterogeneous dataset including 227,768 online news articles and 13,448 Youtube videos published by mainstream media, 107,898 users posts and 3,829,309 comments on the social media platform Reddit, and 278,456,892 views to COVID-19 related Wikipedia pages. Our results show that public attention, quantified as users activity on Reddit and active searches on Wikipedia pages, is mainly driven by media coverage and declines rapidly, while news exposure and COVID-19 incidence remain high. Furthermore, by using an unsupervised, dynamical topic modeling approach, we show that while the attention dedicated to different topics by media and online users are in good accordance, interesting deviations emerge in their temporal patterns. Overall, our findings offer an additional key to interpret public perception/response to the current global health emergency and raise questions about the effects of attention saturation on collective awareness, risk perception and thus on tendencies towards behavioural changes.

cs.SI

Towards a data-driven characterization of behavioral changes induced by the seasonal flu

In this work, we aim to determine the main factors driving behavioral change during the seasonal flu. To this end, we analyze a unique dataset comprised of 599 surveys completed by 434 Italian users of Influweb, a Web platform for participatory surveillance, during the 2017-18 and 2018-19 seasons. The data provide socio-demographic information, level of concerns about the flu, past experience with illnesses, and the type of behavioral changes implemented by each participant. We describe each response with a set of features and divide them in three target categories. These describe those that report i) no (26 %), ii) only moderately (36 %), iii) significant (38 %) changes in behaviors. In these settings, we adopt machine learning algorithms to investigate the extent to which target variables can be predicted by looking only at the set of features. Notably, $66\%$ of the samples in the category describing more significant changes in behaviors are correctly classified through Gradient Boosted Trees. Furthermore, we investigate the importance of each feature in the classification task and uncover complex relationships between individuals' characteristics and their attitude towards behavioral change. We find that intensity, recency of past illnesses, perceived susceptibility to and perceived severity of an infection are the most significant features in the classification task. Interestingly, the last two match the theoretical constructs suggested by the Health-Belief Model. Overall, the research contributes to the small set of empirical studies devoted to the data-driven characterization of behavioral changes induced by infectious diseases.

physics.soc-ph

Explore with caution: mapping the evolution of scientific interest in Physics

In the book The Essential Tension Thomas Kuhn described the conflict between tradition and innovation in scientific research --i.e., the desire to explore new promising areas, counterposed to the need to capitalize on the work done in the past. While it is true that along their careers many scientists probably felt this tension, only few works have tried to quantify it. Here, we address this question by analyzing a large-scale dataset, containing all the papers published by the American Physical Society (APS) in more than $25$ years, which allows for a better understanding of scientists' careers evolution in Physics. We employ the Physics and Astronomy Classification Scheme (PACS) present in each paper to map the scientific interests of $181,397$ authors and their evolution along the years. Our results indeed confirm the existence of the `essential tension' with scientists balancing between exploring the boundaries of their area and exploiting previous work. In particular, we found that although the majority of physicists change the topics of their research, they stay within the same broader area thus exploring with caution new scientific endeavors. Furthermore, we quantify the flows of authors moving between different subfields and pinpoint which areas are more likely to attract or donate researchers to the other ones. Overall, our results depict a very distinctive portrait of the evolution of research interests in Physics and can help in designing specific policies for the future.

physics.soc-ph