Searcharxiv⌕ Search

arXiv subjects

Fabio Saracco

Publications and source records attributed to Fabio Saracco.

At least 19 recordsLinked to original sources

Beyond Direct Retweets: Multi-Step Pathways in Italian COVID-19 Twitter

We study how retweet interactions in large-scale Twitter debates are organized beyond direct links alone. Focusing on Twitter debate in Italy during the first phase of the COVID-19 pandemic, we combine a validated community-reconstruction pipeline with a higher-order random-walk framework to examine how short multi-step pathways redistribute attention across discursive communities. Rather than reconstructing observed cascades of individual tweets, we use motif-based random-walk paths as a structural device to compare direct community-to-community connectivity with the distribution of multi-step endpoints. We find that attention is initially concentrated within communities, but that this concentration weakens as path length increases. At the same time, the resulting cross-community redistribution is not uniform: some communities become increasingly relevant as endpoints of longer pathways, while others lose relative prominence. These differences are not fully captured by community size or by first-order retweet connectivity alone, and they also display important directional asymmetries when the network is analyzed under the reversed orientation. In summary, the results show that moving beyond direct retweets changes the community-level representation of online debate and reveals higher-order structural patterns that remain invisible in first-order analyses.

physics.soc-ph↗

Leveraging Content Producer Networks and User Perception to Detect Online Discursive Communities

Online discussions are often characterized by strong behavioral asymmetries: a relatively small fraction of users actively produces content, while the majority primarily consumes and redistributes it. Here we propose a community-detection framework for online social networks that exploits this asymmetry by first identifying and clustering a set of leading users, and then extending the resulting labels to the broader user base. We introduce two complementary strategies to cluster leaders, one based on their mutual interactions and the other on audience overlap, both relying on entropy-based filtering to separate signal from noise. We evaluate the framework on three major Italian political debates on Twitter/X, using public figures--identified through the pre-2022 verification system--as leaders, and known affiliations of political actors as ground truth labels. Compared with standard baselines, the proposed approach yields more coherent and interpretable communities aligned with political structures, with the two variants respectively recovering parties and coalitions. Activity-based criteria for selecting leaders produce qualitatively similar but consistently weaker results, particularly at the coalition level. Overall, our findings show that creating statistically validated networks of publicly recognized figures, whose off-platform roles constrain and stabilize their online behavior, provide a strong basis to identify discursive communities on social media. Although developed for Twitter/X, the approach is conceptually general, as it leverages structural asymmetries common to many online platforms.

cs.SI↗

Impact of behavioral heterogeneity on epidemic outcome and its mapping into effective network topologies

Human behavior plays a critical role in shaping epidemic trajectories. During health crises, people respond in diverse ways in terms of self-protection and adherence to recommended measures, largely reflecting differences in how individuals assess risk. This behavioral variability induces effective heterogeneity into key epidemic parameters, such as infectivity and susceptibility. We introduce a minimal extension of the susceptible-infected-removed~(SIR) model, denoted HeSIR, that captures these effects through a simple bimodal scheme, where individuals may have higher or lower transmission--related traits. We derive a closed-form expression for the epidemic threshold in terms of the model parameters, and the network's degree distribution and homophily, defined as the tendency of like--risk individuals to preferentially interact. We identify a resurgence regime just beyond the classical threshold, where the number of infected individuals may initially decline before surging into large-scale transmission. Through simulations on homogeneous and heterogeneous network topologies we corroborate the analytical results and highlight how variations in susceptibility and infectivity influence the epidemic dynamics. We further show that, under suitable assumptions, the HeSIR model maps onto a standard SIR process on an appropriately modified contact network, providing a unified interpretation in terms of structural connectivity. Our findings quantify the effect of heterogeneous behavioral responses, especially in the presence of homophily, and caution against underestimating epidemic potential in fragmented populations, which may undermine timely containment efforts. The results also extend to heterogeneity arising from biological or other non-behavioral sources.

physics.soc-ph↗

Polarization and echo chambers in Reddit's political discourse

Political debate nowadays takes place mainly on online social media, with election periods amplifying ideological engagement. Reddit is generally considered more resistant to polarization and echo chamber effects than platforms like Twitter or Facebook. Here, we challenge this assumption through a case study across the 2016 US presidential election. We use statistical validation techniques to extract ideologically distinct communities of subreddits, in terms of their contributing user base and news consumption, which we use to analyze the dynamics of political debate. We thus reveal clear polarization in both interaction-based and topic-based communities, with clusters of Democratic, Conservative, and Banned subreddits. Election periods intensify cross-group engagement, align Banned and Conservative content, and reduce linguistic diversity within groups. Overall we characterize Reddit as a polarized environment marked by the presence of echo chambers, highlighting network validation as a key method for identifying behavioral and interaction patterns on online social media.

physics.soc-ph↗

The Physics of News, Rumors, and Opinions

The boundaries between physical and social networks have narrowed with the advent of the Internet and its pervasive platforms. This has given rise to a complex adaptive information ecosystem where individuals and machines compete for attention, leading to emergent collective phenomena. The flow of information in this ecosystem is often non-trivial and involves complex user strategies from the forging or strategic amplification of manipulative content to large-scale coordinated behavior that trigger misinformation cascades, echo-chamber reinforcement, and opinion polarization. We argue that statistical physics provides a suitable and necessary framework for analyzing the unfolding of these complex dynamics on socio-technological systems. This review systematically covers the foundational and applied aspects of this framework. The review is structured to first establish the theoretical foundation for analyzing these complex systems, examining both structural models of complex networks and physical models of social dynamics (e.g., epidemic and spin models). We then ground these concepts by describing the modern media ecosystem where these dynamics currently unfold, including a comparative analysis of platforms and the challenge of information disorders. The central sections proceed to apply this framework to two central phenomena: first, by analyzing the collective dynamics of information spreading, with a dedicated focus on the models, the main empirical insights, and the unique traits characterizing misinformation; and second, by reviewing current models of opinion dynamics, spanning discrete, continuous, and coevolutionary approaches. In summary, we review both empirical findings based on massive data analytics and theoretical advances, highlighting the valuable insights obtained from physics-based efforts to investigate these phenomena of high societal impact.

physics.soc-ph↗

Language bubbles in online social networks

Social media platforms have become essential spaces for public discourse. While political polarisation and limited communication across different groups are widely acknowledged, the connection between social network fragmentation and the language features and quality used by various communities has received insufficient attention. This study aims to fill this gap by examining the social structure and linguistic richness of the Italian debate on Twitter/X. We analyse tweets and retweets from Italian politicians and news outlets between 2018 and 2022, characterising the retweet network and evaluating the language used within different communities through various lexical metrics. Our analysis uncovers two systematic patterns: communities closer in the network tend to use more similar vocabulary, while isolated communities consistently demonstrate lower lexical diversity and richness. Together, these patterns illustrate what we call ``language bubbles''. These findings indicate that socially isolated communities interact less with others and develop distinct and poorer linguistic profiles, highlighting a structural link between social fragmentation and linguistic divergence.

physics.soc-ph↗

Entropy-based models to randomize real-world hypergraphs

Network theory has often disregarded many-body relationships, solely focusing on pairwise interactions: neglecting them, however, can lead to misleading representations of complex systems. Hypergraphs represent a suitable framework for describing polyadic interactions. Here, we leverage the representation of hypergraphs based on the incidence matrix for extending the entropy-based approach to higher-order structures: in analogy with the Exponential Random Graphs, we introduce the Exponential Random Hypergraphs (ERHs). After exploring the asymptotic behaviour of thresholds generalising the percolation one, we apply ERHs to study real-world data. First, we generalise key network metrics to hypergraphs; then, we compute their expected value and compare it with the empirical one, in order to detect deviations from random behaviours. Our method is analytically tractable, scalable and capable of revealing structural patterns of real-world hypergraphs that differ significantly from those emerging as a consequence of simpler constraints.

cs.SI↗

Statistically validated projection of bipartite signed networks

Bipartite networks provide a major insight into the organisation of many real-world systems. One of the most relevant issues encountered when modelling a bipartite network is that of facing the information shortage concerning intra-layer linkages. In the present contribution, we propose an unsupervised algorithm to obtain statistically validated projections of bipartite signed networks, according to which any two nodes sharing a statistically significant number of concordant (discordant) relationships are connected by a positive (negative) edge. Our algorithm outputs a matrix of link-specific $p-$values, from which a validated projection can be obtained upon running a multiple-hypothesis testing procedure. After testing our method on synthetic configurations output by a fully controllable generative model, we apply it to several real-world configurations: in all cases, non-trivial mesoscopic structures, induced by relationships that cannot be traced back to the constraints defining the employed benchmarks, hence revealing genuine traces of self-organisation, are detected.

physics.soc-ph↗

Maximum entropy modeling of Optimal Transport: the sub-optimality regime and the transition from dense to sparse networks

We present a bipartite network model that captures intermediate stages of optimization by blending the Maximum Entropy approach with Optimal Transport. In this framework, the network's constraints define the total mass each node can supply or receive, while an external cost field favors a minimal set of links, driving the system toward a sparse, tree-like structure. By tuning the control parameter, one transitions from uniformly distributed weights to an optimal transport regime in which weights condense onto cost-favorable edges. We quantify this dense-to-sparse transition, showing with numerical analyses that the process does not hinge on specific assumptions about the node-strength or cost distributions. Finite-size analysis confirms that the results persist in the thermodynamic limit. Because the model offers explicit control over the degree of sub-optimality, this approach lends to practical applications in link prediction, network reconstruction, and statistical validation, particularly in systems where partial optimization coexists with other noise-like factors.

cond-mat.stat-mech↗

TROPIC - Trustworthiness Rating of Online Publishers through online Interactions Calculation

Existing methods for assessing the trustworthiness of news publishers face high costs and scalability issues. The tool presented in this paper supports the efforts of specialized organizations by providing a solution that, starting from an online discussion, provides (i) trustworthiness ratings for previously unclassified news publishers and (ii) an interactive platform to guide annotation efforts and improve the robustness of the ratings. The system implements a novel framework for assessing the trustworthiness of online news publishers based on user interactions on social media platforms.

cs.SI↗

Patterns of link reciprocity in directed, signed networks

Most of the analyses concerning signed networks have focused on the balance theory, hence identifying frustration with undirected, triadic motifs having an odd number of negative edges; much less attention has been paid to their directed counterparts. To fill this gap, we focus on signed, directed connections, with the aim of exploring the notion of frustration in such a context. When dealing with signed, directed edges, frustration is a multi-faceted concept, admitting different definitions at different scales: if we limit ourselves to consider cycles of length two, frustration is related to reciprocity, i.e. the tendency of edges to admit the presence of partners pointing in the opposite direction. As the reciprocity of signed networks is still poorly understood, we adopt a principled approach for its study, defining quantities and introducing models to consistently capture empirical patterns of the kind. In order to quantify the tendency of empirical networks to form either mutualistic or antagonistic cycles of length two, we extend the Exponential Random Graphs framework to binary, directed, signed networks with global and local constraints and, then, compare the empirical abundance of the aforementioned patterns with the one expected under each model. We find that the (directed extension of the) balance theory is not capable of providing a consistent explanation of the patterns characterising the directed, signed networks considered in this work. Although part of the ambiguities can be solved by adopting a coarser definition of balance, our results call for a different theory, accounting for the directionality of edges in a coherent manner. In any case, the evidence that the empirical, signed networks can be highly reciprocated leads us to recommend to explicitly account for the role played by bidirectional dyads in determining frustration at higher levels (e.g. the triadic one).

physics.soc-ph↗

Testing structural balance theories in heterogeneous signed networks

The abundance of data about social relationships allows the human behavior to be analyzed as any other natural phenomenon. Here we focus on balance theory, stating that social actors tend to avoid establishing cycles with an odd number of negative links. This statement, however, can be supported only after a comparison with a benchmark. Since the existing ones disregard actors' heterogeneity, we extend Exponential Random Graphs to signed networks with both global and local constraints and employ them to assess the significance of empirical unbalanced patterns. We find that the nature of balance crucially depends on the null model: while homogeneous benchmarks favor the weak balance theory, according to which only triangles with one negative link should be under-represented, heterogeneous benchmarks favor the strong balance theory, according to which also triangles with all negative links should be under-represented. Biological networks, instead, display strong frustration under any benchmark, confirming that structural balance inherently characterizes social networks.

physics.soc-ph↗

Online disinformation in the 2020 U.S. Election: swing vs. safe states

For U.S. presidential elections, most states use the so-called winner-take-all system, in which the state's presidential electors are awarded to the winning political party in the state after a popular vote phase, regardless of the actual margin of victory. Therefore, election campaigns are especially intense in states where there is no clear direction on which party will be the winning party. These states are often referred to as swing states. To measure the impact of such an election law on the campaigns, we analyze the Twitter activity surrounding the 2020 US preelection debate, with a particular focus on the spread of disinformation. We find that about 88% of the online traffic was associated with swing states. In addition, the sharing of links to unreliable news sources is significantly more prevalent in tweets associated with swing states: in this case, untrustworthy tweets are predominantly generated by automated accounts. Furthermore, we observe that the debate is mostly led by two main communities, one with a predominantly Republican affiliation and the other with accounts of different political orientations. Most of the disinformation comes from the former.

cs.SI↗

Unveiling News Publishers Trustworthiness Through Social Interactions

With the primary goal of raising readers' awareness of misinformation phenomena, extensive efforts have been made by both academic institutions and independent organizations to develop methodologies for assessing the trustworthiness of online news publishers. Unfortunately, existing approaches are costly and face critical scalability challenges. This study presents a novel framework for assessing the trustworthiness of online news publishers using user interactions on social media platforms. The proposed methodology provides a versatile solution that serves the dual purpose of i) identifying verifiable online publishers and ii) automatically performing an initial estimation of the trustworthiness of previously unclassified online news outlets.

cs.SI↗

Entropy-based detection of Twitter echo chambers

Echo chambers, i.e. clusters of users exposed to news and opinions in line with their previous beliefs, were observed in many online debates on social platforms. We propose a completely unbiased entropy-based method for detecting echo chambers. The method is completely agnostic to the nature of the data. In the Italian Twitter debate about the Covid-19 vaccination, we find a limited presence of users in echo chambers (about 0.35% of all users). Nevertheless, their impact on the formation of a common discourse is strong, as users in echo chambers are responsible for nearly a third of the retweets in the original dataset. Moreover, in the case study observed, echo chambers appear to be a receptacle for disinformative content.

cs.SI↗

Pattern detection in bipartite networks: a review of terminology, applications and methods

Two dimensional matrices with binary (0/1) entries are a common data structure in many research fields. Examples include ecology, economics, mathematics, physics, psychometrics and others. Because the columns and rows of these matrices represent distinct entities, they can equivalently be expressed as a pair of bipartite networks that are linked by projection. A variety of diversity statistics and network metrics can then be used to quantify patterns in these matrices and networks. But what should these patterns be compared to? In all of these disciplines, researchers have recognized the necessity of comparing an empirical matrix to a benchmark set of "null" matrices created by randomizing certain elements of the original data. This common need has nevertheless promoted the independent development of methodologies by researchers who come from different backgrounds and use different terminology. Here, we provide a multidisciplinary review of randomization techniques for matrices representing binary, bipartite networks. We aim to translate the concepts from different technical domains into a common language that is accessible to a broad scientific audience. Specifically, after briefly reviewing examples of binary matrix structures across different fields, we introduce the major approaches and common strategies for randomizing these matrices. We then explore the details of and performance of specific techniques, and discuss their limitations and computational challenges. In particular, we focus on the conceptual importance and implementation of structural constraints on the randomization, such as preserving row or columns sums of the original matrix in each of the randomized matrices. Our review serves both as a guide for empiricists in different disciplines, as well as a reference point for researchers working on theoretical and methodological developments in matrix randomization methods.

physics.data-an↗

Sustainable Development Goals as unifying narratives in large UK firms' Twitter discussions

To achieve sustainable development worldwide, the United Nations set 17 Sustainable Development Goals (SDGs) for humanity to reach by 2030. Society is involved in the challenge, with firms playing a crucial role. Thus, a key question is to what extent firms engage with the SDGs. Efforts to map firms' contributions have mainly focused on analysing companies' reports based on limited samples and non-real-time data. We present a novel interdisciplinary approach based on analysing big data from an online social network (Twitter) with complex network methods from statistical physics. By doing so, we provide a comprehensive and nearly real-time picture of firms' engagement with SDGs. Results show that: 1) SDGs themes tie conversations among major UK firms together; 2) the social dimension is predominant; 3) the attention to different SDGs themes varies depending on the community and sector firms belong to; 4) stakeholder engagement is higher on posts related to global challenges compared to general ones; 5) large UK companies and stakeholders generally behave differently from Italian ones. This paper provides theoretical contributions and practical implications relevant to firms, policymakers and management education. Most importantly, it provides a novel tool and a set of keywords to monitor the influence of the private sector on the implementation of the 2030 Agenda.

physics.soc-ph↗

Inferring comparative advantage via entropy maximization

We revise the procedure proposed by Balassa to infer comparative advantage, which is a standard tool, in Economics, to analyze specialization (of countries, regions, etc.). Balassa's approach compares the export of a product for each country with what would be expected from a benchmark based on the total volumes of countries and products flows. Based on results in the literature, we show that the implementation of Balassa's idea generates a bias: the prescription of the maximum likelihood used to calculate the parameters of the benchmark model conflicts with the model's definition. Moreover, Balassa's approach does not implement any statistical validation. Hence, we propose an alternative procedure to overcome such a limitation, based upon the framework of entropy maximisation and implementing a proper test of hypothesis: the `key products' of a country are, now, the ones whose production is significantly larger than expected, under a null-model constraining the same amount of information employed by Balassa's approach. What we found is that countries diversification is always observed, regardless of the strictness of the validation procedure. Besides, the ranking of countries' fitness is only partially affected by the details of the validation scheme employed for the analysis while large differences are found to affect the rankings of products Complexities. The routine for implementing the entropy-based filtering procedures employed here is freely available through the official Python Package Index PyPI.

cs.SI↗