SearcharxivSearch

arXiv subjects

Pietro Panzarasa

Publications and source records attributed to Pietro Panzarasa.

18 recordsLinked to original sources

The Promise of Large Language Models in Digital Health: Evidence from Sentiment Analysis in Online Health Communities

Digital health analytics face critical challenges nowadays. The sophisticated analysis of patient-generated health content, which contains complex emotional and medical contexts, requires scarce domain expertise, while traditional ML approaches are constrained by data shortage and privacy limitations in healthcare settings. Online Health Communities (OHCs) exemplify these challenges with mixed-sentiment posts, clinical terminology, and implicit emotional expressions that demand specialised knowledge for accurate Sentiment Analysis (SA). To address these challenges, this study explores how Large Language Models (LLMs) can integrate expert knowledge through in-context learning for SA, providing a scalable solution for sophisticated health data analysis. Specifically, we develop a structured codebook that systematically encodes expert interpretation guidelines, enabling LLMs to apply domain-specific knowledge through targeted prompting rather than extensive training. Six GPT models validated alongside DeepSeek and LLaMA 3.1 are compared with pre-trained language models (BioBERT variants) and lexicon-based methods, using 400 expert-annotated posts from two OHCs. LLMs achieve superior performance while demonstrating expert-level agreement. This high agreement, with no statistically significant difference from inter-expert agreement levels, suggests knowledge integration beyond surface-level pattern recognition. The consistent performance across diverse LLM models, supported by in-context learning, offers a promising solution for digital health analytics. This approach addresses the critical challenge of expert knowledge shortage in digital health research, enabling real-time, expert-quality analysis for patient monitoring, intervention assessment, and evidence-based health strategies.

cs.CL

The rise and fall of countries in the global value chains

Countries participate in global value chains by engaging in backward and forward transactions connecting multiple geographically dispersed production stages. Inspired by network theory, we model global trade as a multilayer network and study its power structure by investigating the tendency of eigenvector centrality to concentrate on a small fraction of countries, a phenomenon called localization transition. We show that the market underwent a significant structural variation in 2007 just before the global financial crisis. That year witnessed an abrupt repositioning of countries in the global value chains, and in particular a remarkable reversal of leading role between the two major economies, the US and China. We uncover the hierarchical structure of the multilayer network based on countries' time series of eigenvector centralities, and show that trade tends to concentrate between countries with different power dynamics, yet in geographical proximity. We further investigate the contribution of individual industries to countries' economic dominance, and show that also within industry variations in countries' market positioning took place in 2007. Moreover, we shed light on the crucial role that domestic trade played in the geopolitical landscape leading China to overtake the US and cement its status as leading economy of the global value chains. Our study shows how the 2008 crisis can offer insights to policy makers and governments on how to turn early structural signals of upcoming exogenous shocks into opportunities for redesigning countries' global roles in a changing geopolitical landscape.

physics.soc-ph

What does Network Analysis teach us about International Environmental Cooperation?

Over the past 70 years, the number of international environmental agreements (IEAs) has increased substantially, highlighting their prominent role in environmental governance. This paper applies the toolkit of network analysis to identify the network properties of international environmental cooperation based on 546 IEAs signed between 1948 and 2015. We identify four stylised facts that offer topological corroboration for some key themes in the IEA literature. First, we find that a statistically significant cooperation network did not emerge until early 1970, but since then the network has grown continuously in strength, resulting in higher connectivity and intensity of cooperation between signatory countries. Second, over time the network has become closer, denser and more cohesive, allowing more effective policy coordination and knowledge diffusion. Third, the network, while global, has a noticeable European imprint: initially the United Kingdom and more recently France and Germany have been the most strategic players to broker environmental cooperation. Fourth, international environmental coordination started with the management of fisheries and the sea, but is now most intense on waste and hazardous substances. The network of air and atmosphere treaties is weaker on a number of metrics and lacks the hierarchical structure found in other networks. It is the only network whose topological properties are shaped significantly by UN-sponsored treaties.

econ.GN

Geometric graphs from data to aid classification tasks with graph convolutional networks

Traditional classification tasks learn to assign samples to given classes based solely on sample features. This paradigm is evolving to include other sources of information, such as known relations between samples. Here we show that, even if additional relational information is not available in the data set, one can improve classification by constructing geometric graphs from the features themselves, and using them within a Graph Convolutional Network. The improvement in classification accuracy is maximized by graphs that capture sample similarity with relatively low edge density. We show that such feature-derived graphs increase the alignment of the data to the ground truth while improving class separation. We also demonstrate that the graphs can be made more efficient using spectral sparsification, which reduces the number of edges while still improving classification performance. We illustrate our findings using synthetic and real-world data sets from various scientific domains.

cs.LG

Quantifying the Alignment of Graph and Features in Deep Learning

We show that the classification performance of graph convolutional networks (GCNs) is related to the alignment between features, graph, and ground truth, which we quantify using a subspace alignment measure (SAM) corresponding to the Frobenius norm of the matrix of pairwise chordal distances between three subspaces associated with features, graph, and ground truth. The proposed measure is based on the principal angles between subspaces and has both spectral and geometrical interpretations. We showcase the relationship between the SAM and the classification performance through the study of limiting cases of GCNs and systematic randomizations of both features and graph structure applied to a constructive example and several examples of citation networks of different origins. The analysis also reveals the relative importance of the graph and features for classification purposes.

cs.LG

Early warnings of COVID-19 outbreaks across Europe from social media?

We analyze data from Twitter to uncover early-warning signals of COVID-19 outbreaks in Europe in the winter season 2019-2020, before the first public announcements of local sources of infection were made. We show evidence that unexpected levels of concerns about cases of pneumonia were raised across a number of European countries. Whistleblowing came primarily from the geographical regions that eventually turned out to be the key breeding grounds for infections. These findings point to the urgency of setting up an integrated digital surveillance system in which social media can help geo-localize chains of contagion that would otherwise proliferate almost completely undetected.

econ.GN

A network approach to expertise retrieval based on path similarity and credit allocation

With the increasing availability of online scholarly databases, publication records can be easily extracted and analysed. Researchers can promptly keep abreast of others' scientific production and, in principle, can select new collaborators and build new research teams. A critical factor one should consider when contemplating new potential collaborations is the possibility of unambiguously defining the expertise of other researchers. While some organisations have established database systems to enable their members to manually produce a profile, maintaining such systems is time-consuming and costly. Therefore, there has been a growing interest in retrieving expertise through automated approaches. Indeed, the identification of researchers' expertise is of great value in many applications, such as identifying qualified experts to supervise new researchers, assigning manuscripts to reviewers, and forming a qualified team. Here, we propose a network-based approach to the construction of authors' expertise profiles. Using the MEDLINE corpus as an example, we show that our method can be applied to a number of widely used data sets and outperforms other methods traditionally used for expertise identification.

cs.SI

The nested structural organization of the worldwide trade multi-layer network

Nestedness has traditionally been used to detect assembly patterns in meta-communities and networks of interacting species. Attempts have also been made to uncover nested structures in international trade, typically represented as bipartite networks in which connections can be established between countries (exporters or importers) and industries. A bipartite representation of trade, however, inevitably neglects transactions between industries. To fully capture the organization of the global value chain, we draw on the World Input-Output Database and construct a multi-layer network in which the nodes are the countries, the layers are the industries, and links can be established from sellers to buyers within and across industries. We define the buyers' and sellers' participation matrices in which the rows are the countries and the columns are all possible pairs of industries, and then compute nestedness based on buyers' and sellers' involvement in transactions between and within industries. Drawing on appropriate null models that preserve the countries' or layers' degree distributions in the original multi-layer network, we uncover variations of country- and transaction-based nestedness over time, and identify the countries and industries that most contributed to nestedness. We discuss the implications of our findings for the study of the international production network and other real-world systems.

physics.soc-ph

Predicting success in the worldwide start-up network

By drawing on large-scale online data we construct and analyze the time-varying worldwide network of professional relationships among start-ups. The nodes of this network represent companies, while the links model the flow of employees and the associated transfer of know-how across companies. We use network centrality measures to assess, at an early stage, the likelihood of the long-term positive performance of a start-up, showing that the start-up network has predictive power and provides valuable recommendations doubling the current state of the art performance of venture funds. Our network-based approach not only offers an effective alternative to the labour-intensive screening processes of venture capital firms, but can also enable entrepreneurs and policy-makers to conduct a more objective assessment of the long-term potentials of innovation ecosystems and to target interventions accordingly.

physics.soc-ph

Unfolding the complexity of the global value chain: Strengths and entropy in the single-layer, multiplex, and multi-layer international trade networks

The worldwide trade network has been widely studied through different data sets and network representations with a view to better understanding interactions among countries and products. Here we investigate international trade through the lenses of the single-layer, multiplex, and multi-layer networks. We discuss differences among the three network frameworks in terms of their relative advantages in capturing salient topological features of trade. We draw on the World Input-Output Database to build the three networks. We then uncover sources of heterogeneity in the way strength is allocated among countries and transactions by computing the strength distribution and entropy in each network. Additionally, we trace how entropy evolved, and show how the observed peaks can be associated with the onset of the global economic downturn. Findings suggest how more complex representations of trade, such as the multi-layer network, enable us to disambiguate the distinct roles of intra- and cross-industry transactions in driving the evolution of entropy at a more aggregate level. We discuss our results and the implications of our comparative analysis of networks for research on international trade and other empirical domains across the natural and social sciences.

physics.soc-ph

The advantages of interdisciplinarity in modern science

As the increasing complexity of large-scale research requires the combined efforts of scientists with expertise in different fields, the advantages and costs of interdisciplinary scholarship have taken center stage in current debates on scientific production. Here we conduct a comparative assessment of the scientific success of specialized and interdisciplinary researchers in modern science. Drawing on comprehensive data sets on scientific production, we propose a two-pronged approach to interdisciplinarity. For each scientist, we distinguish between background interdisciplinarity, rooted in knowledge accumulated over time, and social interdisciplinarity, stemming from exposure to collaborators' knowledge. We find that, while abandoning specialization in favor of moderate degrees of background interdisciplinarity deteriorates performance, very interdisciplinary scientists outperform specialized ones, at all career stages. Moreover, successful scientists tend to intensify the heterogeneity of collaborators and to match the diversity of their network with the diversity of their background. Collaboration sustains performance by facilitating knowledge diffusion, acquisition and creation. Successful scientists tend to absorb a larger fraction of their collaborators' knowledge, and at a faster pace, than less successful ones. Collaboration also provides successful scientists with opportunities for the cross-fertilization of ideas and the synergistic creation of new knowledge. These results can inspire scientists to shape successful careers, research institutions to develop effective recruitment policies, and funding agencies to award grants of enhanced impact.

physics.soc-ph

Homophily and missing links in citation networks

Citation networks have been widely used to study the evolution of science through the lenses of the underlying patterns of knowledge flows among academic papers, authors, research sub-fields, and scientific journals. Here we focus on citation networks to cast light on the salience of homophily, namely the principle that similarity breeds connection, for knowledge transfer between papers. To this end, we assess the degree to which citations tend to occur between papers that are concerned with seemingly related topics or research problems. Drawing on a large data set of articles published in the journals of the American Physical Society between 1893 and 2009, we propose a novel method for measuring the similarity between articles through the statistical validation of the overlap between their bibliographies. Results suggest that the probability of a citation made by one article to another is indeed an increasing function of the similarity between the two articles. Our study also enables us to uncover missing citations between pairs of highly related articles, and may thus help identify barriers to effective knowledge flows. By quantifying the proportion of missing citations, we conduct a comparative assessment of distinct journals and research sub-fields in terms of their ability to facilitate or impede the dissemination of knowledge. Findings indicate that knowledge transfer seems to be more effectively facilitated by journals of wide visibility, such as Physical Review Letters, than by lower-impact ones. Our study has important implications for authors, editors and reviewers of scientific journals, as well as public preprint repositories, as it provides a procedure for recommending relevant yet missing references and properly integrating bibliographies of papers.

physics.soc-ph

Degree correlations in signed social networks

We investigate degree correlations in two online social networks where users are connected through different types of links. We find that, while subnetworks in which links have a positive connotation, such as endorsement and trust, are characterized by assortative mixing by degree, networks in which links have a negative connotation, such as disapproval and distrust, are characterized by disassortative patterns. We introduce a class of simple theoretical models to analyze the interplay between network topology and the superimposed structure based on the sign of links. Results uncover the conditions that underpin the emergence of the patterns observed in the data, namely the assortativity of positive subnetworks and the disassortativity of negative ones. We discuss the implications of our study for the analysis of signed complex networks.

physics.soc-ph

A Unifying Framework for Measuring Weighted Rich Clubs

Network analysis can help uncover meaningful regularities in the organization of complex systems. Among these, rich clubs are a functionally important property of a variety of social, technological and biological networks. Rich clubs emerge when nodes that are somehow prominent or 'rich' (e.g., highly connected) interact preferentially with one another. The identification of rich clubs is non-trivial, especially in weighted networks, and to this end multiple distinct metrics have been proposed. Here we describe a unifying framework for detecting rich clubs which intuitively generalizes various metrics into a single integrated method. This generalization rests upon the explicit incorporation of randomized control networks into the measurement process. We apply this framework to real-life examples, and show that, depending on the selection of randomized controls, different kinds of rich-club structures can be detected, such as topological and weighted rich clubs.

physics.soc-ph

Weighted Multiplex Networks

One of the most important challenges in network science is to quantify the information encoded in complex network structures. Disentangling randomness from organizational principles is even more demanding when networks have a multiplex nature. Multiplex networks are multilayer systems of $N$ nodes that can be linked in multiple interacting and co-evolving layers. In these networks, relevant information might not be captured if the single layers were analyzed separately. Here we demonstrate that such partial analysis of layers fails to capture significant correlations between weights and topology of complex multiplex networks. To this end, we study two weighted multiplex co-authorship and citation networks involving the authors included in the American Physical Society. We show that in these networks weights are strongly correlated with multiplex structure, and provide empirical evidence in favor of the advantage of studying weighted measures of multiplex networks, such as multistrength and the inverse multiparticipation ratio. Finally, we introduce a theoretical framework based on the entropy of multiplex ensembles to quantify the information stored in multiplex networks that would remain undetected if the single layers were analyzed in isolation.

physics.soc-ph

Multiplex PageRank

Many complex systems can be described as multiplex networks in which the same nodes can interact with one another in different layers, thus forming a set of interacting and co-evolving networks. Examples of such multiplex systems are social networks where people are involved in different types of relationships and interact through various forms of communication media. The ranking of nodes in multiplex networks is one of the most pressing and challenging tasks that research on complex networks is currently facing. When pairs of nodes can be connected through multiple links and in multiple layers, the ranking of nodes should necessarily reflect the importance of nodes in one layer as well as their importance in other interdependent layers. In this paper, we draw on the idea of biased random walks to define the Multiplex PageRank centrality measure in which the effects of the interplay between networks on the centrality of nodes are directly taken into account. In particular, depending on the intensity of the interaction between layers, we define the Additive, Multiplicative, Combined, and Neutral versions of Multiplex PageRank, and show how each version reflects the extent to which the importance of a node in one layer affects the importance the node can gain in another layer. We discuss these measures and apply them to an online multiplex social network. Findings indicate that taking the multiplex nature of the network into account helps uncover the emergence of rankings of nodes that differ from the rankings obtained from one single layer. Results provide support in favor of the salience of multiplex centrality measures, like Multiplex PageRank, for assessing the prominence of nodes embedded in multiple interacting networks, and for shedding a new light on structural properties that would otherwise remain undetected if each of the interacting networks were analyzed in isolation.

physics.soc-ph

Social cohesion, structural holes, and a tale of two measures

In the social sciences, the debate over the structural foundations of social capital has long vacillated between two positions on the relative benefits associated with two types of social structures: closed structures, rich in third-party relationships, and open structures, rich in structural holes and brokerage opportunities. In this paper, we engage with this debate by focusing on the measures typically used for formalising the two conceptions of social capital: clustering and effective size. We show that these two measures are simply two sides of the same coin, as they can be expressed one in terms of the other through a simple functional relation. Building on this relation, we then attempt to reconcile closed and open structures by proposing a new measure, Simmelian brokerage, that captures opportunities of brokerage between otherwise disconnected cohesive groups of contacts. Implications of our findings for research on social capital and complex networks are discussed.

physics.soc-ph

Prominence and control: The weighted rich-club effect

Complex systems are often characterized by large-scale hierarchical organizations. Whether the prominent elements, at the top of the hierarchy, share and control resources or avoid one another lies at the heart of a system's global organization and functioning. Inspired by network perspectives, we propose a new general framework for studying the tendency of prominent elements to form clubs with exclusive control over the majority of a system's resources. We explore associations between prominence and control in the fields of transportation, scientific collaboration, and online communication.

physics.soc-ph