SearcharxivSearch

arXiv subjects

Samuel V. Scarpino

Publications and source records attributed to Samuel V. Scarpino.

At least 19 recordsLinked to original sources

Hypergraph Representations of scRNA-seq Data for Improved Clustering with Random Walks

Analysis of single-cell RNA sequencing data is often conducted through network projections such as coexpression networks, primarily due to the abundant availability of network analysis tools for downstream tasks. However, this approach has several limitations: loss of higher-order information, inefficient data representation caused by converting a sparse dataset to a fully connected network, and overestimation of coexpression due to zero-inflation. To address these limitations, we propose conceptualizing scRNA-seq expression data as hypergraphs, which are generalized graphs in which the hyperedges can connect more than two vertices. In the context of scRNA-seq data, the hypergraph nodes represent cells and the edges represent genes. Each hyperedge connects all cells where its corresponding gene is actively expressed and records the expression of the gene across different cells. This hypergraph conceptualization enables us to explore multi-way relationships beyond the pairwise interactions in coexpression networks without loss of information. We propose two novel clustering methods: (1) the Dual-Importance Preference Hypergraph Walk (DIPHW) and (2) the Coexpression and Memory-Integrated Dual-Importance Preference Hypergraph Walk (CoMem-DIPHW). They outperform established methods on both simulated and real scRNA-seq datasets. The improvement brought by our proposed methods is especially significant when data modularity is weak. Furthermore, CoMem-DIPHW incorporates the gene coexpression network, cell coexpression network, and the cell-gene expression hypergraph from the single-cell abundance counts data altogether for embedding computation. This approach accounts for both the local level information from single-cell level gene expression and the global level information from the pairwise similarity in the two coexpression networks.

q-bio.QM

Deterministic construction of typical networks in network models

It is often desirable to assess how well a given dataset is described by a given model. In network science, for instance, one often wants to say that a given real-world network appears to come from a particular network model. In statistical physics, the corresponding problem is about how typical a given state, representing real-world data, is in a particular statistical ensemble. One way to address this problem is to measure the distance between the data and the most typical state in the ensemble. Here, we identify the conditions that allow us to define this most typical state. These conditions hold in a wide class of grand canonical ensembles and their random mixtures. Our main contribution is a deterministic construction of a state that converges to this most typical state in the thermodynamic limit. This construction involves rounds of derandomization procedures, some of which deal with derandomizing point processes, an uncharted territory. We illustrate the construction on one particular network model, deterministic hyperbolic graphs, and its application to real-world networks, many of which we find are close to the most typical network in the model. While our main focus is on network models, our results are very general and apply to any grand canonical ensembles and their random mixtures satisfying certain niceness requirements.

physics.soc-ph

One pathogen does not an epidemic make: A review of interacting contagions, diseases, beliefs, and stories

From pathogens and computer viruses to genes and memes, contagion models have found widespread utility across the natural and social sciences. Despite their success and breadth of adoption, the approach and structure of these models remain surprisingly siloed by field. Given the siloed nature of their development and widespread use, one persistent assumption is that a given contagion can be studied in isolation, independently from what else might be spreading in the population. In reality, countless contagions of biological and social nature interact within hosts (interacting with existing beliefs, or the immune system) and across hosts (interacting in the environment, or affecting transmission mechanisms). Additionally, from a modeling perspective, we know that relaxing these assumptions has profound effects on the physics and translational implications of the models. Here, we review mechanisms for interactions in social and biological contagions, as well as the models and frameworks developed to include these interactions in the study of the contagions. We highlight existing problems related to the inference of interactions and to the scalability of mathematical models and identify promising avenues of future inquiries. In doing so, we highlight the need for interdisciplinary efforts under a unified science of contagions and for removing a common dichotomy between social and biological contagions.

physics.soc-ph

Structural causal influence (SCI) captures the forces of social inequality in models of disease dynamics

Mathematical modeling has played a central role in understanding how infectious disease transmission manifests in populations. These models have demonstrated the importance of key community-level factors in structuring epidemic risk, and are now routinely used in public health for decision support. One barrier to their broader utility is that the existing canon does not often accommodate social inequalities as distinct formal drivers of variability in transmission dynamics. Given decades of evidence supporting the organizational effects of inequalities in structuring society more generally, and infectious disease risk more specifically, addressing this modeling gap is of critical importance. In this study, we build on previous efforts to integrate social forces into computational epidemiology by introducing a metric, the structural causal influence (SCI). The SCI uses causal analysis to provide a measure of the relative vulnerability of sub-communities within a susceptible population, shaped by differences in characteristics such as access to therapy, exposure to disease, and other determinants driven by social forces. We develop our metric in a simple case and apply it to a context of public health importance: Hepatitis C virus in a population of persons who inject drugs. In addition, we demonstrate the flexibility of the SCI using an agent-based model of an infectious disease. Our use of the SCI reveals that, under specific parameters in a multi-community model, the "less vulnerable" community may achieve a basic reproduction number below one, ensuring disease extinction. However, even minimal transmission between communities can increase this number, leading to sustained epidemics within both communities.

q-bio.QM

A Misclassification Network-Based Method for Comparative Genomic Analysis

Classifying genome sequences based on metadata has been an active area of research in comparative genomics for decades with many important applications across the life sciences. Established methods for classifying genomes can be broadly grouped into sequence alignment-based and alignment-free models. Conventional alignment-based models rely on genome similarity measures calculated based on local sequence alignments or consistent ordering among sequences. However, such methods are computationally expensive when dealing with large ensembles of even moderately sized genomes. In contrast, alignment-free (AF) approaches measure genome similarity based on summary statistics in an unsupervised setting and are efficient enough to analyze large datasets. However, both alignment-based and AF methods typically assume fixed scoring rubrics that lack the flexibility to assign varying importance to different parts of the sequences based on prior knowledge. In this study, we integrate AI and network science approaches to develop a comparative genomic analysis framework that addresses these limitations. Our approach, termed the Genome Misclassification Network Analysis (GMNA), simultaneously leverages misclassified instances, a learned scoring rubric, and label information to classify genomes based on associated metadata and better understand potential drivers of misclassification. We evaluate the utility of the GMNA using Naive Bayes and convolutional neural network models, supplemented by additional experiments with transformer-based models, to construct SARS-CoV-2 sampling location classifiers using over 500,000 viral genome sequences and study the resulting network of misclassifications. We demonstrate the global health potential of the GMNA by leveraging the SARS-CoV-2 genome misclassification networks to investigate the role human mobility played in structuring geographic clustering of SARS-CoV-2.

q-bio.GN

Spatial scales of COVID-19 transmission in Mexico

During outbreaks of emerging infectious diseases, internationally connected cities often experience large and early outbreaks, while rural regions follow after some delay. This hierarchical structure of disease spread is influenced primarily by the multiscale structure of human mobility. However, during the COVID-19 epidemic, public health responses typically did not take into consideration the explicit spatial structure of human mobility when designing non-pharmaceutical interventions (NPIs). NPIs were applied primarily at national or regional scales. Here we use weekly anonymized and aggregated human mobility data and spatially highly resolved data on COVID-19 cases, deaths and hospitalizations at the municipality level in Mexico to investigate how behavioural changes in response to the pandemic have altered the spatial scales of transmission and interventions during its first wave (March - June 2020). We find that the epidemic dynamics in Mexico were initially driven by SARS-CoV-2 exports from Mexico State and Mexico City, where early outbreaks occurred. The mobility network shifted after the implementation of interventions in late March 2020, and the mobility network communities became more disjointed while epidemics in these communities became increasingly synchronised. Our results provide actionable and dynamic insights into how to use network science and epidemiological modelling to inform the spatial scale at which interventions are most impactful in mitigating the spread of COVID-19 and infectious diseases in general.

physics.soc-ph

Characterizing collective physical distancing in the U.S. during the first nine months of the COVID-19 pandemic

The COVID-19 pandemic offers an unprecedented natural experiment providing insights into the emergence of collective behavioral changes of both exogenous (government mandated) and endogenous (spontaneous reaction to infection risks) origin. Here, we characterize collective physical distancing -- mobility reductions, minimization of contacts, shortening of contact duration -- in response to the COVID-19 pandemic in the pre-vaccine era by analyzing de-identified, privacy-preserving location data for a panel of over 5.5 million anonymized, opted-in U.S. devices. We define five indicators of users' mobility and proximity to investigate how the emerging collective behavior deviates from the typical pre-pandemic patterns during the first nine months of the COVID-19 pandemic. We analyze both the dramatic changes due to the government mandated mitigation policies and the more spontaneous societal adaptation into a new (physically distanced) normal in the fall 2020. The indicators defined here allow the quantification of behavior changes across the rural/urban divide and highlight the statistical association of mobility and proximity indicators with metrics characterizing the pandemic's social and public health impact such as unemployment and deaths. This study provides a framework to study massive social distancing phenomena with potential uses in analyzing and monitoring the effects of pandemic mitigation plans at the national and international level.

physics.soc-ph

Effective Resistance for Pandemics: Mobility Network Sparsification for High-Fidelity Epidemic Simulation

Network science has increasingly become central to the field of epidemiology and our ability to respond to infectious disease threats. However, many networks derived from modern datasets are not just large, but dense, with a high ratio of edges to nodes. This includes human mobility networks where most locations have a large number of links to many other locations. Simulating large-scale epidemics requires substantial computational resources and in many cases is practically infeasible. One way to reduce the computational cost of simulating epidemics on these networks is sparsification, where a representative subset of edges is selected based on some measure of their importance. We test several sparsification strategies, ranging from naive thresholding to random sampling of edges, on mobility data from the U.S. Following recent work in computer science, we find that the most accurate approach uses the effective resistances of edges, which prioritizes edges that are the only efficient way to travel between their endpoints. The resulting sparse network preserves many aspects of the behavior of an SIR model, including both global quantities, like the epidemic size, and local details of stochastic events, including the probability each node becomes infected and its distribution of arrival times. This holds even when the sparse network preserves fewer than $10\%$ of the edges of the original network. In addition to its practical utility, this method helps illuminate which links of a weighted, undirected network are most important to disease spread.

q-bio.PE

The role of directionality, heterogeneity and correlations in epidemic risk and spread

Most models of epidemic spread, including many designed specifically for COVID-19, implicitly assume mass-action contact patterns and undirected contact networks, meaning that the individuals most likely to spread the disease are also the most at risk to receive it from others. Here, we review results from the theory of random directed graphs which show that many important quantities, including the reproduction number and the epidemic size, depend sensitively on the joint distribution of in- and out-degrees ("risk" and "spread"), including their heterogeneity and the correlation between them. By considering joint distributions of various kinds, we elucidate why some types of heterogeneity cause a deviation from the standard Kermack-McKendrick analysis of SIR models, i.e., so-called mass-action models where contacts are homogeneous and random, and some do not. We also show that some structured SIR models informed by realistic complex contact patterns among types of individuals (age or activity) are simply mixtures of Poisson processes and tend not to deviate significantly from the simplest mass-action model. Finally, we point out some possible policy implications of this directed structure, both for contact tracing strategy and for interventions designed to prevent superspreading events. In particular, directed graphs have a forward and backward version of the classic "friendship paradox" -- forward edges tend to lead to individuals with high risk, while backward edges lead to individuals with high spread -- such that a combination of both forward and backward contact tracing is necessary to find superspreading events and prevent future cascades of infection.

physics.soc-ph

The unintended consequences of inconsistent pandemic control policies

Controlling the spread of COVID-19 - even after a licensed vaccine is available - requires the effective use of non-pharmaceutical interventions: physical distancing, limits on group sizes, mask wearing, etc. To date, such interventions have neither been uniformly nor systematically implemented in most countries. For example, even when under strict stay-at-home orders, numerous jurisdictions granted exceptions and/or were in close proximity to locations with entirely different regulations in place. Here, we investigate the impact of such geographic inconsistencies in epidemic control policies by coupling search and mobility data to a simple mathematical model of SARS-COV2 transmission. Our results show that while stay-at-home orders decrease contacts in most areas of the US, some specific activities and venues often see an increase in attendance. Indeed, over the month of March 2020, between 10 and 30% of churches in the US saw increases in attendance; even as the total number of visits to churches declined nationally. This heterogeneity, where certain venues see substantial increases in attendance while others close, suggests that closure can cause individuals to find an open venue, even if that requires longer-distance travel. And, indeed, the average distance travelled to churches in the US rose by 13% over the same period. Strikingly, our model reveals that across a broad range of model parameters, partial measures can often be worse than none at all where individuals not complying with policies by traveling to neighboring areas can create epidemics when the outbreak would otherwise have been controlled. Taken together, our data analysis and modelling results highlight the potential unintended consequences of inconsistent epidemic control policies and stress the importance of balancing the societal needs of a population with the risk of an outbreak growing into a large epidemic.

q-bio.PE

Stochasticity and heterogeneity in the transmission dynamics of SARS-CoV-2

SARS-CoV-2 causing COVID-19 disease has moved rapidly around the globe, infecting millions and killing hundreds of thousands. The basic reproduction number, which has been widely used and misused to characterize the transmissibility of the virus, hides the fact that transmission is stochastic, is dominated by a small number of individuals, and is driven by super-spreading events (SSEs). The distinct transmission features, such as high stochasticity under low prevalence, and the central role played by SSEs on transmission dynamics, should not be overlooked. Many explosive SSEs have occurred in indoor settings stoking the pandemic and shaping its spread, such as long-term care facilities, prisons, meat-packing plants, fish factories, cruise ships, family gatherings, parties and night clubs. These SSEs demonstrate the urgent need to understand routes of transmission, while posing an opportunity that outbreak can be effectively contained with targeted interventions to eliminate SSEs. Here, we describe the potential types of SSEs, how they influence transmission, and give recommendations for control of SARS-CoV-2.

q-bio.PE

Beyond $R_0$: Heterogeneity in secondary infections and probabilistic epidemic forecasting

The basic reproductive number -- $R_0$ -- is one of the most common and most commonly misapplied numbers in public health. Although often used to compare outbreaks and forecast pandemic risk, this single number belies the complexity that two different pathogens can exhibit, even when they have the same $R_0$. Here, we show how to predict outbreak size using estimates of the distribution of secondary infections, leveraging both its average $R_0$ and the underlying heterogeneity. To do so, we reformulate and extend a classic result from random network theory that relies on contact tracing data to simultaneously determine the first moment ($R_0$) and the higher moments (representing the heterogeneity) in the distribution of secondary infections. Further, we show the different ways in which this framework can be implemented in the data-scarce reality of emerging pathogens. Lastly, we demonstrate that without data on the heterogeneity in secondary infections for emerging infectious diseases like COVID-19, the uncertainty in outbreak size ranges dramatically. Taken together, our work highlights the critical need for contact tracing during emerging infectious disease outbreaks and the need to look beyond $R_0$ when predicting epidemic size.

q-bio.PE

Mobile phone data and COVID-19: Missing an opportunity?

This paper describes how mobile phone data can guide government and public health authorities in determining the best course of action to control the COVID-19 pandemic and in assessing the effectiveness of control measures such as physical distancing. It identifies key gaps and reasons why this kind of data is only scarcely used, although their value in similar epidemics has proven in a number of use cases. It presents ways to overcome these gaps and key recommendations for urgent action, most notably the establishment of mixed expert groups on national and regional level, and the inclusion and support of governments and public authorities early on. It is authored by a group of experienced data scientists, epidemiologists, demographers and representatives of mobile network operators who jointly put their work at the service of the global effort to combat the COVID-19 pandemic.

cs.CY

Interacting contagions are indistinguishable from social reinforcement

From fake news to innovative technologies, many contagions spread via a process of social reinforcement, where multiple exposures are distinct from prolonged exposure to a single source. Contrarily, biological agents such as Ebola or measles are typically thought to spread as simple contagions. Here, we demonstrate that interacting simple contagions are indistinguishable from complex contagions. In the social context, our results highlight the challenge of identifying and quantifying mechanisms, such as social reinforcement, in a world where an innumerable amount of ideas, memes and behaviors interact. In the biological context, this parallel allows the use of complex contagions to effectively quantify the non-trivial interactions of infectious diseases.

physics.soc-ph

On the predictability of infectious disease outbreaks

Infectious disease outbreaks recapitulate biology: they emerge from the multi-level interaction of hosts, pathogens, and their shared environment. As a result, predicting when, where, and how far diseases will spread requires a complex systems approach to modeling. Recent studies have demonstrated that predicting different components of outbreaks--e.g., the expected number of cases, pace and tempo of cases needing treatment, demand for prophylactic equipment, importation probability etc.--is feasible. Therefore, advancing both the science and practice of disease forecasting now requires testing for the presence of fundamental limits to outbreak prediction. To investigate the question of outbreak prediction, we study the information theoretic limits to forecasting across a broad set of infectious diseases using permutation entropy as a model independent measure of predictability. Studying the predictability of a diverse collection of historical outbreaks--including, chlamydia, dengue, gonorrhea, hepatitis A, influenza, measles, mumps, polio, and whooping cough--we identify a fundamental entropy barrier for infectious disease time series forecasting. However, we find that for most diseases this barrier to prediction is often well beyond the time scale of single outbreaks. We also find that the forecast horizon varies by disease and demonstrate that both shifting model structures and social network heterogeneity are the most likely mechanisms for the observed differences across contagions. Our results highlight the importance of moving beyond time series forecasting, by embracing dynamic modeling approaches, and suggest challenges for performing model selection across long time series. We further anticipate that our findings will contribute to the rapidly growing field of epidemiological forecasting and may relate more broadly to the predictability of complex adaptive systems.

physics.soc-ph

The Interhospital Transfer Network for Very Low Birth Weight Infants in the United States

Very low birth weight (VLBW) infants require specialized care in neonatal intensive care units. In the United States (U.S.), such infants frequently are transferred between hospitals. Although these neonatal transfer networks are important, both economically and for infant morbidity and mortality, the national-level pattern of neonatal transfers is largely unknown. Using data from Vermont Oxford Network on 44,753 births, 2,122 hospitals, and 9,722 inter-hospital infant transfers from 2015, we performed the largest analysis to date on the inter-hospital transfer network for VLBW infants in the U.S. We find that transfers are organized around regional communities, but that despite being largely within state boundaries, most communities often contain at least two hospitals in different states. To classify the structural variation in transfer pattern amongst these communities, we applied a spectral measure for regionalization and found an association between a community's degree of regionalization and their infant transfer rate, which was not utilized in detecting communities. We also demonstrate that the established measures of network centrality and hierarchy, e.g., the community-wide entropy in PageRank or betweenness centrality and number of distinct `layers' within a community, correlate weakly with our regionalization index and were not significantly associated with metrics on infant transfer rate. Our results suggest that the regionalization index captures novel information about the structural properties of VLBW infant transfer networks, have the practical implication of characterizing neonatal care in the U.S., and may apply more broadly to the role of centralizing forces in organizing complex adaptive systems.

physics.soc-ph

Socioeconomic bias in influenza surveillance

Individuals in low socioeconomic brackets are considered at-risk for developing influenza-related complications and often exhibit higher than average influenza-related hospitalization rates. This disparity has been attributed to various factors, including restricted access to preventative and therapeutic health care, limited sick leave, and household structure. Adequate influenza surveillance in these at-risk populations is a critical precursor to accurate risk assessments and effective intervention. However, the United States of America's primary national influenza surveillance system (ILINet) monitors outpatient healthcare providers, which may be largely inaccessible to lower socioeconomic populations. Recent initiatives to incorporate internet-source and hospital electronic medical records data into surveillance systems seek to improve the timeliness, coverage, and accuracy of outbreak detection and situational awareness. Here, we use a flexible statistical framework for integrating multiple surveillance data sources to evaluate the adequacy of traditional (ILINet) and next generation (BioSense 2.0 and Google Flu Trends) data for situational awareness of influenza across poverty levels. We find that zip codes in the highest poverty quartile are a critical blind-spot for ILINet that the integration of next generation data fails to ameliorate.

stat.AP

Asymmetric percolation drives a double transition in sexual contact networks

Zika virus (ZIKV) exhibits unique transmission dynamics in that it is concurrently spread by a mosquito vector and through sexual contact. We show that this sexual component of ZIKV transmission induces novel processes on networks through the highly asymmetric durations of infectiousness between males and females -- it is estimated that males are infectious for periods up to ten times longer than females -- leading to an asymmetric percolation process on the network of sexual contacts. We exactly solve the properties of this asymmetric percolation on random sexual contact networks and show that this process exhibits two epidemic transitions corresponding to a core-periphery structure. This structure is not present in the underlying contact networks, which are not distinguishable from random networks, and emerges because of the asymmetric percolation. We provide an exact analytical description of this double transition and discuss the implications of our results in the context of ZIKV epidemics. Most importantly, our study suggests a bias in our current ZIKV surveillance as the community most at risk is also one of the least likely to get tested.

physics.soc-ph