SearcharxivSearch

arXiv subjects

Jean-Gabriel Young

Publications and source records attributed to Jean-Gabriel Young.

At least 19 recordsLinked to original sources

Emergent contagion complexity: Disentangling mechanistic complexity from correlated heterogeneity

Simple and complex contagions differ mechanistically; multiple exposures act synergistically in the latter but independently in the former. Yet correlated mixtures of simple contagions may appear complex when inferring global contagion rules, a phenomenon we call "emergent complexity." We present a measure of contagion complexity and an inferential framework for estimating mixtures of nonparametric contagion rules from time-series data. Our work reframes past studies on complex contagion by offering heterogeneous mixtures of simple contagions as an alternative explanation.

physics.soc-ph

The nestedness of higher-order networks

In contrast to dyadic interactions, higher-order interactions may contain one another, with subgroups naturally embedded within larger groups. These containment patterns arise empirically in ecology, sociology, computer science and the science of science, and have been studied under the names nestedness, simpliciality, encapsulation, and inclusion. In this chapter, we review each of these measures and unify them through a mathematical object known as the encapsulation directed acyclic graph, formulating each measure as a function of its properties. We demonstrate that nested structure is prevalent in social systems across several domains, show that different measures capture complementary aspects of this structure, and find that the absence of nestedness can itself be a powerful indicator of the mesoscale organization of a system.

physics.soc-ph

Simpson's paradox explains the ubiquity of nonlinear, threshold, and complex contagions

Complex contagions describe systems where the probability or rate of contagious transmission is a nonlinear function of the exposure to contagious agents. These models were first studied theoretically but have since been used to capture effects such as nonconformism, social reinforcement or peer pressure in empirical data. However, recent studies have shown that local correlations (e.g., group structure or temporal burstiness) and heterogeneity (e.g., diversity of parameters or covariates) can give the illusion of nonlinear effects even when the dynamics is actually linear. We briefly review these studies to inform a new model and explanation for these effective models of complex contagions. We find global threshold dynamics and superlinear complex contagions even in populations where agents are distributed across social groups described solely by linear or even sublinear contagions. This effect can be understood as a manifestation of Simpson's paradox. Incidence data from heterogeneous groups can look superlinear once averaged over all groups, since the sampling of groups represented at high incidence is biased towards those with stronger local transmission. We then define what we call a Simpson's contagion: a contagion process that looks superlinear when observed over an entire population, but is mechanistically linear or even sublinear in all of its subgroups. By exploring these Simpson's contagions over mathematical case studies, our work contributes to the growing body of literature on the ubiquity of threshold and complex contagions as effective models, and our results stress the pitfall of model selection that ignores correlations and heterogeneity in populations.

physics.soc-ph

Inferring signed social networks from contact patterns

Social networks are typically inferred from indirect observations, such as proximity data; yet, most methods cannot distinguish between absent relationships and actual negative ties, as both can result in few or no interactions. We address the challenge of inferring signed networks from contact patterns while accounting for whether lack of interactions reflect a lack of opportunity as opposed to active avoidance. We develop a Bayesian framework with MCMC inference that models interaction groups to separate chance from choice when no interactions are observed. Validation on synthetic data demonstrates superior performance compared to natural baselines, particularly in detecting negative edges. We apply our method to French high school contact data to reveal a structure consistent with friendship surveys and demonstrate the model's adequacy through posterior predictive checks.

cs.SI

One pathogen does not an epidemic make: A review of interacting contagions, diseases, beliefs, and stories

From pathogens and computer viruses to genes and memes, contagion models have found widespread utility across the natural and social sciences. Despite their success and breadth of adoption, the approach and structure of these models remain surprisingly siloed by field. Given the siloed nature of their development and widespread use, one persistent assumption is that a given contagion can be studied in isolation, independently from what else might be spreading in the population. In reality, countless contagions of biological and social nature interact within hosts (interacting with existing beliefs, or the immune system) and across hosts (interacting in the environment, or affecting transmission mechanisms). Additionally, from a modeling perspective, we know that relaxing these assumptions has profound effects on the physics and translational implications of the models. Here, we review mechanisms for interactions in social and biological contagions, as well as the models and frameworks developed to include these interactions in the study of the contagions. We highlight existing problems related to the inference of interactions and to the scalability of mathematical models and identify promising avenues of future inquiries. In doing so, we highlight the need for interdisciplinary efforts under a unified science of contagions and for removing a common dichotomy between social and biological contagions.

physics.soc-ph

Message passing for epidemiological interventions on networks with loops

Spreading models capture key dynamics on networks, such as cascading failures in economic systems, (mis)information diffusion, and pathogen transmission. Here, we focus on design intervention problems -- for example, designing optimal vaccination rollouts or wastewater surveillance systems -- which can be solved by comparing outcomes under various counterfactuals. A leading approach to computing these outcomes is message passing, which allows for the rapid and direct computation of the marginal probabilities for each node. However, despite its efficiency, classical message passing tends to overestimate outbreak sizes on real-world networks, leading to incorrect predictions and, thus, interventions. Here, we improve these estimates by using the neighborhood message passing (NMP) framework for the epidemiological calculations. We evaluate the quality of the improved algorithm and demonstrate how it can be used to test possible solutions to three intervention design problems: influence maximization, optimal vaccination, and sentinel surveillance.

cs.SI

The simpliciality of higher-order networks

Higher-order networks are widely used to describe complex systems in which interactions can involve more than two entities at once. In this paper, we focus on inclusion within higher-order networks, referring to situations where specific entities participate in an interaction, and subsets of those entities also interact with each other. Traditional modeling approaches to higher-order networks tend to either not consider inclusion at all (e.g., hypergraph models) or explicitly assume perfect and complete inclusion (e.g., simplicial complex models). To allow for a more nuanced assessment of inclusion in higher-order networks, we introduce the concept of "simpliciality" and several corresponding measures. Contrary to current modeling practice, we show that empirically observed systems rarely lie at either end of the simpliciality spectrum. In addition, we show that generative models fitted to these datasets struggle to capture their inclusion structure. These findings suggest new modeling directions for the field of higher-order network science.

physics.soc-ph

Sensitivity analysis of epidemic forecasting and spreading on networks with probability generating functions

Epidemic forecasting tools embrace the stochasticity and heterogeneity of disease spread to predict the growth and size of outbreaks. Conceptually, stochasticity and heterogeneity are often modeled as branching processes or as percolation on contact networks. Mathematically, probability generating functions provide a flexible and efficient tool to describe these models and quickly produce forecasts. While their predictions are probabilistic-i.e., distributions of outcome-they depend deterministically on the input distribution of transmission statistics and/or contact structure. Since these inputs can be noisy data or models of high dimension, traditional sensitivity analyses are computationally prohibitive and are therefore rarely used. Here, we use statistical condition estimation to measure the sensitivity of stochastic polynomials representing noisy generating functions. In doing so, we can separate the stochasticity of their forecasts from potential noise in their input. For standard epidemic models, we find that predictions are most sensitive at the critical epidemic threshold (basic reproduction number $R_0 = 1$) only if the transmission is sufficiently homogeneous (dispersion parameter $k > 0.3$). Surprisingly, in heterogeneous systems ($k \leq 0.3$), the sensitivity is highest for values of $R_{0} > 1$. We expect our methods will improve the transparency and applicability of the growing utility of probability generating functions as epidemic forecasting tools.

q-bio.PE

Governance as a complex, networked, democratic, satisfiability problem

Democratic governments comprise a subset of a population whose goal is to produce coherent decisions, solving societal challenges while respecting the will of the people. New governance frameworks represent this as a social network rather than as a hierarchical pyramid with centralized authority. But how should this network be structured? We model the decisions a population must make as a satisfiability problem and the structure of information flow involved in decision-making as a social hypergraph. This framework allows to consider different governance structures, from dictatorships to direct democracy. Between these extremes, we find a regime of effective governance where small overlapping decision groups make specific decisions and share information. Effective governance allows even incoherent or polarized populations to make coherent decisions at low coordination costs. Beyond simulations, our conceptual framework can explore a wide range of governance strategies and their ability to tackle decision problems that challenge standard governments.

physics.soc-ph

Network compression with configuration models and the minimum description length

Random network models, constrained to reproduce specific statistical features, are often used to represent and analyze network data and their mathematical descriptions. Chief among them, the configuration model constrains random networks by their degree distribution and is foundational to many areas of network science. However, configuration models and their variants are often selected based on intuition or mathematical and computational simplicity rather than on statistical evidence. To evaluate the quality of a network representation, we need to consider both the amount of information required to specify a random network model and the probability of recovering the original data when using the model as a generative process. To this end, we calculate the approximate size of network ensembles generated by the popular configuration model and its generalizations, including versions accounting for degree correlations and centrality layers. We then apply the minimum description length principle as a model selection criterion over the resulting nested family of configuration models. Using a dataset of over 100 networks from various domains, we find that the classic Configuration Model is generally preferred on networks with an average degree above ten, while a Layered Configuration Model constrained by a centrality metric offers the most compact representation of the majority of sparse networks.

cs.SI

Reconstructing networks from simple and complex contagions

Network scientists often use complex dynamic processes to describe network contagions, but tools for fitting contagion models typically assume simple dynamics. Here, we address this gap by developing a nonparametric method to reconstruct a network and dynamics from a series of node states, using a model that breaks the dichotomy between simple pairwise and complex neighborhood-based contagions. We then show that a network is more easily reconstructed when observed through the lens of complex contagions if it is dense or the dynamic saturates, and that simple contagions are better otherwise.

cs.SI

Symmetry-driven embedding of networks in hyperbolic space

Hyperbolic models are known to produce networks with properties observed empirically in most network datasets, including heavy-tailed degree distribution, high clustering, and hierarchical structures. As a result, several embeddings algorithms have been proposed to invert these models and assign hyperbolic coordinates to network data. Current algorithms for finding these coordinates, however, do not quantify uncertainty in the inferred coordinates. We present BIGUE, a Markov chain Monte Carlo (MCMC) algorithm that samples the posterior distribution of a Bayesian hyperbolic random graph model. We show that the samples are consistent with current algorithms while providing added credible intervals for the coordinates and all network properties. We also show that some networks admit two or more plausible embeddings, a feature that an optimization algorithm can easily overlook.

stat.CO

Exact and rapid linear clustering of networks with dynamic programming

We study the problem of clustering networks whose nodes have imputed or physical positions in a single dimension, for example prestige hierarchies or the similarity dimension of hyperbolic embeddings. Existing algorithms, such as the critical gap method and other greedy strategies, only offer approximate solutions to this problem. Here, we introduce a dynamic programming approach that returns provably optimal solutions in polynomial time -- O(n^2) steps -- for a broad class of clustering objectives. We demonstrate the algorithm through applications to synthetic and empirical networks and show that it outperforms existing heuristics by a significant margin, with a similar execution time.

cs.SI

Compressing network populations with modal networks reveals structural diversity

Analyzing relational data consisting of multiple samples or layers involves critical challenges: How many networks are required to capture the variety of structures in the data? And what are the structures of these representative networks? We describe efficient nonparametric methods derived from the minimum description length principle to construct the network representations automatically. The methods input a population of networks or a multilayer network measured on a fixed set of nodes and output a small set of representative networks together with an assignment of each network sample or layer to one of the representative networks. We identify the representative networks and assign network samples to them with an efficient Monte Carlo scheme that minimizes our description length objective. For temporally ordered networks, we use a polynomial time dynamic programming approach that restricts the clusters of network layers to be temporally contiguous. These methods recover planted heterogeneity in synthetic network populations and identify essential structural heterogeneities in global trade and fossil record networks. Our methods are principled, scalable, parameter-free, and accommodate a wide range of data, providing a unified lens for exploratory analyses and preprocessing large sets of network samples.

physics.soc-ph

Accurately summarizing an outbreak using epidemiological models takes time

Recent outbreaks of monkeypox and Ebola, and worrying waves of COVID-19, influenza and respiratory syncytial virus, have all led to a sharp increase in the use of epidemiological models to estimate key epidemiological parameters. The feasibility of this estimation task is known as the practical identifiability (PI) problem. Here, we investigate the PI of eight commonly reported statistics of the classic Susceptible-Infectious-Recovered model using a new measure that shows how much a researcher can expect to learn in a model-based Bayesian analysis of prevalence data. Our findings show that the basic reproductive number and final outbreak size are often poorly identified, with learning exceeding that of individual model parameters only in the early stages of an outbreak. The peak intensity, peak timing, and initial growth rate are better identified, being in expectation over 20 times more probable having seen the data by the time the underlying outbreak peaks. We then test PI for a variety of true parameter combinations, and find that PI is especially problematic in slow-growing or less-severe outbreaks. These results add to the growing body of literature questioning the reliability of inferences from epidemiological models when limited data are available.

q-bio.PE

Latent Network Models to Account for Noisy, Multiply-Reported Social Network Data

Social network data are often constructed by incorporating reports from multiple individuals. However, it is not obvious how to reconcile discordant responses from individuals. There may be particular risks with multiply-reported data if people's responses reflect normative expectations -- such as an expectation of balanced, reciprocal relationships. Here, we propose a probabilistic model that incorporates ties reported by multiple individuals to estimate the unobserved network structure. In addition to estimating a parameter for each reporter that is related to their tendency of over- or under-reporting relationships, the model explicitly incorporates a term for ``mutuality,'' the tendency to report ties in both directions involving the same alter. Our model's algorithmic implementation is based on variational inference, which makes it efficient and scalable to large systems. We apply our model to data from 75 Indian villages collected with a name-generator design, and a Nicaraguan community collected with a roster-based design. We observe strong evidence of ``mutuality'' in both datasets, and find that this value varies by relationship type. Consequently, our model estimates networks with reciprocity values that are substantially different than those resulting from standard deterministic aggregation approaches, demonstrating the need to consider such issues when gathering, constructing, and analysing survey-based network data.

cs.SI

Hypergraph reconstruction from noisy pairwise observations

The network reconstruction task aims to estimate a complex system's structure from various data sources such as time series, snapshots, or interaction counts. Recent work has examined this problem in networks whose relationships involve precisely two entities-the pairwise case. Here we investigate the general problem of reconstructing a network in which higher-order interactions are also present. We study a minimal example of this problem, focusing on the case of hypergraphs with interactions between pairs and triplets of vertices, measured imperfectly and indirectly. We derive a Metropolis-Hastings-within-Gibbs algorithm for this model and use the algorithms to highlight the unique challenges that come with estimating higher-order models. We show that this approach tends to reconstruct empirical and synthetic networks more accurately than an equivalent graph model without higher-order interactions.

cs.SI

The Promise of Cross-Species Coexpression Analysis in Studying the Coevolution and Ecology of Host-Symbiont Interactions

Measuring gene expression simultaneously in both hosts and symbionts offers a powerful approach to explore the biology underlying species interactions. Such dual or simultaneous RNAseq approaches have primarily been used to gain insight into gene function in model systems, but there is opportunity to expand and apply these tools in new ways to understand ecological and evolutionary questions. By incorporating genetic diversity in both hosts and symbionts and studying how gene expression is correlated between partner species, we can gain new insight into host-symbiont coevolution and the ecology of species interactions. In this perspective, we explore how these relatively new tools could be applied to study such questions. We review the mechanisms that could be generating patterns of cross-species gene coexpression, including indirect genetic effects and selective filters, how these tools could be applied across different biological and temporal scales, and outline other methodological considerations and experiment possibilities.

q-bio.PE