SearcharxivSearch

arXiv subjects

Carter T. Butts

Publications and source records attributed to Carter T. Butts.

At least 19 recordsLinked to original sources

Motional Degrees of Freedom in Network Hamiltonian Models

Network Hamiltonian Models (NHMs) provide an efficient framework for modeling the aggregation of interacting particles (e.g., the condensation of proteins into gel-like, oligomeric, or fibrillar states), representing the system as a network whose edges represent bound interactions. Terms within the network Hamiltonian represent multi-body interactions governing aggregation behavior, and are specified via topological degrees of freedom. Because motional degrees of freedom are not explicitly represented within the NHM, their influence must be indirectly accounted for by introduction of corresponding terms to the Hamiltonian. Here, we describe specifications for terms representing motional degrees of freedom of two, three, and four-body interactions. We also consider the impact of these terms on the aggregation states of a minimal system governed only by a pairwise edge potential, showing that three and four-body interactions favor condensation of the system into small droplet-like structures at low temperature.

q-bio.MN

A Behavioral Micro-foundation for Cross-sectional Network Models

Models for cross-sectional network data have become increasingly well-developed in recent decades, and are widely used. This has led to a growing interest in the connection between such cross-sectional models and the behavioral processes from which the corresponding networks were presumably generated. Here, we build on prior work in this area to present a behavioral micro-foundation for cross-sectional network models, based on a continuous time stochastic choice mechanism, that can accommodate highly general classes of cases (including agents who are not themselves in the network, and multilateral edge control). As we show, the equilibrium behavior of this process under appropriate conditions can be expressed in exponential family form, allowing estimation of individual preferences using existing methods; the graph potential separates naturally into a preference-based term reflecting agent utilities, and an entropic term reflecting the rules of tie formation. We illustrate our approach via an analysis of friendship in a professional organization, and modeling of phase transitions in the structure of small groups.

cs.SI

The Decay of Impact with Network Distance in Linear Diffusion Processes

Many processes related to status, power, and influence within social networks have been modeled using forced linear diffusion models; examples include the highly successful Friedkin-Johnsen model of social influence, the status/power scores of Katz and Bonacich, and the widely used network autocorrelation model. While a basic assumption of such models is that the impact of one individual on another through any given path falls exponentially with path length, the total impact of the first individual on the second involves contributions from walks of all lengths; thus, while total impact is expected to decline with network distance, the relationship is not trivial. Here, we provide an approximate solution for the total impact of one node on another as a function of network distance, showing that the total impact is given to first order by a product of eigenvector centrality scores together with an expression in terms of the graph spectrum (eigenvalues of the adjacency matrix) that falls exponentially with distance. We also show how this solution can be refined using higher-order eigenvectors of the adjacency matrix. A numerical study on interpersonal networks drawn from educational settings verifies an average exponential decline in impact strength under the linear diffusion model, and shows that the first-order eigenvector approximation can often be a good proxy for total impact as obtained from the exact solution. This suggests a simple model that can be used to approximate total impact for social influence or status processes in a range of settings.

cs.SI

Transition State Theory for Network Dynamics

Many classic questions of structural theory concern discrete changes, such as the formation or dissolution of groups, role turnover, or faction realignment. Here, we consider a basic framework combining prior work on change paths and recent advances in dynamic network modeling with ideas from transition state theory. This framework facilitates both characterizing the process of structural change and, in some cases, predicting it. Notably, this approach allows approximate prediction of network change from cross-sectional models, under limited assumptions regarding the underlying microdynamics. We apply this framework to a simple model of faction realignment in small groups, showing that the process through which realignment occurs can be well-predicted ex ante for a number of different network micro-processes.

cs.SI

A Return to Biased Nets: New Specifications and Approximate Bayesian Inference

The biased net paradigm was the first general and empirically tractable scheme for parameterizing complex patterns of dependence in networks, expressing deviations from uniform random graph structure in terms of latent ``bias events,'' whose realizations enhance reciprocity, transitivity, or other structural features. Subsequent developments have introduced local specifications of biased nets, which reduce the need for approximations required in early specifications based on tracing processes. Here, we show that while one such specification leads to inconsistencies, a closely related Markovian specification both evades these difficulties and can be extended to incorporate new types of effects. We introduce the notion of inhibitory bias events, with satiation as an example, which are useful for avoiding degeneracies that can arise from closure bias terms. Although our approach does not lead to a computable likelihood, we provide a strategy for approximate Bayesian inference using random forest prevision. We demonstrate our approach on a network of friendship ties among college students, recapitulating a relationship between the sibling bias and tie strength posited in earlier work by Fararo.

stat.ME

Rooted America: Immobility and Segregation of the Intercounty Migration Network

Despite the popular narrative that the United States is a "land of mobility," the country may have become a "rooted America" after a decades-long decline in migration rates. This article interrogates the lingering question about the social forces that limit migration, with an empirical focus on internal migration in the United States. We propose a systemic, network model of migration flows, combining demographic, economic, political, and geographic factors and network dependence structures that reflect the internal dynamics of migration systems. Using valued temporal exponential-family random graph models, we model the network of intercounty migration flows from 2011 to 2015. Our analysis reveals a pattern of segmented immobility, where fewer people migrate between counties with dissimilar political contexts, levels of urbanization, and racial compositions. Probing our model using "knockout experiments" suggests one would have observed approximately 4.6 million (27 percent) more intercounty migrants each year were the segmented immobility mechanisms inoperative. This article offers a systemic view of internal migration and reveals the social and political cleavages that underlie geographic immobility in the United States.

cs.SI

A Perturbative Solution to the Linear Influence/Network Autocorrelation Model Under Network Dynamics

Known by many names and arising in many settings, the forced linear diffusion model is central to the modeling of power and influence within social networks (while also serving as the mechanistic justification for the widely used spatial/network autocorrelation models). The standard equilibrium solution to the diffusion model depends on strict timescale separation between network dynamics and attribute dynamics, such that the diffusion network can be considered fixed with respect to the diffusion process. Here, we consider a relaxation of this assumption, in which the network changes only slowly relative to the diffusion dynamics. In this case, we show that one can obtain a perturbative solution to the diffusion model, which depends on knowledge of past states in only a minimal way.

cs.SI

Common Ground In Crisis: Causal Narrative Networks of Public Official Communications During the COVID-19 Pandemic

This study investigates the use of causal narratives in public social media communications by U.S. public agencies over the first fifteen months of the COVID-19 pandemic. We extract causal narratives in the form of cause/effect pairs from official communications, analyzing the resulting semantic network to understand the structure and dependencies among concepts within agency discourse and the evolution of that discourse over time. We show that although the semantic network of causally-linked claims is complex and dynamic, there is considerable consistency across agencies in their causal assertions. We also show that the position of concepts within the structure of causal discourse has a significant impact on message retransmission net of controls, an important engagement outcome.

cs.SI

Calling The Dead: Resilience In The WTC Communication Networks

Organizations in emergency settings must cope with various sources of disruption, most notably personnel loss. Death, incapacitation, or isolation of individuals within an organizational communication network can impair information passing, coordination, and connectivity, and may drive maladaptive responses such as repeated attempts to contact lost personnel (``calling the dead'') that themselves consume scarce resources. At the same time, organizations may respond to such disruption by reorganizing to restore function, a behavior that is fundamental to organizational resilience. Here, we use empirically calibrated models of communication for 17 groups of responders to the World Trade Center Disaster to examine the impact of exogenous removal of personnel on communication activity and network resilience. We find that removal of high-degree personnel and those in institutionally coordinative roles is particularly damaging to these organizations, with specialist responders being slower to adapt to losses. However, all organizations show adaptations to disruption, in some cases becoming better connected and making more complete use of personnel relative to control after experiencing losses.

cs.SI

California Exodus? A Network Model of Population Redistribution in the United States

Motivated by debates about California's net migration loss, we employ valued exponential-family random graph models to analyze the inter-county migration flow networks in the United States. We introduce a protocol that visualizes the complex effects of potential underlying mechanisms, and perform in silico knockout experiments to quantify their contribution to the California Exodus. We find that racial dynamics contribute to the California Exodus, urbanization ameliorates it, and political climate and housing costs have little impact. Moreover, the severity of the California Exodus depends on how one measures it, and California is not the state with the most substantial population loss. The paper demonstrates how generative statistical models can provide mechanistic insights beyond simple hypothesis-testing.

cs.SI

Parameter Estimation Procedures for Exponential-Family Random Graph Models on Count-Valued Networks: A Comparative Simulation Study

The exponential-family random graph models (ERGMs) have emerged as an important framework for modeling social networks for a wide variety of relational types. ERGMs for valued networks are less well-developed than their unvalued counterparts, and pose particular computational challenges. Network data with edge values on the non-negative integers (count-valued networks) is an important such case, with examples ranging from the magnitude of migration and trade flows between places to the frequency of interactions and encounters between individuals. Here, we propose an efficient parallelable subsampled maximum pseudo-likelihood estimation (MPLE) scheme for count-valued ERGMs, and compare its performance with existing Contrastive Divergence (CD) and Monte Carlo Maximum Likelihood Estimation (MCMLE) approaches via a simulation study based on migration flow networks in two U.S. states. Our results suggest that edge value variance is a key factor in method performance, while network size mainly influences their relative merits in computational time. For small-variance networks, all methods perform well in point estimations while CD greatly overestimates uncertainties, and MPLE underestimates them for dependence terms; all methods have fast estimation for small networks, but CD and subsampled multi-core MPLE provides speed advantages as network size increases. For large-variance networks, both MPLE and MCMLE offer high-quality estimates of coefficients and their uncertainty, but MPLE is significantly faster than MCMLE; MPLE is also a better seeding method for MCMLE than CD, as the latter makes MCMLE more prone to convergence failure.

stat.ME

Continuous Time Graph Processes with Known ERGM Equilibria: Contextual Review, Extensions, and Synthesis

Graph processes that unfold in continuous time are of obvious theoretical and practical interest. Particularly useful are those whose long-term behavior converges to a graph distribution of known form. Here, we review some of the conditions for such convergence, and provide examples of novel and/or known processes that do so. These include subfamilies of the well-known stochastic actor oriented models, as well as continuum extensions of temporal and separable temporal exponential family random graph models. We also comment on some related threads in the broader work on network dynamics, which provide additional context for the continuous time case.

stat.ME

Modeling Complex Interactions in a Disrupted Environment: Relational Events in the WTC Response

When subjected to a sudden, unanticipated threat, human groups characteristically self-organize to identify the threat, determine potential responses, and act to reduce its impact. Central to this process is the challenge of coordinating information sharing and response activity within a disrupted environment. In this paper, we consider coordination in the context of responses to the 2001 World Trade Center disaster. Using records of communications among 17 organizational units, we examine the mechanisms driving communication dynamics, with an emphasis on the emergence of coordinating roles. We employ relational event models (REMs) to identify the mechanisms shaping communications in each unit, finding a consistent pattern of behavior across units with very different characteristics. Using a simulation-based "knock-out" study, we also probe the importance of different mechanisms for hub formation. Our results suggest that, while preferential attachment and pre-disaster role structure generally contribute to the emergence of hub structure, temporally local conversational norms play a much larger role. We discuss broader implications for the role of microdynamics in driving macroscopic outcomes, and for the emergence of coordination in other settings.

stat.AP

A Unified Prediction Framework for Signal Maps

Signal maps are essential for the planning and operation of cellular networks. However, the measurements needed to create such maps are expensive, often biased, not always reflecting the metrics of interest, and posing privacy risks. In this paper, we develop a unified framework for predicting cellular signal maps from limited measurements. Our framework builds on a state-of-the-art random-forest predictor, or any other base predictor. We propose and combine three mechanisms that deal with the fact that not all measurements are equally important for a particular prediction task. First, we design quality-of-service functions ($Q$), including signal strength (RSRP) but also other metrics of interest to operators, i.e., coverage and call drop probability. By implicitly altering the loss function employed in learning, quality functions can also improve prediction for RSRP itself where it matters (e.g., MSE reduction up to 27% in the low signal strength regime, where errors are critical). Second, we introduce weight functions ($W$) to specify the relative importance of prediction at different locations and other parts of the feature space. We propose re-weighting based on importance sampling to obtain unbiased estimators when the sampling and target distributions are different. This yields improvements up to 20% for targets based on spatially uniform loss or losses based on user population density. Third, we apply the Data Shapley framework for the first time in this context: to assign values ($ϕ$) to individual measurement points, which capture the importance of their contribution to the prediction task. This improves prediction (e.g., from 64% to 94% in recall for coverage loss) by removing points with negative values, and can also enable data minimization. We evaluate our methods and demonstrate significant improvement in prediction performance, using several real-world datasets.

cs.LG

Highly Scalable Maximum Likelihood and Conjugate Bayesian Inference for ERGMs on Graph Sets with Equivalent Vertices

The exponential family random graph modeling (ERGM) framework provides a flexible approach for the statistical analysis of networks. As ERGMs typically involve normalizing factors that are costly to compute, practical inference relies on a variety of approximations or other workarounds. Markov Chain Monte Carlo maximum likelihood (MCMC MLE) provides a powerful tool to approximate the MLE of ERGM parameters, and is feasible for typical models on single networks with as many as a few thousand nodes. MCMC-based algorithms for Bayesian analysis are more expensive, and high-quality answers are challenging to obtain on large graphs. For both strategies, extension to the pooled case - in which we observe multiple networks from a common generative process - adds further computational cost, with both time and memory scaling linearly in the number of graphs. This becomes prohibitive for large networks, or where large numbers of graph observations are available. Here, we exploit some basic properties of the discrete exponential families to develop an approach for ERGM inference in the pooled case that (where applicable) allows an arbitrarily large number of graph observations to be fit at no additional computational cost beyond preprocessing the data itself. Moreover, a variant of our approach can also be used to perform Bayesian inference under conjugate priors, again with no additional computational cost in the estimation phase. As we show, the conjugate prior is easily specified, and is well-suited to applications such as regularization. Simulation studies show that the pooled method leads to estimates with good frequentist properties, and posterior estimates under the conjugate prior are well-behaved. We demonstrate our approach with applications to pooled analysis of brain functional connectivity networks and to replicated x-ray crystal structures of hen egg-white lysozyme.

stat.ME

Neural Upscaling from Residue-level Protein Structure Networks to Atomistic Structure

Coarse-graining is a powerful tool for extending the reach of dynamic models of proteins and other biological macromolecules. Topological coarse-graining, in which biomolecules or sets thereof are represented via graph structures, is a particularly useful way of obtaining highly compressed representations of molecular structure, and simulations operating via such representations can achieve substantial computational savings. A drawback of coarse-graining, however, is the loss of atomistic detail - an effect that is especially acute for topological representations such as protein structure networks (PSNs). Here, we introduce an approach based on a combination of machine learning and physically-guided refinement for inferring atomic coordinates from PSNs. This "neural upscaling" procedure exploits the constraints implied by PSNs on possible configurations, as well as differences in the likelihood of observing different configurations with the same PSN. Using a 1 $μ$s atomistic molecular dynamics trajectory of A$β_{1-40}$, we show that neural upscaling is able to effectively recapitulate detailed structural information for intrinsically disordered proteins, being particularly successful in recovering features such as transient secondary structure. These results suggest that scalable network-based models for protein structure and dynamics may be used in settings where atomistic detail is desired, with upscaling employed to impute atomic coordinates from PSNs.

q-bio.BM

Bayesian Estimation of the Hydroxyl Radical Diffusion Coefficient at Low Temperature and High Pressure from Atomistic Molecular Dynamics

The hydroxyl radical is the primary reactive oxygen species produced by the radiolysis of water, and is a significant source of radiation damage to living organisms. Mobility of the hydroxyl radical at low temperatures and/or high pressures is hence a potentially important factor in determining the challenges facing psychrophilic and/or barophilic organisms in high-radiation environments (e.g., ice-interface or undersea environments in which radiative heating is a potential heat and energy source). Here, we estimate the diffusion coefficient for the hydroxyl radical in aqueous solution, using a hierarchical Bayesian model based on atomistic molecular dynamics trajectories in TIP4P/2005 water over a range of temperatures and pressures.

stat.AP

Bayesian Analysis of Static Light Scattering Data for Globular Proteins

Static light scattering is a popular physical chemistry technique that enables calculation of physical attributes such as the radius of gyration and the second virial coefficient for a macromolecule (e.g., a polymer or a protein) in solution. The second virial coefficient is a physical quantity that characterizes the magnitude and sign of pairwise interactions between particles, and hence is related to aggregation propensity, a property of considerable scientific and practical interest. Estimating the second virial coefficient from experimental data is challenging due both to the degree of precision required and the complexity of the error structure involved. In contrast to conventional approaches based on heuristic OLS estimates, Bayesian inference for the second virial coefficient allows explicit modeling of error processes, incorporation of prior information, and the ability to directly test competing physical models. Here, we introduce a fully Bayesian model for static light scattering experiments on small-particle systems, with joint inference for concentration, index of refraction, oligomer size, and the second virial coefficient. We apply our proposed model to study the aggregation behavior of hen egg-white lysozyme and human gammaS-crystallin using in-house experimental data. Based on these observations, we also perform a simulation study on the primary drivers of uncertainty in this family of experiments, showing in particular the potential for improved monitoring and control of concentration to aid inference.

stat.AP