SearcharxivSearch

arXiv subjects

Jacob G. Foster

Publications and source records attributed to Jacob G. Foster.

14 recordsLinked to original sources

The content and structure of dreams are coupled to affect

Dreams offer a unique window into the cognitive and affective dynamics of the sleeping and the waking mind. Recent quantitative linguistic approaches have shown promise in obtaining corpus-level measures of dream sentiment and topic occurrence. However, it is currently unclear how the affective content of individual dreams relates to their semantic content and structure. Here, we combine word embedding, topic modeling, and network analysis to investigate this relationship. By applying Discourse Atom Topic Modeling (DATM) to the DreamBank corpus of >18K dream reports, we represent the latent themes arising within dream reports as a sparse dictionary of topics and identify the affective associations of those topics. We show that variation in dream report affect (valence and arousal) is associated with changes in topical content. By representing each dream report as a network of topics, we demonstrate that the affective content of dream narratives is also coupled to semantic structure. Positively valenced dream reports exhibit more coherent, structured, and linear narratives, whilst negatively valenced dreams have more narrative loops and dominant topics. Topic networks of high arousal dream reports are structurally dominated by few high arousal topics and incoherent topical connections, whereas low arousal dream reports contain more loops. These findings suggest that affective processes are associated with both the content and structure of dreams. Our approach showcases the potential of integrating natural language processing and network analysis with psychology to elucidate the interplay of affect, cognition and narrative in dreams. This methodology has broad applications for the study of narrated experience and psychiatric symptomatology.

q-bio.NC

Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research

Benchmark datasets play a central role in the organization of machine learning research. They coordinate researchers around shared research problems and serve as a measure of progress towards shared goals. Despite the foundational role of benchmarking practices in this field, relatively little attention has been paid to the dynamics of benchmark dataset use and reuse, within or across machine learning subcommunities. In this paper, we dig into these dynamics. We study how dataset usage patterns differ across machine learning subcommunities and across time from 2015-2020. We find increasing concentration on fewer and fewer datasets within task communities, significant adoption of datasets from other tasks, and concentration across the field on datasets that have been introduced by researchers situated within a small number of elite institutions. Our results have implications for scientific evaluation, AI ethics, and equity/access within the field.

cs.LG

Adapting Coreference Resolution for Processing Violent Death Narratives

Coreference resolution is an important component in analyzing narrative text from administrative data (e.g., clinical or police sources). However, existing coreference models trained on general language corpora suffer from poor transferability due to domain gaps, especially when they are applied to gender-inclusive data with lesbian, gay, bisexual, and transgender (LGBT) individuals. In this paper, we analyzed the challenges of coreference resolution in an exemplary form of administrative text written in English: violent death narratives from the USA's Centers for Disease Control's (CDC) National Violent Death Reporting System. We developed a set of data augmentation rules to improve model performance using a probabilistic data programming framework. Experiments on narratives from an administrative database, as well as existing gender-inclusive coreference datasets, demonstrate the effectiveness of data augmentation in training coreference models that can better handle text data about LGBT individuals.

cs.CL

Machine learning as a model for cultural learning: Teaching an algorithm what it means to be fat

As we navigate our cultural environment, we learn cultural biases, like those around gender, social class, health, and body weight. It is unclear, however, exactly how public culture becomes private culture. In this paper, we provide a theoretical account of such cultural learning. We propose that neural word embeddings provide a parsimonious and cognitively plausible model of the representations learned from natural language. Using neural word embeddings, we extract cultural schemata about body weight from New York Times articles. We identify several cultural schemata that link obesity to gender, immorality, poor health, and low socioeconomic class. Such schemata may be subtly but pervasively activated in public culture; thus, language can chronically reproduce biases. Our findings reinforce ongoing concerns that machine learning can also encode, and reproduce, harmful human biases.

cs.CY

Why Scientists Chase Big Problems: Individual Strategy and Social Optimality

Scientists pursue collective knowledge, but they also seek personal recognition from their peers. When scientists decide whether or not to work on a big new problem, they weigh the potential rewards of a major discovery against the costs of setting aside other projects. These self-interested choices can potentially spread researchers across problems in an efficient manner, but efficiency is not guaranteed. We use simple economic models to understand such decisions and their collective consequences. Academic science differs from industrial R&D in that academics often share partial solutions to gain reputation. This convention of Open Science is thought to accelerate collective discovery, but we find that it need not do so. The ability to share partial results influences which scientists work on a particular problem; consequently, Open Science can slow down the solution of a problem if it deters entry by important actors.

physics.soc-ph

Tradition and Innovation in Scientists' Research Strategies

What factors affect a scientist's choice of research problem? Qualitative research in the history, philosophy, and sociology of science suggests that this choice is shaped by an "essential tension" between the professional demand for productivity and a conflicting drive toward risky innovation. We examine this tension empirically in the context of biomedical chemistry. We use complex networks to represent the evolving state of scientific knowledge, as expressed in publications. We then define research strategies relative to these networks. Scientists can introduce novel chemicals or chemical relationships--or delve deeper into known ones. They can consolidate existing knowledge clusters, or bridge distant ones. Analyzing such choices in aggregate, we find that the distribution of strategies remains remarkably stable, even as chemical knowledge grows dramatically. High-risk strategies, which explore new chemical relationships, are less prevalent in the literature, reflecting a growing focus on established knowledge at the expense of new opportunities. Research following a risky strategy is more likely to be ignored but also more likely to achieve high impact and recognition. While the outcome of a risky strategy has a higher expected reward than the outcome of a conservative strategy, the additional reward is insufficient to compensate for the additional risk. By studying the winners of 137 different prizes in biomedicine and chemistry, we show that the occasional "gamble" for extraordinary impact is the most plausible explanation for observed levels of risk-taking. Our empirical demonstration and unpacking of the "essential tension" suggests policy interventions that may foster more innovative research.

physics.soc-ph

Dynamic Landscapes: A Model of Context and Contingency in Evolution

The basic mechanics of evolution have been understood since Darwin. But debate continues over whether macroevolutionary phenomena are driven primary by the fitness structure of genotype space or by ecological interaction. In this paper we propose a simple, abstract model capturing some key features of fitness-landscape and ecological models of evolution. Our model describes evolutionary dynamics in a high-dimensional, structured genotype space with a significant role for interspecific interaction. We find some promising qualitative similarity with the empirical facts about macroevolution, including broadly distributed extinction sizes and realistic exploration of the genotype space. The abstraction of our model permits numerous interpretations and applications beyond macroevolution, including molecular evolution and technological innovation.

q-bio.PE

Clustering Drives Assortativity and Community Structure in Ensembles of Networks

Clustering, assortativity, and communities are key features of complex networks. We probe dependencies between these attributes and find that ensembles with strong clustering display both high assortativity by degree and prominent community structure, while ensembles with high assortativity are much less biased towards clustering or community structure. Further, clustered networks can amplify small homophilic bias for trait assortativity. This marked asymmetry suggests that transitivity, rather than homophily, drives the standard nonsocial/social network dichotomy.

physics.soc-ph

Edge direction and the structure of networks

Directed networks are ubiquitous and are necessary to represent complex systems with asymmetric interactions---from food webs to the World Wide Web. Despite the importance of edge direction for detecting local and community structure, it has been disregarded in studying a basic type of global diversity in networks: the tendency of nodes with similar numbers of edges to connect. This tendency, called assortativity, affects crucial structural and dynamic properties of real-world networks, such as error tolerance or epidemic spreading. Here we demonstrate that edge direction has profound effects on assortativity. We define a set of four directed assortativity measures and assign statistical significance by comparison to randomized networks. We apply these measures to three network classes---online/social networks, food webs, and word-adjacency networks. Our measures (i) reveal patterns common to each class, (ii) separate networks that have been previously classified together, and (iii) expose limitations of several existing theoretical models. We reject the standard classification of directed networks as purely assortative or disassortative. Many display a class-specific mixture, likely reflecting functional or historical constraints, contingencies, and forces guiding the system's evolution.

physics.soc-ph

Physics With Two Time Dimensions

We explore the properties of physical theories in space-times with two time dimensions. We show that the common arguments used to rule such theories out do not apply if the dynamics associated with the additional time dimension is thermal or chaotic and does not permit long-lived time-like excitations. We discuss several possible realizations of such theories, including holographic representations and the possibility that quantum dynamics emerges as a consequence of a second time dimension.

hep-th

Clustering Phase Transitions and Hysteresis: Pitfalls in Constructing Network Ensembles

Ensembles of networks are used as null models in many applications. However, simple null models often show much less clustering than their real-world counterparts. In this paper, we study a model where clustering is enhanced by means of a fugacity term as in the Strauss (or "triangle") model, but where the degree sequence is strictly preserved -- thus maintaining the quenched heterogeneity of nodes found in the original degree sequence. Similar models had been proposed previously in [R. Milo et al., Science 298, 824 (2002)]. We find that our model exhibits phase transitions as the fugacity is changed. For regular graphs (identical degrees for all nodes) with degree k > 2 we find a single first order transition. For all non-regular networks that we studied (including Erdos - Renyi and scale-free networks) we find multiple jumps resembling first order transitions, together with strong hysteresis. The latter transitions are driven by the sudden emergence of "cluster cores": groups of highly interconnected nodes with higher than average degrees. To study these cluster cores visually, we introduce q-clique adjacency plots. We find that these cluster cores constitute distinct communities which emerge spontaneously from the triangle generating process. Finally, we point out that cluster cores produce pitfalls when using the present (and similar) models as null models for strongly clustered networks, due to the very strong hysteresis which effectively leads to broken ergodicity on realistic time scales.

cond-mat.stat-mech

A Model of Sequential Branching in Hierarchical Cell Fate Determination

Multipotent stem or progenitor cells undergo a sequential series of binary fate decisions, which ultimately generate the diversity of differentiated cells. Efforts to understand cell fate control have focused on simple gene regulatory circuits that predict the presence of multiple stable states, bifurcations and switch-like transitions. However, existing gene network models do not explain more complex properties of cell fate dynamics such as the hierarchical branching of developmental paths. Here, we construct a generic minimal model of the genetic regulatory network controlling cell fate determination, which exhibits five elementary characteristics of cell differentiation: stability, directionality, branching, exclusivity, and promiscuous expression. We argue that a modular architecture comprising repeated network elements reproduces these features of differentiation by sequentially repressing selected modules and hence restricting the dynamics to lower dimensional subspaces of the high-dimensional state space. We implement our model both with ordinary differential equations (ODEs), to explore the role of bifurcations in producing the one-way character of differentiation, and with stochastic differential equations (SDEs), to demonstrate the effect of noise on the system. We further argue that binary cell fate decisions are prevalent in cell differentiation due to general features of the underlying dynamical system. This minimal model makes testable predictions about the structural basis for directional, discrete and diversifying cell phenotype development and thus can guide the evaluation of real gene regulatory networks that govern differentiation.

q-bio.MN

Reinforced walks in two and three dimensions

In probability theory, reinforced walks are random walks on a lattice (or more generally a graph) that preferentially revisit neighboring `locations' (sites or bonds) that have been visited before. In this paper, we consider walks with one-step reinforcement, where one preferentially \emph{revisits} locations irrespective of the number of visits. Previous numerical simulations [A. Ordemann {\it et al.}, Phys. Rev. E {\bf 64}, 046117 (2001)] suggested that the site model on the lattice shows a phase transition at finite reinforcement between a random-walk like and a collapsed phase, in both 2 and 3 dimensions. The very different mathematical structure of bond and site models might also suggest different phenomenology (critical properties, etc.). We use high statistics simulations and heuristic arguments to suggest that site and bond reinforcement are in the same universality class, and that the purported phase transition in 2 dimensions actually occurs at zero coupling constant. We also show that a quasi-static approximation predicts the large time scaling of the end-to-end distance in the collapsed phase of both site and bond reinforcement models, in excellent agreement with simulation results.

cond-mat.stat-mech

Link and subgraph likelihoods in random undirected networks with fixed and partially fixed degree sequence

The simplest null models for networks, used to distinguish significant features of a particular network from {\it a priori} expected features, are random ensembles with the degree sequence fixed by the specific network of interest. These "fixed degree sequence" (FDS) ensembles are, however, famously resistant to analytic attack. In this paper we introduce ensembles with partially-fixed degree sequences (PFDS) and compare analytic results obtained for them with Monte Carlo results for the FDS ensemble. These results include link likelihoods, subgraph likelihoods, and degree correlations. We find that local structural features in the FDS ensemble can be reasonably well estimated by simultaneously fixing only the degrees of few nodes, in addition to the total number of nodes and links. As test cases we use a food web, two protein interaction networks (\textit{E. coli, S. cerevisiae}), the internet on the autonomous system (AS) level, and the World Wide Web. Fixing just the degrees of two nodes gives the mean neighbor degree as a function of node degree, $ _k$, in agreement with results explicitly obtained from rewiring. For power law degree distributions, we derive the disassortativity analytically. In the PFDS ensemble the partition function can be expanded diagrammatically. We obtain an explicit expression for the link likelihood to lowest order, which reduces in the limit of large, sparse undirected networks with $L$ links and with $k_{\rm max} \ll L$ to the simple formula $P(k,k') = kk'/(2L + kk')$. In a similar limit, the probability for three nodes to be linked into a triangle reduces to the factorized expression $P_Δ(k_1,k_2,k_3) = P(k_1,k_2)P(k_1,k_3)P(k_2,k_3)$.

cond-mat.stat-mech