SearcharxivSearch

arXiv subjects

Daniel B. Larremore

Publications and source records attributed to Daniel B. Larremore.

At least 19 recordsLinked to original sources

How large should academic departments be?

Academic departments are the primary unit of scholarship and education at universities, and they vary vastly in their sizes. However, the consequences and natural dynamics of department size are poorly understood. Small departments face disproportionate teaching and administrative overhead per faculty member, while large ones face coordination costs and thematic incoherence. Here, we characterize and model the dynamics of academic department sizes using $14,000$ U.S.-based departments in eight academic domains. Across all domains, similar broad-tailed distributions reveal a common size range from 4 to 23 faculty members, widening across domains at its upper border. Annual size-dependent closure risks and growth rates indicate that stability is greatest in this size range: below it, small departments either close or grow quickly; within it, closure risk is low and sizes stabilize; above it, large departments can persist, with marginal attrition and minimal closure risk. An analytically tractable model of size-dependent coagulation and fragmentation, informed only by the aggregated size distribution, reproduces department dynamics across the full size range. Rescaling each domain by its most stable size reveals a common regression toward the stable range across most domains. Our results establish academic departments as organizations with natural size dynamics defined by a grow-or-close pattern for the smallest departments, and a weak pressure against unlimited growth for the largest.

physics.soc-ph

Consensus and fragmentation in academic publication preferences

Academic publishing requires solving a collective coordination problem: among thousands of possible publication venues, which deserve a community's attention? A clear consensus helps scholars allocate attention, match submissions to appropriate outlets, and evaluate scholars for hiring and promotion. Yet preferences are not centrally coordinated--they emerge within each field over time. Here we ask whether all fields have arrived at similar solutions to this coordination problem, and whether preferences vary systematically with individual characteristics. Using an adaptive survey of 3,510 US tenure-track faculty yielding 163,002 pairwise comparisons across 8,044 venues, we show that fields occupy a wide spectrum of coordination. Economics, Chemistry, and Physics exhibit strong consensus, with respondents agreeing on elite venues and accurately predicting one another's choices. Computer Science and Engineering show fragmented preferences distributed across hundreds of outlets with minimal overlap. Within fields, preferences correlate with institutional prestige--faculty at elite institutions prefer higher-ranked venues--and with gender, as men prefer higher-ranked venues than women even after accounting for prestige and career stage. Scholars realize their personal preferences more successfully than their respective fields' consensus preferences, indicating that heterogeneity, not just selective hierarchy, shapes publishing outcomes. Journal Impact Factors explain only 64% of preference choices, systematically undervaluing what fields actually prefer. These results quantify how publication preferences vary across the structural diversity of academic fields.

cs.DL

Scientific productivity as a random walk

The expectation that scientific productivity follows regular patterns over a career underpins many scholarly evaluations. However, recent studies of individual productivity patterns reveal a puzzle: the average number of papers published per year robustly follows the ``canonical trajectory'' of a rapid rise followed by a gradual decline, yet only about 20\% of individual productivity trajectories follow this pattern. We resolve this puzzle by modeling scientific productivity as a random walk, showing that the canonical pattern can be explained as a decrease in the variance in changes to productivity in the early-to-mid career. By empirically characterizing the variable structure of 2,085 productivity trajectories of computer science faculty at 205 PhD-granting institutions, spanning 29,119 publications over 1980--2016, we (i) discover remarkably simple patterns in both early-career and year-to-year changes to productivity, and (ii) show that a random walk model of productivity both reproduces the canonical trajectory in the average productivity and captures much of the diversity of individual-level trajectories, including the lognormal distribution of cumulative productivity observed by William Shockley in 1957. We confirm that these results generalize across fields by fitting our model to a separate panel of 22,952 faculty across 12 fields from 2011 to 2023. These results highlight the importance of variance in shaping individual scientific productivity, opening up new avenues for characterizing how systemic incentives and opportunities can be directed for aggregate effect.

stat.AP

Misere Connect Four is Solved

Connect Four is a two-player game where each player attempts to be the first to create a sequence of four of their pieces, arranged horizontally, vertically, or diagonally, by dropping pieces into the columns of a grid of width seven and height six, in alternating turns. Misere Connect Four is played by the same rules, but with the opposite objective: do not connect four. This paper announces that Misere Connect Four is solved: perfect play by both sides leads to a second-player win. More generally, this paper also announces that Misere Connect $k$ played on a $w \times h$ board is also solved, but the outcome depends on the game's parameters $k$, $w$, and $h$, and may be a first-player win, a second-player win, or a draw. These results are constructive, meaning that we provide explicit strategies, thus enabling readers to impress their friends and foes alike with provably optimal play in the misere form of a table-top game for children.

math.CO

A model for efficient dynamical ranking in networks

We present a physics-inspired method for inferring dynamic rankings in directed temporal networks - networks in which each directed and timestamped edge reflects the outcome and timing of a pairwise interaction. The inferred ranking of each node is real-valued and varies in time as each new edge, encoding an outcome like a win or loss, raises or lowers the node's estimated strength or prestige, as is often observed in real scenarios including sequences of games, tournaments, or interactions in animal hierarchies. Our method works by solving a linear system of equations and requires only one parameter to be tuned. As a result, the corresponding algorithm is scalable and efficient. We test our method by evaluating its ability to predict interactions (edges' existence) and their outcomes (edges' directions) in a variety of applications, including both synthetic and real data. Our analysis shows that in many cases our method's performance is better than existing methods for predicting dynamic rankings and interaction outcomes.

physics.soc-ph

Infectious disease surveillance needs for the United States: lessons from COVID-19

The COVID-19 pandemic has highlighted the need to upgrade systems for infectious disease surveillance and forecasting and modeling of the spread of infection, both of which inform evidence-based public health guidance and policies. Here, we discuss requirements for an effective surveillance system to support decision making during a pandemic, drawing on the lessons of COVID-19 in the U.S., while looking to jurisdictions in the U.S. and beyond to learn lessons about the value of specific data types. In this report, we define the range of decisions for which surveillance data are required, the data elements needed to inform these decisions and to calibrate inputs and outputs of transmission-dynamic models, and the types of data needed to inform decisions by state, territorial, local, and tribal health authorities. We define actions needed to ensure that such data will be available and consider the contribution of such efforts to improving health equity.

cs.CY

An Open-Source Cultural Consensus Approach to Name-Based Gender Classification

Name-based gender classification has enabled hundreds of otherwise infeasible scientific studies of gender. Yet, the lack of standardization, proliferation of ad hoc methods, reliance on paid services, understudied limitations, and conceptual debates cast a shadow over many applications. To address these problems we develop and evaluate an ensemble-based open-source method built on publicly available data of empirical name-gender associations. Our method integrates 36 distinct sources-spanning over 150 countries and more than a century-via a meta-learning algorithm inspired by Cultural Consensus Theory (CCT). We also construct a taxonomy with which names themselves can be classified. We find that our method's performance is competitive with paid services and that our method, and others, approach the upper limits of performance; we show that conditioning estimates on additional metadata (e.g. cultural context), further combining methods, or collecting additional name-gender association data is unlikely to meaningfully improve performance. This work definitively shows that name-based gender classification can be a reliable part of scientific research and provides a pair of tools, a classification method and a taxonomy of names, that realize this potential.

cs.SI

Subfield prestige and gender inequality in computing

Women and people of color remain dramatically underrepresented among computing faculty, and improvements in demographic diversity are slow and uneven. Effective diversification strategies depend on quantifying the correlates, causes, and trends of diversity in the field. But field-level demographic changes are driven by subfield hiring dynamics because faculty searches are typically at the subfield level. Here, we quantify and forecast variations in the demographic composition of the subfields of computing using a comprehensive database of training and employment records for 6882 tenure-track faculty from 269 PhD-granting computing departments in the United States, linked with 327,969 publications. We find that subfield prestige correlates with gender inequality, such that faculty working in computing subfields with more women tend to hold positions at less prestigious institutions. In contrast, we find no significant evidence of racial or socioeconomic differences by subfield. Tracking representation over time, we find steady progress toward gender equality in all subfields, but more prestigious subfields tend to be roughly 25 years behind the less prestigious subfields in gender representation. These results illustrate how the choice of subfield in a faculty search can shape a department's gender diversity.

cs.CY

Labor advantages drive the greater productivity of faculty at elite universities

Faculty at prestigious institutions dominate scientific discourse, with the small proportion of researchers at elite universities producing a disproportionate share of all research publications. Environmental prestige is known to drive such epistemic disparity, but the mechanisms by which it causes increased faculty productivity remain unknown. Here we combine employment, publication, and federal survey data for 78,802 tenure-track faculty at 262 PhD-granting institutions in the American university system between 2008--2017 to show through multiple lines of evidence that the greater availability of funded graduate and postdoctoral labor at more prestigious institutions drives the environmental effect of prestige on productivity. In particular, we show that greater environmental prestige leads to larger faculty-led research groups, which drive higher faculty productivity, primarily in disciplines with research group collaboration norms. In contrast, we show that productivity does not increase substantially with prestige for either faculty papers published without group members, nor group members themselves. The disproportionate scientific productivity of elite researchers is thus largely explained by their substantial labor advantage, indicating a more limited role for prestige itself in predicting scientific contributions.

cs.DL

Emergence of Hierarchy in Networked Endorsement Dynamics

Many social and biological systems are characterized by enduring hierarchies, including those organized around prestige in academia, dominance in animal groups, and desirability in online dating. Despite their ubiquity, the general mechanisms that explain the creation and endurance of such hierarchies are not well understood. We introduce a generative model for the dynamics of hierarchies using time-varying networks in which new links are formed based on the preferences of nodes in the current network and old links are forgotten over time. The model produces a range of hierarchical structures, ranging from egalitarianism to bistable hierarchies, and we derive critical points that separate these regimes in the limit of long system memory. Importantly, our model supports statistical inference, allowing for a principled comparison of generative mechanisms using data. We apply the model to study hierarchical structures in empirical data on hiring patterns among mathematicians, dominance relations among parakeets, and friendships among members of a fraternity, observing several persistent patterns as well as interpretable differences in the generative mechanisms favored by each. Our work contributes to the growing literature on statistically grounded models of time-varying networks.

cs.SI

The Dynamics of Faculty Hiring Networks

Faculty hiring networks-who hires whose graduates as faculty-exhibit steep hierarchies, which can reinforce both social and epistemic inequalities in academia. Understanding the mechanisms driving these patterns would inform efforts to diversify the academy and shed new light on the role of hiring in shaping which scientific discoveries are made. Here, we investigate the degree to which structural mechanisms can explain hierarchy and other network characteristics observed in empirical faculty hiring networks. We study a family of adaptive rewiring network models, which reinforce institutional prestige within the hierarchy in five distinct ways. Each mechanism determines the probability that a new hire comes from a particular institution according to that institution's prestige score, which is inferred from the hiring network's existing structure. We find that structural inequalities and centrality patterns in real hiring networks are best reproduced by a mechanism of global placement power, in which a new hire is drawn from a particular institution in proportion to the number of previously drawn hires anywhere. On the other hand, network measures of biased visibility are better recapitulated by a mechanism of local placement power, in which a new hire is drawn from a particular institution in proportion to the number of its previous hires already present at the hiring institution. These contrasting results suggest that the underlying structural mechanism reinforcing hierarchies in faculty hiring networks is a mixture of global and local preference for institutional prestige. Under these dynamics, we show that each institution's position in the hierarchy is remarkably stable, due to a dynamic competition that overwhelmingly favors more prestigious institutions.

physics.soc-ph

A guide to choosing and implementing reference models for social network analysis

Analyzing social networks is challenging. Key features of relational data require the use of non-standard statistical methods such as developing system-specific null, or reference, models that randomize one or more components of the observed data. Here we review a variety of randomization procedures that generate reference models for social network analysis. Reference models provide an expectation for hypothesis-testing when analyzing network data. We outline the key stages in producing an effective reference model and detail four approaches for generating reference distributions: permutation, resampling, sampling from a distribution, and generative models. We highlight when each type of approach would be appropriate and note potential pitfalls for researchers to avoid. Throughout, we illustrate our points with examples from a simulated social system. Our aim is to provide social network researchers with a deeper understanding of analytical approaches to enhance their confidence when tailoring reference models to specific research questions.

cs.SI

Community Detection in Bipartite Networks with Stochastic Blockmodels

In bipartite networks, community structures are restricted to being disassortative, in that nodes of one type are grouped according to common patterns of connection with nodes of the other type. This makes the stochastic block model (SBM), a highly flexible generative model for networks with block structure, an intuitive choice for bipartite community detection. However, typical formulations of the SBM do not make use of the special structure of bipartite networks. Here we introduce a Bayesian nonparametric formulation of the SBM and a corresponding algorithm to efficiently find communities in bipartite networks which parsimoniously chooses the number of communities. The biSBM improves community detection results over general SBMs when data are noisy, improves the model resolution limit by a factor of $\sqrt{2}$, and expands our understanding of the complicated optimization landscape associated with community detection tasks. A direct comparison of certain terms of the prior distributions in the biSBM and a related high-resolution hierarchical SBM also reveals a counterintuitive regime of community detection problems, populated by smaller and sparser networks, where nonhierarchical models outperform their more flexible counterpart.

physics.soc-ph

Explaining Gender Differences in Academics' Career Trajectories

Academic fields exhibit substantial levels of gender segregation. To date, most attempts to explain this persistent global phenomenon have relied on limited cross-sections of data from specific countries, fields, or career stages. Here we used a global longitudinal dataset assembled from profiles on ORCID.org to investigate which characteristics of a field predict gender differences among the academics who leave and join that field. Only two field characteristics consistently predicted such differences: (1) the extent to which a field values raw intellectual talent ("brilliance") and (2) whether a field is in Science, Technology, Engineering, and Mathematics (STEM). Women more than men moved away from brilliance-oriented and STEM fields, and men more than women moved toward these fields. Our findings suggest that stereotypes associating brilliance and other STEM-relevant traits with men more than women play a key role in maintaining gender segregation across academia.

cs.SI

Control of excitable systems is optimal near criticality

Experiments suggest that cerebral cortex gains several functional advantages by operating in a dynamical regime near the critical point of a phase transition. However, a long-standing criticism of this hypothesis is that critical dynamics are rather noisy, which might be detrimental to aspects of brain function that require precision. If the cortex does operate near criticality, how might it mitigate the noisy fluctuations? One possibility is that other parts of the brain may act to control the fluctuations and reduce cortical noise. To better understand this possibility, here we numerically and analytically study a network of binary neurons. We determine how efficacy of controlling the population firing rate depends on proximity to criticality as well as different structural properties of the network. We found that control is most effective - errors are minimal for the widest range of target firing rates - near criticality. Optimal control is slightly away from criticality for networks with heterogeneous degree distributions. Thus, while criticality is the noisiest dynamical regime, it is also the regime that is easiest to control, which may offer a way to mitigate the noise.

q-bio.NC

Community detection, link prediction, and layer interdependence in multilayer networks

Complex systems are often characterized by distinct types of interactions between the same entities. These can be described as a multilayer network where each layer represents one type of interaction. These layers may be interdependent in complicated ways, revealing different kinds of structure in the network. In this work we present a generative model, and an efficient expectation-maximization algorithm, which allows us to perform inference tasks such as community detection and link prediction in this setting. Our model assumes overlapping communities that are common between the layers, while allowing these communities to affect each layer in a different way, including arbitrary mixtures of assortative, disassortative, or directed structure. It also gives us a mathematically principled way to define the interdependence between layers, by measuring how much information about one layer helps us predict links in another layer. In particular, this allows us to bundle layers together to compress redundant information, and identify small groups of layers which suffice to predict the remaining layers accurately. We illustrate these findings by analyzing synthetic data and two real multilayer networks, one representing social support relationships among villagers in South India and the other representing shared genetic substrings material between genes of the malaria parasite.

cs.SI

A physical model for efficient ranking in networks

We present a physically-inspired model and an efficient algorithm to infer hierarchical rankings of nodes in directed networks. It assigns real-valued ranks to nodes rather than simply ordinal ranks, and it formalizes the assumption that interactions are more likely to occur between individuals with similar ranks. It provides a natural statistical significance test for the inferred hierarchy, and it can be used to perform inference tasks such as predicting the existence or direction of edges. The ranking is obtained by solving a linear system of equations, which is sparse if the network is; thus the resulting algorithm is extremely efficient and scalable. We illustrate these findings by analyzing real and synthetic data, including datasets from animal behavior, faculty hiring, social support networks, and sports tournaments. We show that our method often outperforms a variety of others, in both speed and accuracy, in recovering the underlying ranks and predicting edge directions.

physics.soc-ph

Robust entropy requires strong and balanced excitatory and inhibitory synapses

It is widely appreciated that well-balanced excitation and inhibition are necessary for proper function in neural networks. However, in principle, such balance could be achieved by many possible configurations of excitatory and inhibitory strengths, and relative numbers of excitatory and inhibitory neurons. For instance, a given level of excitation could be balanced by either numerous inhibitory neurons with weak synapses, or few inhibitory neurons with strong synapses. Among the continuum of different but balanced configurations, why should any particular configuration be favored? Here we address this question in the context of the entropy of network dynamics by studying an analytically tractable network of binary neurons. We find that entropy is highest at the boundary between excitation-dominant and inhibition-dominant regimes. Entropy also varies along this boundary with a trade-off between high and robust entropy: weak synapse strengths yield high network entropy which is fragile to parameter variations, while strong synapse strengths yield a lower, but more robust, network entropy. In the case where inhibitory and excitatory synapses are constrained to have similar strength, we find that a small, but non-zero fraction of inhibitory neurons, like that seen in mammalian cortex, results in robust and relatively high entropy.

q-bio.NC