SearcharxivSearch

arXiv subjects

Chad M. Topaz

Publications and source records attributed to Chad M. Topaz.

At least 19 recordsLinked to original sources

What exam scores can and cannot prove about unauthorized AI assistance: Evidence from a highly public classroom episode

In spring 2026, an economics professor at Brown University gave a take-home midterm and, after unusually high scores, made the final exam proctored. Among the 59 students who completed the course, average scores fell from 95.7 out of 100 to 48.8. The instructor attributed the drop to unauthorized use of generative AI on the midterm; others proposed test anxiety, a harder final, student withdrawals, and regression to the mean. The publicly released scores, one midterm-final pair per student, show two striking patterns: the correlation between a student's two scores is only 0.06, and individual changes range from a 4-point gain to a 100-point loss. Permutation tests find no statistically detectable association between students' midterm and final scores, yet would detect association of the strength the alternative explanations predict more than 90 percent of the time. Once model complexity is accounted for, a model in which a student's midterm carries no information about that student's final describes the data as well as any model that permits an association. Neither result rules out a weak association, but together they show that the data do not require any. In simulated classes where each student's two scores remain linked through that student's own proficiency, as the alternative explanations imply, the two patterns almost never appear together; they appear together regularly only when that link is nearly severed. More than one mechanism could have severed it: midterm answers that were not the students' own work would have done so, and so would a final testing substantially different material, with no assistance involved. The scores cannot distinguish these possibilities or identify which students, if any, used unauthorized assistance. Our study offers a framework for evaluating statistical evidence in future disputes over generative AI use in academic assessment.

stat.AP

Institutional Harm through Threshold Cascades

Can a population of people not individually inclined to harm others nonetheless produce harmful collective outcomes, purely because of the institutional structure they inhabit? Social scientists have long argued yes, but existing accounts are largely qualitative and provide no precise condition distinguishing safe institutions from unsafe ones. We develop a threshold cascade model in which agents have positive activation thresholds, harmful behavior is irreversible, and the institution exerts both standing pressure and peer influence along a weighted network. We give a necessary and sufficient condition, checkable from the institution's structure and its members' thresholds, for resistance to any shock up to a given size. The criterion extends to signed influence, in which some peer effects counteract harm, and yields a convex optimization formulation for least-cost repair. It also reveals a sharp frontier between functionality and safety. An institution can coordinate its members and remain safe if and only if the exposure that coordination creates stays below the weakest member's net threshold. A further tension arises when coordination requires responsiveness to peer influence, which can make it impossible to prevent the most exposed group from cascading. We then analyze a mean-field model of two groups differing in how easily their members are pushed into harm. When one group is unstable in isolation but the system is stable under full mixing, disproportionate within-group influence creates a sharp homophily threshold beyond which the harm-free state becomes unstable. In the model, identical treatment of both groups does not generally equalize their cascade robustness.

physics.soc-ph

Topological summaries of fingerprint ridge patterns carry identity information

Fingerprints are the most widely deployed biometric. Verifying whether two impressions come from the same finger typically relies on minutiae, small landmarks such as skin ridge endings and bifurcations. These landmarks are extracted through a multi-stage pipeline of image enhancement, skeletonization, minutiae detection, and alignment. We investigate an alternative: using topological data analysis to represent the full pattern of skin ridges and valleys directly, bypassing minutiae detection and the downstream matching pipeline. We apply persistent homology, a topological tool that tracks how loops in the ridge pattern form and fill in across spatial scales, producing multi-scale summaries of ridge geometry. We develop and compare a range of verification methods on a standard benchmark dataset, FVC2000 DB1. Even the simplest topological summaries, with no trained parameters, substantially outperform geometry-only baselines. A trained method achieves an AUC of 0.91, while an optimal-transport method excels at the strictest false-accept thresholds, suggesting they capture different aspects of the ridge pattern. Fusing these two approaches yields the best performance at every low false-accept threshold we examine. Our results establish that these topological summaries capture substantial fingerprint identity information, far more effective for verification than raw pixel-level geometry. Because the entire pipeline is openly specified, it offers a transparent complement to minutiae-based systems, and we provide a modular framework for constructing, evaluating, and combining topological verification methods.

cs.CV

A Null Model for Mapper Subtype Claims

The Mapper algorithm from topological data analysis constructs a graph summarizing the shape of a high-dimensional dataset, and groups of data points identified within this graph are widely interpreted as evidence of distinct subtypes. However, the covariance structure of the data alone can make such groups appear differentiated, even when no subtypes are present. Existing validation approaches do not account for this effect and thus cannot distinguish covariance artifacts from genuine subtypes. We propose a Gaussian null model that generates reference data matching the sample covariance matrix. We pair it with a test statistic that measures mean-level differentiation between communities. In an idealized setting, we prove that covariance geometry alone causes Mapper communities to differ in their average feature profiles, and we show that a simpler label-permutation baseline cannot detect this effect. Simulations confirm well-controlled Type I error under Gaussian data. We apply the framework to four published Mapper analyses spanning breast cancer gene expression, Congressional voting, NBA player performance, and lower-grade glioma genomics. In every case, once outlier singleton communities are accounted for, the observed differentiation does not exceed what the null produces at the α = 0.05 level. This result does not rule out subtypes in these datasets, but it does indicate that the observed structure is consistent with what covariance geometry alone can produce. Stronger evidence would be needed to support a subtype claim.

stat.ME

Vegetation Pattern Formation via Energy-Balance-Constrained Modeling

Vegetation in semi-arid environments self-organizes into striking spatial patterns -- bands, spots, labyrinths, and gaps -- with characteristic wavelengths on the order of tens to hundreds of meters. Existing reaction-diffusion models postulate nonlinearities and transport laws from qualitative physical reasoning, making it hard to distinguish essential structural features from artifacts of the chosen forms. Here we show how energy-balance and water-conservation principles can constrain the admissible model class before a specific closure is chosen. These constraints motivate a family of semilinear closures; an Euler--Lagrange representative yields a fourth-order vegetation equation coupled to quasi-steady water transport on a one-dimensional hillslope. Linear stability analysis identifies three instability mechanisms: classical water-mediated feedback, energy-balance spatial coupling, and water deflection by vegetation gradients. Their balance depends on terrain geometry. On slopes, the water-mediated coupling dominates and the model reproduces two empirical observations: pattern wavelength increases with aridity, and vegetation bands migrate uphill. On flat terrain, the energy-balance spatial coupling can drive instability independently. Numerical simulations confirm the linear predictions, and exploratory continuation reveals a narrow hysteresis region consistent with subcritical bifurcation.

nlin.PS

A Discrete-Time Model of the Academic Pipeline in Mathematical Sciences with Constrained Hiring in the United States

The field of the mathematical sciences relies on a continuous academic pipeline in which individuals progress from undergraduate study through graduate training and postdoctoral program to long term faculty employment. National statistics report trends in bachelor's, master's, and doctoral degree awards, but these data alone do not explain how individuals move through the academic system or how structural constraints shape downstream career outcomes. Persistent growth in postdoctoral appointments alongside relatively stable faculty employment indicates that degree production alone is insufficient to characterize workforce dynamics. In this study, we develop a discrete time compartmental model of the academic pipeline in the field of the mathematical sciences that links observed degree flows to latent population stocks. Undergraduate and graduate populations are reconstructed directly from nationally reported degree data, allowing postdoctoral and faculty dynamics to be examined under completion, exit, and hiring processes. Advancement to faculty positions is modeled as vacancy limited, with competition for permanent positions depending on downstream population size. Numerical simulations show that increases in degree inflow do not translate into proportional faculty growth when hiring is constrained by limited turnover. Instead, excess supply accumulates primarily at the postdoctoral stage, leading to sustained congestion and elevated competition. Sensitivity analyses indicate that long run workforce outcomes are governed mainly by faculty exit rates and hiring capacity rather than by degree production alone. These results demonstrate the central role of vacancy limited hiring in shaping academic career trajectories within the field of the mathematical sciences.

math.DS

Reliable Topology for Dynamic Data: Mathematical Foundations and Applications

Across many scientific domains, practitioners rely on coarse, discretized summaries to track the evolving structure of complex systems under noise, measurement error, and changing system size. Understanding when such summaries are reliable -- and when apparent robustness is illusory -- remains a fundamental challenge. Topological data analysis (TDA) provides a case study: Crocker diagrams track the number of topological features across spatial scale and time, and because they are computationally efficient and easy to interpret, they have been widely used for exploratory analysis, bifurcation detection, model selection, and parameter inference. Despite their popularity, Crocker diagrams have lacked rigorous stability guarantees ensuring robustness to small data distortions. We develop a conditional stability theory for Crocker diagrams constructed from evolving point clouds. Our main results include deterministic conditions guaranteeing exact invariance when pairwise distances are well separated from the diagram's discretization thresholds, together with bounds on how much the diagrams can change when these conditions fail. We also establish probabilistic stability guarantees under Gaussian noise and bounds on topological change caused by adding or removing points, scaling linearly with the number of modified points. We illustrate these results using two complementary examples: an analytically tractable breathing polygon model that reveals how stability thresholds depend on geometry, and a feasibility analysis of epithelial cell imaging data showing when bounded-change guarantees provide the appropriate robustness framework. Together, these results reveal a two-tier stability structure for coarse, discretized topological summaries: exact invariance under verifiable geometric separation conditions, and geometry-controlled bounded change otherwise.

math.AT

A dynamical model of the U.S. mathematics graduate degree pipeline

We present a latent-stock compartmental framework for modeling degree production systems when only completion flows, rather than enrollments, are observed. Applied to U.S.\ mathematics degrees from 1969 to 2017, the model treats master's and PhD populations as latent compartments -- unobserved state variables that are inferred indirectly because they generate the observed completion flows -- with time-varying routing fractions and completion hazards. Using information-criterion model comparison across a grid of specifications, we find strong support for smooth nonlinear time variation in routing fractions and hazards, while models with explicit international forcing are disfavored. The preferred model achieves a log-scale root mean squared error of approximately 0.036, corresponding to a typical multiplicative error of about 4\% in fitted degree counts, and highlights key structural shifts in the graduate pipeline: the master's pathway became increasingly central to PhD production through the late twentieth century before weakening, while direct bachelor's-to-PhD entry remained small but persistent. Estimated completion hazards for both degrees rise over time, indicating faster effective turnover in the graduate compartments. Methodologically, our main contribution is a latent stock dynamical approach that recasts linked degreecompletion time series as a coherent stock-flow system when intermediate enrollments are unobserved, making explicit both what features of pipeline dynamics are identifiable from completion data alone and what limitations such data impose.

math.DS

How Withheld Punishment Enables Authoritarian Persistence: An Evolutionary Dynamics Approach

Democratic backsliding is often framed as a contest between pro-democratic defenders and anti-institutional norm-breakers. That framing can miss a third behavior, a public that withholds punishment from norm-breakers while penalizing those who confront them. We study a minimal three-strategy evolutionary game, with institutional defenders, anti-institutional disruptors, and this non-punishing public evolving under replicator dynamics. We grant defenders a head-to-head advantage over disruptors and ask whether it guarantees their long-run success. It does not. Two payoff regimes, differing only in how the public and disruptors interact, produce two failure modes. In an exploitation regime, the public is harmed by disruptors yet withholds sanction, so the three strategies exhibit cyclic dominance. When the losses around the cycle outweigh the gains, every interior trajectory approaches a boundary heteroclinic cycle in which disruptors repeatedly resurge. In an accommodation regime, the public and disruptors each gain from their interaction. When the public's gain is large enough, every interior trajectory converges to a stable public-disruptor coalition that excludes defenders. A pro-democratic advantage is therefore not enough. Weak sanction and penalized confrontation can leave anti-institutional disruption recurring or entrenched.

physics.soc-ph

Capturing Dynamics of Time-Varying Data via Topology

One approach to understanding complex data is to study its shape through the lens of algebraic topology. While the early development of topological data analysis focused primarily on static data, in recent years, theoretical and applied studies have turned to data that varies in time. A time-varying collection of metric spaces as formed, for example, by a moving school of fish or flock of birds, can contain a vast amount of information. There is often a need to simplify or summarize the dynamic behavior. We provide an introduction to topological summaries of time-varying metric spaces including vineyards [19], crocker plots [56], and multiparameter rank functions [37]. We then introduce a new tool to summarize time-varying metric spaces: a crocker stack. Crocker stacks are convenient for visualization, amenable to machine learning, and satisfy a desirable continuity property which we prove. We demonstrate the utility of crocker stacks for a parameter identification task involving an influential model of biological aggregations [58]. Altogether, we aim to bring the broader applied mathematics community up-to-date on topological summaries of time-varying metric spaces.

cs.LG

Connecting the Dots: Discovering the "Shape" of Data

Scientists use a mathematical subject called 'topology' to study the shapes of objects. An important part of topology is counting the numbers of pieces and holes in objects, and people use this information to group objects into different types. For example, a doughnut has the same number of holes and the same number of pieces as a teacup with one handle, but it is different from a ball. In studies that resemble activities like "connect the dots", scientists use ideas from topology to study the shape of data. Data can take many possible forms: a picture made of dots, a large collection of numbers from a scientific experiment, or something else. The approach in these studies is called 'topological data analysis', and it has been used to study the branching structures of veins in leaves, how people vote in elections, flight patterns in models of bird flocking, and more. Scientists can take data on the way veins branch on leaves and use topological data analysis to divide the leaves into different groups and discover patterns that may otherwise be hard to find.

math.HO

Impacts of California Proposition 47 on Crime in Santa Monica, CA

We examine crime patterns in Santa Monica, California before and after passage of Proposition 47, a 2014 initiative that reclassified some non-violent felonies to misdemeanors. We also study how the 2016 opening of four new light rail stations, and how more community-based policing starting in late 2018, impacted crime. A series of statistical analyses are performed on reclassified (larceny, fraud, possession of narcotics, forgery, receiving/possessing stolen property) and non-reclassified crimes by probing publicly available databases from 2006 to 2019. We compare data before and after passage of Proposition 47, city-wide and within eight neighborhoods. Similar analyses are conducted within a 450 meter radius of the new transit stations. Reports of monthly reclassified crimes increased city-wide by approximately 15% after enactment of Proposition 47, with a significant drop observed in late 2018. Downtown exhibited the largest overall surge. The reported incidence of larceny intensified throughout the city. Two new train stations, including Downtown, reported significant crime increases in their vicinity after service began. While the number of reported reclassified crimes increased after passage of Proposition 47, those not affected by the new law decreased or stayed constant, suggesting that Proposition 47 strongly impacted crime in Santa Monica. Reported crimes decreased in late 2018 concurrent with the adoption of new policing measures that enhanced outreach and patrolling. These findings may be relevant to law enforcement and policy-makers. Follow-up studies needed to confirm long-term trends may be affected by the COVID-19 pandemic that drastically changed societal conditions.

stat.AP

An Unpublished Manuscript of John von Neumann on Shock Waves in Boostered Detonations

We report on an unpublished and previously unknown manuscript of John von Neumann and contextualize it within the development of the theory of shock waves and detonations during the nineteenth and twentieth centuries. Von Neumann studies bombs comprising a primary explosive charge along with explosive booster material. His goal is to calculate the minimal amount of booster needed to create a sustainable detonation, presumably because booster material is often more expensive and more volatile. In service of this goal, he formulates and analyzes a partial differential equation based model describing a moving shock wave at the interface of detonated and undetonated material. We provide a complete transcription of von Neumann's work and give our own accompanying explanations and analyses, including the correction of two small errors in his calculations. Today, detonations are typically modeled using a combination of experimental results and numerical simulations particular to the shape and materials of the explosive, as the complex three dimensional dynamics of detonations are analytically intractable. Although von Neumann's manuscript will not revolutionize our modern understanding of detonations, the document is a valuable historical record of the state of hydrodynamics research during and after World War II.

physics.hist-ph

Comparing demographics of signatories to public letters on diversity in the mathematical sciences

In its December 2019 edition, the \textit{Notices of the American Mathematical Society} published an essay critical of the use of diversity statements in academic hiring. The publication of this essay prompted many responses, including three public letters circulated within the mathematical sciences community. Each letter was signed by hundreds of people and was published online, also by the American Mathematical Society. We report on a study of the signatories' demographics, which we infer using a crowdsourcing approach. Letter A highlights diversity and social justice. The pool of signatories contains relatively more individuals inferred to be women and/or members of underrepresented ethnic groups. Moreover, this pool is diverse with respect to the levels of professional security and types of academic institutions represented. Letter B does not comment on diversity, but rather, asks for discussion and debate. This letter was signed by a strong majority of individuals inferred to be white men in professionally secure positions at highly research intensive universities. Letter C speaks out specifically against diversity statements, calling them "a mistake," and claiming that their usage during early stages of faculty hiring "diminishes mathematical achievement." Individuals who signed both Letters B and C, that is, signatories who both privilege debate and oppose diversity statements, are overwhelmingly inferred to be tenured white men at highly research intensive universities. Our empirical results are consistent with theories of power drawn from the social sciences.

math.HO

Spatiotemporal chaos and quasipatterns in coupled reaction-diffusion systems

In coupled reaction-diffusion systems, modes with two different length scales can interact to produce a wide variety of spatiotemporal patterns. Three-wave interactions between these modes can explain the occurrence of spatially complex steady patterns and time-varying states including spatiotemporal chaos. The interactions can take the form of two short waves with different orientations interacting with one long wave, or vice-versa. We investigate the role of such three-wave interactions in a coupled Brusselator system. As well as finding simple steady patterns when the waves reinforce each other, we can also find spatially complex but steady patterns, including quasipatterns. When the waves compete with each other, time varying states such as spatiotemporal chaos are also possible. The signs of the quadratic coefficients in three-wave interaction equations distinguish between these two cases. By manipulating parameters of the chemical model, the formation of these various states can be encouraged, as we confirm through extensive numerical simulation. Our arguments allow us to predict when spatiotemporal chaos might be found: standard nonlinear methods fail in this case. The arguments are quite general and apply to a wide class of pattern-forming systems, including the Faraday wave experiment.

nlin.PS

Analyzing Collective Motion with Machine Learning and Topology

We use topological data analysis and machine learning to study a seminal model of collective motion in biology [D'Orsogna et al., Phys. Rev. Lett. 96 (2006)]. This model describes agents interacting nonlinearly via attractive-repulsive social forces and gives rise to collective behaviors such as flocking and milling. To classify the emergent collective motion in a large library of numerical simulations and to recover model parameters from the simulation data, we apply machine learning techniques to two different types of input. First, we input time series of order parameters traditionally used in studies of collective motion. Second, we input measures based in topology that summarize the time-varying persistent homology of simulation data over multiple scales. This topological approach does not require prior knowledge of the expected patterns. For both unsupervised and supervised machine learning methods, the topological approach outperforms the one that is based on traditional order parameters.

math.AT

Diversity of Artists in Major U.S. Museums

The U.S. art museum sector is grappling with diversity. While previous work has investigated the demographic diversity of museum staffs and visitors, the diversity of artists in their collections has remained unreported. We conduct the first large-scale study of artist diversity in museums. By scraping the public online catalogs of 18 major U.S. museums, deploying a sample of 10,000 artist records comprising over 9,000 unique artists to crowdsourcing, and analyzing 45,000 responses, we infer artist genders, ethnicities, geographic origins, and birth decades. Our results are threefold. First, we provide estimates of gender and ethnic diversity at each museum, and overall, we find that 85% of artists are white and 87% are men. Second, we identify museums that are outliers, having significantly higher or lower representation of certain demographic groups than the rest of the pool. Third, we find that the relationship between museum collection mission and artist diversity is weak, suggesting that a museum wishing to increase diversity might do so without changing its emphases on specific time periods and regions. Our methodology can be used to broadly and efficiently assess diversity in other fields.

stat.AP

Assessing biological models using topological data analysis

We use topological data analysis as a tool to analyze the fit of mathematical models to experimental data. This study is built on data obtained from motion tracking groups of aphids in [Nilsen et al., PLOS One, 2013] and two random walk models that were proposed to describe the data. One model incorporates social interactions between the insects, and the second model is a control model that excludes these interactions. We compare data from each model to data from experiment by performing statistical tests based on three different sets of measures. First, we use time series of order parameters commonly used in collective motion studies. These order parameters measure the overall polarization and angular momentum of the group, and do not rely on a priori knowledge of the models that produced the data. Second, we use order parameter time series that do rely on a priori knowledge, namely average distance to nearest neighbor and percentage of aphids moving. Third, we use computational persistent homology to calculate topological signatures of the data. Analysis of the a priori order parameters indicates that the interactive model better describes the experimental data than the control model does. The topological approach performs as well as these a priori order parameters and better than the other order parameters, suggesting the utility of the topological approach in the absence of specific knowledge of mechanisms underlying the data.

q-bio.QM